MODELS
Unlimited Basic subscription: unlimited Qwen 3.8 27B for €14.95/month, with no per-token billing

LLM MODELROUTING.

Turn multiple models into one system.

LLM Model Routing lets you work with multiple models from one integration. Your application connects to QDivZero instead of connecting directly to each model, keeping routing in an independent layer. This lets you expand or change the models you use without accumulating integrations and model-specific logic in your code.

Want to try it with your models?Configure your routing

DESIGN HOW THEY WORK TOGETHER.

Decide what role each model plays in your application and define how they should coordinate. One routing layer can combine general-purpose, specialised, or differently capable models depending on what you are building.

Plan the cost of your models

Routing decisions live in QDivZero instead of being spread throughout your application code. You can review and change how requests are distributed across your models from one central configuration.

HOW IT WORKS

ONE WAY TO CALL YOUR MODELS.

Configure the models you want to use in QDivZero and group them into a single route. Your application always calls the same endpoint through an OpenAI-compatible API, and QDivZero directs each request to one of those models according to the strategy you define.

Start

INTENT-BASED MODEL ROUTING

QDivZero identifies the intent of each request—such as summarising, coding, or analysing an image—and routes it to the model you assigned to that task. This lets you combine specialised models and use faster or more economical options when the most powerful model is not needed.

What it optimizes

01 —

Cost

Do not pay for capacity the request does not need. Reserve the most expensive models for tasks where they add real value and use more efficient options for the rest.

02 —

Latency

Not every response needs to wait for the heaviest model. Tasks that work with faster models can return sooner without forcing your whole application into the same performance profile.

03 —

Quality

Routing lets you keep higher-capability models for requests that genuinely need them instead of lowering the quality of the whole system to reduce costs.

Adapt every decision to the request.

The most capable model can also be more expensive or slower. LLM routing lets you use different models according to the balance you need between quality, cost, and latency instead of applying the same choice to every request.

Try LLM Routing

What LLM Model Routing includes

Intent-Based Routing

Detect the intent of each request and send it to the model you assigned to that type of task.

Failover Routing

Define alternative models in priority order to continue with another destination when the primary model cannot handle a request.

Model groups

Group multiple models within one route and manage how you want to use them from one place.

Routing priorities

Choose which model to use first and the order in which QDivZero should try the alternatives you configured.

OpenAI-compatible API

Keep the OpenAI API format so you can use your routes without creating a different integration for every model.

The right model for every request

Route each task to the model that offers the right balance of quality, speed, and cost instead of always using the same option.

FAQ about LLM Model Routing in QDivZero

Answers about intent-based routing, model failover, OpenAI compatibility, latency, and the cost of routing AI requests.

What is LLM Model Routing and when should I use it?

LLM Model Routing connects multiple AI models behind one endpoint and determines which destination handles each request. It is useful when an application combines tasks with different requirements, needs to reduce costs by reserving the most capable models for complex work, or needs alternatives when a model is unavailable.

What is the difference between intent-based routing and failover routing?

Intent-based routing identifies the goal of a request and sends it to the model configured for that task, such as summarisation, coding, or image analysis. Failover routing follows a priority order: it tries the primary model first and uses the next available destination when the previous one cannot respond.

Can I use LLM Model Routing with an OpenAI-compatible application?

Yes. The application calls an OpenAI-compatible endpoint and uses the routing name as the model. You can keep familiar SDKs and request formats while QDivZero applies the strategy and selects a destination behind that stable contract.

How does model routing affect cost and latency?

Intent-based routing adds an evaluation step to classify the request, so it can introduce additional latency and cost. In return, it can send simple tasks to faster or less expensive models. Ordered routing avoids semantic classification and directly applies the configured priorities and fallback routes.

Can I combine my own deployments and Public Models in one route?

Yes. QDivZero can build routes with compatible destinations backed by Compute and Public Models. Compute destinations deduct balance for active compute time, while Public Models deduct it according to input and output tokens. In the same account, you can also activate Unlimited Basic with Qwen 3.8 27B for a fixed monthly fee without using balance.

What happens when the primary model is unavailable?

When priorities and alternatives are configured, QDivZero tries the next available destination in the defined order. The application continues to use the same endpoint, so failover logic stays in the platform instead of being implemented in every client.

Make your models work as a system.

Configure your LLM Model Routing, distribute requests across different models, and adjust routing logic according to your application's needs.