MODELS
Unlimited Basic subscription: unlimited Qwen 3.8 27B for €14.95/month, with no per-token billing

COMPUTE & DEPLOYMENTOF MODELS.

Multi-provider compute to run your models.

QDivZero analyzes your model's requirements and looks for compatible capacity across different providers. It compares the available options and selects the right infrastructure for your deployment, without you having to manage servers or GPUs.

Not sure where to start?Deploy a model

ONE MODEL IS JUST THE BEGINNING.

Run text, image, video, audio, embedding, and multimodal models without setting up your own infrastructure. From Hugging Face models to your own models, QDivZero prepares everything needed to run them in production.

Price your deployment

Try new models, compare them, and take the ones that best fit your applications to production. All from the same environment, without having to build different infrastructure for each model.

HOW IT WORKS

A single way to deploy, regardless of the model.

QDivZero unifies the entire deployment process. The workflow stays the same even when you change models or the infrastructure behind them changes. You define the requirements you need, and QDivZero manages them across the infrastructure, so you can always deploy and connect your applications in the same way.

Start
01 —

Define your deployment

Select your model and define its performance, region, and availability requirements.

02 —

QDivZero finds capacity

It compares compatible capacity available across providers and selects the best option for your requirements.

03 —

Your deployment is ready

QDivZero prepares the environment, starts the model, and exposes it through an OpenAI-compatible API.

Deployment requirements

Configure how your model responds

Adjust context, concurrency and inference behavior to match the way you need to use it.
Selected modelQwen3.8-27B

Serving profile

Fast
Balanced
Agent
Validated profile
Context window
Configured
Concurrency
Optimized
Runtime
Ready

Ready for capacity matching

Want to know what your deployment would cost?Compare prices
Bee on an orange flower.

Configure each deployment based on how you plan to use your model.

Adjust inference behavior, define where you want to run it, and decide when it should be available. From serving configuration to region, capacity, or availability, you stay in control of the decisions that matter in production.

Launch a deployment

OPERATE YOUR AI DEPLOYMENT

Monarch butterfly on a leaf.

Control what happens once your model is in production.

Monitor GPU and CPU usage, memory, throughput, tokens and deployment health from one place. Adjust capacity, availability or scheduling as your workload changes, without rebuilding the deployment.

Open deployment dashboard

Qwen3.8-Flash-Next-NVFP4 instance

dep_qwen38_eu_01

RunningMonitoring
Resource usage
GPU usage72%
CPU usage34%
Memory67%
24.8000MTokens processed
Deployment healthHealthy
Throughput46 tok/s
Latency128 ms
Availability99.99%
Want to see what happens after your model is deployed?Explore monitoring

Everything QDivZero manages for you

Model deployment

Deploy public, Hugging Face, private, or fine-tuned models. QDivZero manages the infrastructure.

Multi-provider routing

The scheduler selects capacity across available providers for cost, availability, and latency.

Persistent storage

Mount persistent disks across workloads, share state between launches, and keep it ready for every run.

OpenAI-compatible API

Expose deployed models behind the standard OpenAI contract so existing clients keep working.

Scheduled operations

Use cron-based start and stop rules so workloads run when needed and stay off when they do not.

Security controls

Apply regional controls, verified runtime checks, and private-by-default operations for production workloads.

Observability

See processed tokens, CPU, RAM, GPU, and token throughput in one place to optimise usage and troubleshoot workloads.

Serverless

Start models on demand and stop them after a configurable idle period to avoid unused capacity.

FAQ about QDivZero deployments

Answers about deploying open-weight Hugging Face models, custom and fine-tuned models, GPU infrastructure costs, migration from OpenAI, and inference APIs.

How is QDivZero different from a GPU cloud or AI model API?

A GPU cloud gives you machines that you must configure and operate. A model API removes that work but usually limits you to a closed catalogue. QDivZero combines freedom and simplicity: you can choose from the full Hugging Face catalogue and launch the model through a stable API without managing the infrastructure.

What infrastructure does QDivZero manage when deploying an AI model?

QDivZero validates VRAM and runtime requirements, selects available GPU capacity, and configures the inference environment, storage, and availability. It also centralizes deployment monitoring so your team does not have to manage servers and GPUs.

Can I migrate an application compatible with the OpenAI API?

Yes. Every deployment provides an OpenAI-compatible endpoint so you can reuse your current SDKs, clients, and workflows. You can change the model connected to your application without rebuilding the integration or modifying its architecture.

How is the price of deploying an AI model calculated?

With QDivZero deployments, active compute is prorated by the second and deducted from your account balance. You do not pay for the number of users or the tokens processed by that deployment. In the same account, you can use Public Models, which deduct balance according to input and output tokens, and activate Unlimited Basic with Qwen 3.8 27B for a fixed monthly fee without using balance.

Can I deploy a custom, private, or fine-tuned model on QDivZero?

Yes. You can deploy models developed by your team, keep their weights private, or run fine-tuned versions on dedicated compute. A model does not need to be part of a public catalog to use QDivZero infrastructure.

How does deploying a custom model work?

You define the model and its requirements. QDivZero validates the required runtime and VRAM, finds compatible capacity across providers, and prepares the inference environment. When it is ready, you receive an endpoint for connecting the model to your application through an OpenAI-compatible API.

Why use open-weight Hugging Face models in a business?

Open-weight models can handle use cases such as RAG, classification, code generation, and internal assistants while giving you more control and freedom to choose. QDivZero makes them ready to deploy in a few steps so you can evaluate different options and select the best fit for each application.

Ready to deploy your model?

Choose what you want to run and define your requirements. QDivZero finds compatible capacity, prepares the environment, and keeps your deployment ready to use. Forget about managing AI infrastructure and deploy your first private inference endpoint.