MODELS
Unlimited Basic subscription: unlimited Qwen 3.8 27B for €14.95/month, with no per-token billing

Use cases / Audio and voice

Voice agents

Build AI voice agents that listen and help people complete tasks.

Turn spoken conversations into queries, answers and actions. Build voice agents for customer service, bookings and product features connected to your data and services.

With QDivZero, you can combine speech recognition, a tool-using LLM and audio generation. Choose the models for each part of the conversation and coordinate turns from your application.

From conversation to action

Connect spoken requests to useful actions.

Connect spoken conversations to data and tools. Help users resolve questions and complete tasks within your application.

Voice customer support

Answer by voice using documentation and support tools. Define clarification steps and a support route when a task needs further attention.

Bookings and appointments

Check availability and prepare bookings through conversation. Your application validates details and confirms the operation with the relevant tool.

Voice product interfaces

Offer a spoken way to complete specific tasks. Combine recognition and context, and define how to handle interruptions or unclear inputs.

Conversations connected to processes

Coordinate queries and actions from a voice request. Communicate actual results and request any missing information or confirmation.

Build it with QDivZero

Coordinate each step of a voice conversation.

Coordinate speech recognition, an LLM and audio synthesis. Your application manages turns, context and authorised tools so the conversation communicates what the workflow has completed.

Your application flow

  1. Listen

    ASR converts the spoken request into text.

  2. Understand

    The LLM uses instructions and conversation context.

  3. Act

    Your application executes tools requested by the agent.

  4. Respond

    TTS converts the agent response into speech.

  5. Continue

    Play the audio and coordinate the next conversation turn.

Compute

Open-weight / Hugging Face

You can start with…

Combine ASR to listen, an LLM to interpret and TTS to speak. Evaluate each component for language, quality and latency, and measure the complete conversation with your agent’s tools. Compare complete-conversation latency rather than only the LLM. Verify name and data recognition, tool selection, and spoken answer clarity using the tasks your agent needs.

Text and vision

Qwen3.8-27B

For conversation, code, and tasks combining text, images, and your own context.

View model on Hugging Face

Text to speech

Qwen3-TTS 1.7B CustomVoice

Generate speech with predefined voices and control its style through instructions.

View model on Hugging Face

Voice agents

Voice agents: frequently asked questions

What is an AI voice agent?

A voice agent combines speech recognition, an LLM, tools, and speech generation for conversations and tasks. It can listen to a request, query data or request actions, and respond aloud. Your application coordinates models and conversation logic.

How do I build a voice agent for customer support?

Connect ASR to transcribe requests, an LLM with support documentation and tools, and TTS for spoken responses. QDivZero deploys compatible models for each component. Your application provides the audio interface and coordinates turns.

Can a voice agent manage bookings and appointments?

Yes, when you connect APIs for availability queries and bookings. The LLM requests tools and your application executes authorized actions. Return the result to the agent so it can generate a spoken response about the operation.

Which models do I need to build a voice agent?

A modular flow needs ASR for listening, an LLM for interpretation and tools, and TTS for speaking. Start with Qwen3-ASR, a compatible LLM such as Qwen3.8, and Qwen3-TTS. Choose each component for language, quality, and latency.

How do I reduce voice assistant latency?

Measure speech recognition, LLM response, tool-call, and speech synthesis times. Compare models and capacity using representative conversations. Use streaming where models, runtimes, and your integration support it to start processing or playback earlier.

Can I integrate a voice agent into my app or calling system?

Connect model endpoints to your application and chosen audio or telephony service. Your integration manages input, playback, and conversation turns. QDivZero serves the models; connecting each channel is part of your application.

Can I build a voice agent in Spanish or other languages?

Yes, by combining speech recognition, language and synthesis models that support your chosen language. Evaluate accents, business vocabulary and pronunciation with representative conversations. Your application maintains context and handles interruptions or turn changes according to the experience you design.

How much does deploying an AI voice agent cost?

Calculate ASR, LLM and TTS compute, plus the tools and audio or telephony service you connect. Conversation duration and concurrency help size deployments. Check QDivZero pricing and measure latency and resource use across a complete dialogue.

Ready to build your own AI voice agent?

Connect speech recognition, language models and audio synthesis with your business tools. Deploy compatible components in QDivZero and build a conversation experience that helps users complete tasks.