MODELS
Unlimited Basic subscription: unlimited Qwen 3.8 27B for €14.95/month, with no per-token billing

Use cases / Audio and voice

Voice generation

Turn text into AI speech and give your product a voice.

Generate spoken answers, narration and audio from your text. Text-to-speech models, or TTS, add a listening experience to assistants, applications and content tools.

With QDivZero, you can deploy voice generation models and integrate them through an API. Choose a variant with the languages, voices and style controls your product needs.

Give your application a voice

Generate speech dynamically from text.

Give your application’s content and answers a voice. Turn text into audio with a compatible TTS model.

Assistants that speak

Connect chatbot answers to speech synthesis. Adapt length and pronunciation so users can listen clearly.

Text and script narration

Turn scripts into narration with available voices. Review names, numbers and pacing before incorporating audio into the final piece.

Speech inside your product

Generate audio dynamically through an API and prepare it for your interface. Check format and playback on the target channel or device.

Voice styles and design

Compare voices and style controls in your chosen variant. Evaluate pronunciation and pacing using your product’s languages and text.

Build it with QDivZero

From text to audio, with your choice of voice.

Prepare text and choose a voice compatible with your languages. Send the request to the TTS model and connect audio to your product’s player or channel.

Your application flow

  1. Text

    Send the message or script to convert to speech.

  2. Voice and style

    Select controls supported by the TTS model.

  3. Synthesis

    The model generates audio from the text.

  4. Playback

    Your application receives audio for playback or storage.

Compute

Open-weight / Hugging Face

You can start with…

Choose a TTS variant for predefined voices, style controls or voice design. Compare audio with your texts and each language, and check the runtime’s formats and streaming support. Evaluate voices through the channel users will hear. Compare your texts and devices alongside original audio and verify each variant’s supported style controls.

Text to speech

Qwen3-TTS 1.7B CustomVoice

Generate speech with predefined voices and control its style through instructions.

View model on Hugging Face

Text to speech

Qwen3-TTS 0.6B CustomVoice

A lighter alternative for generating speech in different languages.

View model on Hugging Face

Voice generation

Voice generation: frequently asked questions

How do I convert text to speech with AI?

Use a text-to-speech model, or TTS, to generate audio from a message or script. Send text and supported voice controls to the model endpoint. QDivZero prepares a compatible deployment for connection to your application.

Which open-weight models can generate speech?

Explore Qwen3-TTS and choose a variant for your voice and control requirements. CustomVoice provides predefined voices and style control; VoiceDesign lets you describe voice characteristics. Compare versions using your texts and compute requirements.

Can I generate speech in Spanish and other languages?

Yes, with TTS models supporting the required languages. Check each variant model card and compare pronunciation, rhythm, and naturalness using your texts. Quality can vary by language, selected voice, and generation instructions.

How do I choose a voice style or design an AI voice?

It depends on the model variant. Use style controls with predefined voices or describe vocal characteristics with a voice-design model. Specify the tone and speaking style your application needs and compare generated audio.

Can I add speech to an AI chatbot or assistant?

Yes. Connect your LLM text output to a TTS model and play the audio in your interface. Choose the answer-generation and speech models separately, adapting each component to the user experience.

How do I integrate a speech generation API into my application?

Deploy a compatible TTS model on QDivZero and send text from your server to its endpoint. Your application receives audio and prepares it for the player or channel. Formats and available streaming depend on the model and runtime.

How does TTS differ from speech recognition?

TTS, or text-to-speech, converts text into spoken audio. ASR, or speech recognition, does the reverse and converts audio into text. Combine both with an LLM to build an assistant that listens to a request and answers with speech.

How much does integrating a text-to-speech API cost?

It depends on the model, text length and how much audio you need to generate concurrently. Compare QDivZero pricing and measure quality and generation time with your content. For a voice assistant, also include the speech recognition and language models used in the conversation.

Ready to give your application a voice?

Turn text into audio for assistants, narration and product experiences. Deploy a compatible TTS model in QDivZero, choose the voice controls you need and connect speech synthesis with your application through an API.