MODELS
Unlimited Basic subscription: unlimited Qwen 3.8 27B for €14.95/month, with no per-token billing

Use cases / Audio and voice

Audio analysis

Analyze audio with AI and find information in sound.

Classify recordings, detect acoustic events and ask questions about audio content. Use specialized models to extract information from speech and signals that a transcript does not capture.

With QDivZero, you can deploy compatible audio understanding models or your own task-specific models. Connect their labels, events and answers with your application’s processes through an API.

Understand the audio itself

There is information in the sound as well as the words.

Turn recordings into information about sounds and events. Choose models with the categories and outputs your application needs.

Sound classification

Organise recordings using your environment’s categories. Retain audio and labels to review ambiguous or unknown sounds.

Acoustic event detection

Identify relevant patterns and events in a recording. If you need their timing, choose a model that supports those outputs.

Speech and audio understanding

Query recordings with audio-understanding models. Provide specific questions and retain the file to verify results.

Models trained for your environment

Deploy a private or fine-tuned model compatible with your signals. Connect outputs through an API and evaluate them under real capture conditions.

Build it with QDivZero

From an audio signal to information for your process.

Select a model with the categories or outputs your task requires. Send the recording, validate the result and retain its reference to connect it to your workflow.

Your application flow

  1. Audio

    Prepare the recording or signal in the required format.

  2. Task

    Define sounds, events, or questions to identify.

  3. Model

    The specialist model analyzes the signal directly.

  4. Information

    Use labels, events, or answers in your application.

Compute

Open-weight / Hugging Face

You can start with…

An audio understanding model and a specialized detector have different goals. Choose based on whether you need answers about a recording, sound categories or specific events; compare outputs with audio from your environment. A recording-understanding model does not necessarily return event-detector outputs. Compare task, result formats, and input conditions against your application’s objective.

Audio understanding

Kimi-Audio-7B-Instruct

Understand recordings and work with speech and acoustic content.

View model on Hugging Face

Audio classification

Specialist Hugging Face models

Explore models trained for specific sounds, events, or signals.

Explore audio models

Your own models

Your own fine-tuned model

For a defined task with your own categories, examples, and training data.

Deploy your model

Audio analysis

Audio analysis: frequently asked questions

What is AI audio analysis?

AI audio analysis uses models to understand recordings, classify sounds, or identify acoustic events. It can work directly with signals and spoken information. Deploy compatible audio models on QDivZero and integrate them into your workflows through an API.

How does audio analysis differ from speech-to-text transcription?

Transcription recovers words from recordings. Audio analysis can also use sounds, events, and signal characteristics absent from text. Choose ASR, an audio-understanding model, or both based on the information you need.

Can AI models classify sounds and detect events?

Yes, with models trained for the required categories and events. Work with sound types, patterns, or specific signals. Check whether the model returns labels, temporal localization, or other outputs and evaluate it on recordings from your environment.

Which open-weight models can understand audio?

Explore Kimi-Audio for recording understanding and specialist Hugging Face models for classification or detection. Models support different tasks and outputs. Choose a version compatible with the runtime and audio you intend to process.

Can I deploy a model trained on my own recordings?

Deploy a private or fine-tuned model compatible with the runtime and available compute. Work with signals, labels, and examples from your environment. Check architecture, input formats, and outputs before connecting it to your application.

How do I integrate audio analysis through an API?

Deploy a compatible model on QDivZero, prepare audio input, and send it to the endpoint. Your application receives labels, events, or answers and connects them to your workflow. Match format, duration, and concurrency to model requirements.

How do I choose a sound classification model?

Define the sounds and labels your application needs and check the model’s training data. Evaluate it with recordings from your environment, including noise and similar categories. If you need to locate events in time, verify that the model provides that information rather than only a recording-level label.

Can I combine audio analysis, transcription and search?

Yes. Your application can use one model to analyze the signal and ASR to recover words. It can then index the transcript and metadata to search recordings by content or category. Each stage contributes different information you can bring together in one process.

Ready to turn sound into information for your business?

Integrate sound classification, audio understanding or event detection into your processes. Deploy a compatible Hugging Face model or your own model in QDivZero and connect its results with your application.