MODELS
Unlimited Basic subscription: unlimited Qwen 3.8 27B for €14.95/month, with no per-token billing

Use cases / Data and knowledge

RAG

Build RAG applications with your company’s knowledge.

Connect AI models with private, specialized or up-to-date information. With RAG, your application retrieves relevant content from your data before generating each answer.

QDivZero brings together the components of your RAG architecture: embedding models, vector search and LLMs. Build company assistants, product copilots and search tools with your own context and source references.

The context your model needs

Give the model the information behind the answer.

Connect a language model to your sources. Retrieve relevant context before generating each answer.

Company knowledge

Answer questions about internal procedures and knowledge using company sources. Your application decides what each user can access.

Information that changes

Update sources as content changes. Retrieve current information for each query without retraining the LLM for every document change.

Specialist documentation

Scope answers to your domain’s documentation. Combine retrieval and filters to select the right context for each question.

Answers with sources

Retain references to retrieved passages. Show the documents supporting an answer so users can verify it.

Build it with QDivZero

From your data to the context your LLM needs.

Retrieve relevant source passages and send them to the LLM with the question. Your application retains references, updates the index and defines how to answer when information is missing.

Prepare your data

  1. Documents

    Split content into passages and keep source references.

  2. Embeddings

    Represent each passage using your chosen embedding model.

  3. Vector index

    Store vectors with content and metadata.

For each query

  1. Query

    Embed the question using the same embedding model.

  2. Retrieval

    Find relevant passages in your vector index.

  3. Reranking

    Reorder candidates when you need to refine the context.

  4. Generation

    Send the query and selected context to the LLM.

  5. Answer

    Return the result and keep references to the content.

Retrieval

Open-weight / Hugging Face

You can start with…

Choose embeddings to represent your data, a reranker if retrieval needs refinement and an LLM to generate answers. Evaluate the combination with your queries: RAG quality depends on both the retrieved content and the model that responds. Include unanswerable queries, organisational terms, and questions requiring multiple sources. Compare the entire chain using the same inputs: changing the language model does not fix retrieval that fails to supply necessary context.

Text and vision

Qwen3.8-27B

For conversation, code, and tasks combining text, images, and your own context.

View model on Hugging Face

RAG

RAG: frequently asked questions

What is RAG in artificial intelligence?

RAG stands for retrieval-augmented generation. This architecture searches your data and adds relevant information to an LLM context before generating an answer. Use it to build AI applications that query your own or recently updated documentation.

How do I build a RAG application with my own data?

Prepare and index content using an embedding model. For each query, retrieve relevant passages, select context, and send it to an LLM with the question. QDivZero provides models and Retrieval as components for your RAG application.

What is the difference between RAG and fine-tuning?

RAG provides information through retrieval for each query. Fine-tuning adapts model parameters using training examples. Start with RAG for changing documents, and consider fine-tuning for specific behaviors or tasks. The two approaches can be combined.

How can I improve RAG answer quality?

Review document preparation, passage size, and the embedding model. Compare retrieved content relevance and add reranking where it helps select context. Evaluate LLM instructions and answers using real queries.

Can I use RAG with open-weight models?

Yes. Build RAG with open-weight embedding, reranking, and text generation models. Deploy compatible Hugging Face models through QDivZero and combine them with Retrieval. Change the LLM while keeping your application retrieval logic.

What determines the cost of a RAG application?

Cost depends on deployed models, compute, document volume, and query load. Context length and reranking also affect the workload. Size each component using real traffic and review QDivZero pricing options.

How do I update a RAG knowledge base?

Reindex documents that change, keep their metadata and remove obsolete versions. You can keep the LLM and update the context your application retrieves without retraining it. If you change the embedding model, regenerate the index vectors so documents and queries remain compatible.

How do I integrate RAG into a chatbot or AI agent?

Use the question or agent task to query Retrieval. Send the retrieved passages to the LLM together with instructions and conversation context. Your application decides when to retrieve information, how to show sources and what to do when the documentation does not contain enough information to answer.

Ready to connect your AI with your own data?

Build a RAG application with your company’s documentation, products and knowledge. Choose your models, prepare an index with Retrieval and integrate contextual answers into your product with QDivZero.