MODELS
Unlimited Basic subscription: unlimited Qwen 3.8 27B for €14.95/month, with no per-token billing

Use cases / Image and video

Visual document extraction

Extract data from PDFs and invoices with AI that understands layout.

Turn invoices, forms and scanned documents into fields and tables your systems can use. Multimodal models interpret text and layout to connect labels, values and records.

With QDivZero, you can run visual document understanding models through an API. Prepare pages, define the information you need and connect validated results with your document workflow.

Beyond OCR

A document communicates through its structure.

Process documents where visual layout matters too. Turn tables, forms and scanned pages into structured information.

Invoices in different formats

Extract suppliers, line items and amounts from different layouts. Preserve relationships between headers, rows and totals to validate the invoice.

Forms and scanned documents

Connect labels and values in documents received as images. Flag empty or illegible fields rather than filling gaps with assumptions.

Tables and records

Turn PDF tables into rows and columns. Check headers, cells and continuity across pages before storing the data.

PDFs with visual structure

Retrieve information from columns, images and complex layouts. Retain page references to review the origin of each result.

Build it with QDivZero

From visual documents to fields in your system.

Process pages with a model that interprets their visual structure. Define output fields and tables; your application validates data and retains references to the original document.

Your application flow

  1. Document

    Receive the PDF, image, or scanned document.

  2. Preparation

    Convert pages into inputs supported by the model.

  3. Visual understanding

    The model connects text, fields, and document structure.

  4. Validation

    Check extracted data and keep source references.

  5. Structured data

    Send fields and tables to your destination system.

Compute

Open-weight / Hugging Face

You can start with…

Compare vision models with invoices, tables and forms from your workflow. Assess how they relate labels and values, input resolution and output format; keep page references to review extracted data. Compare complete documents with known fields. Evaluate text reading, column relationships, and structure preservation separately, especially for tables continuing across pages.

Text and vision

Qwen3.8-27B

For conversation, code, and tasks combining text, images, and your own context.

View model on Hugging Face

Multimodal reasoning

DeepSeek V4.1 Flash

Explore reasoning and understanding of text and images in multi-step tasks.

View model on Hugging Face

Visual document extraction

Visual document extraction: frequently asked questions

What is AI document data extraction?

Visual document extraction uses multimodal models to identify information in pages, images, and scanned files. Combine text and layout understanding to extract fields, tables, and relationships for your systems.

How do I extract data from scanned invoices or PDFs?

Prepare pages as images or other inputs supported by the model. Specify fields such as supplier, date, and amount, and define an output structure. Your application validates extracted data and keeps document references.

How does OCR differ from multimodal AI for documents?

OCR converts a text image into characters. A multimodal model can also use layout to interpret fields and connect information. Combine OCR and visual understanding based on document type, scan quality, and required data.

Can I extract PDF tables into JSON?

Yes. Describe columns, data types, and the expected JSON structure. Prepare pages for a compatible model and request extraction. Check rows, amounts, and table continuity in multi-page documents.

Do I need a separate template for each invoice or form?

Use a multimodal model to identify fields from instructions without a fixed template for every design. Accuracy depends on documents and the model. Evaluate real formats and add examples to guide extraction where needed.

How do I automate document processing through an API?

Connect a compatible multimodal model deployed on QDivZero to your document workflow. Your application prepares files, requests extraction, validates results, and sends data to the management system. Size batches according to volume and deployed capacity.

How do I process PDFs with multiple pages?

Prepare pages according to the model’s supported inputs and limits. Your application can divide the work, retain page numbers and combine extracted fields into a single record. Check consistency across pages when a table, line item or amount continues elsewhere in the document.

How do I reduce errors when extracting invoice amounts and tables?

Define fields, types and the structure of each row. Check totals, taxes, dates and references using application validation rules, and route uncertain cases for review. Evaluate the model with the formats and scan quality your business actually receives.

Ready to turn your PDFs into data for your systems?

Automate visual extraction from invoices, forms and scanned documents. Deploy a compatible model in QDivZero, define the fields and tables you need and integrate validated results with your application.