Technical guide · Molecular modelling
Boltz-2biomolecular structures and protein-ligand affinity
What Boltz-2 is, its applications and limitations. A technical guide to predicting 3D structures and protein-ligand affinity with QDivZero and its API.

In this guide
When you are investigating a protein and deciding which ligands to test, two questions need to be considered together. First, examine how each molecule might fit into the target. Then choose a binding estimate that helps prioritize compounds. Boltz-2 combines structural prediction with affinity estimation, allowing you to examine a proposed geometry alongside the outputs you will use for comparison.
This guide follows that process from defining the molecular system to reviewing its results. You will describe the protein and ligand, choose how to supply evolutionary information, configure the calculation and retrieve outputs through a Boltz-2 deployment on QDivZero. Each decision is connected to what it changes in the prediction and the evidence you need to retain for a meaningful comparison. The examples follow the integration based on Boltz 2.2.1.
You can first practise with the page demo, which lets you configure a request and explore a reference structure in 3D. When you are ready to predict your own molecules, the integration examples use your account URL and key to submit a job to your actual service.
What Boltz-2 is and what it is used for
Boltz-2 is an open-source artificial intelligence model for predicting biomolecular structures and protein-ligand binding affinity. To propose how a protein is organized or how it might interact with other molecules, you describe its components and obtain a three-dimensional structure. If your study also involves comparing small molecules against a protein, the affinity module adds binding estimates to that analysis. This capability extends Boltz-1's structural modelling, as explained in the official Boltz repository.
The scope of the task depends on the molecules you are studying. Structural modelling supports proteins, DNA, RNA and ligands, whose sequences or chemical identities are used to generate atomic coordinates that can be saved as mmCIF or PDB. Affinity estimation targets small molecules against proteins, so a system accepted for structural prediction is not necessarily suitable for affinity calculation. Check this compatibility before preparing a campaign, using the official prediction documentation.
Research applications and areas of work
In structural biology, a proposed structure helps you examine interfaces and decide which contacts to investigate experimentally. In drug discovery, you can combine pose inspection with binding estimates to select candidates from a virtual screen. Once you are working with an active compound series, the goal becomes understanding which modifications might improve its behaviour, a medicinal chemistry application. These tasks share the model but require different comparison criteria.
Consider two analogues you want to compare against the same protein. If you also change the protein sequence, MSA or sampling protocol, it becomes harder to attribute a difference to the chemical modification. In this guide, we will keep those factors under control, inspect each ligand pose and choose the affinity output that answers the question. The result helps formulate and prioritize hypotheses for the laboratory. Clinical efficacy, toxicity and pharmacokinetic properties require additional assessments.
Project, authors and official sources
Before incorporating the model into a research workflow, identify its provenance and the version you will run. The Boltz team at MIT Jameel Clinic developed Boltz-2 with Recursion and announced it on 6 June 2025. The official announcement identifies the team and project objective, while the research paper, with DOI 10.1101/2025.06.14.659707, describes its evaluation.
Code and weights are distributed under the MIT license, permitting academic and commercial use subject to its terms. Consult the repository license and model files on Hugging Face to document these components of your experiment. GPU infrastructure is billed separately from the model license.
The examples use Boltz 2.2.1, the software package version in the described integration. Boltz-2 identifies the model, so retaining that name alone is insufficient to repeat a calculation. Record the software version and the specific weight revision as well.
Limitations stated by the developers
Before applying the workflow to a new target, consider the limitations in section 6 of the original paper. Structural prediction can struggle with large complexes and major conformational changes. Affinity depends on the pose and does not explicitly account for ions, water or multimeric partners, while cropping can exclude relevant interactions. Generalization between assays and to proteins or molecules far from training requires further study. Conformational diversity learned from simulations also remains partial.
These limitations help define which systems to investigate. The affinity documentation permits requesting affinity for one small molecule and warns that estimates are unreliable against DNA, RNA or a cofactor. Ligand size also requires review. The developers advise against molecules significantly larger than 56 atoms, although the parser accepts up to 128, counting heavy atoms and hydrogens retained by RDKit's RemoveHs. The Boltz 2.2.1 parser records both thresholds and the warning.
As a practical recommendation, begin with a known system and experimental reference results. This case helps identify sequence, chemical preparation, MSA or pose problems before interpreting new compounds. The following sections explain how to provide that information and distinguish model capabilities from the limits applied by the QDivZero API.
Understand what Boltz-2 calculates
To study a ligand's position relative to a protein, you first need a proposal for their joint geometry. Boltz-2 generates it from the amino acid sequence, ligand chemical identity and additional information you supply. Prediction uses a diffusion process that successively refines initially noisy coordinates into a possible structure of the system.
Since a single proposal may leave uncertainty about chain or ligand orientation, you can request several samples and compare their geometries. Each sample provides a possible structure, accompanied by metrics that help assess model confidence. This review can reveal uncertain regions or interfaces, although plausible geometry and high confidence do not establish that the complex exists under your assay conditions.
If you also need to prioritize compounds, the affinity module provides two outputs with different objectives. One helps distinguish potential binders from non-binding controls, while the other is intended for comparing active compounds and their modifications. Later, we will choose between them according to whether you are conducting an initial screen or working with a known series. In QDivZero, this calculation requires proteins and exactly one small-molecule ligand chain. The platform rejects affinity requests containing DNA, RNA or multi-residue CCD ligands.
Prepare the experiment and deployment
Define the molecular system you want to study
The first decision is to specify the comparison you want to make. To examine a ligand's position and orientation, known as its pose, represent the molecules involved in that complex. To compare two analogues, prepare one system per compound and retain the same protein and protocol. This lets you relate changes in structure or affinity to the chemical modification under investigation.
That comparison begins with the protein sequence. Identify the species, isoform and experimental construct, meaning the version of the protein used in the assay. An experiment involving a truncated domain, variant, tags or modifications may not match the full UniProt sequence. Do not assume the sequence in a Protein Data Bank structure is identical either, because it may contain truncations, mutations or missing residues. Review these differences so the document describes the system you actually want to model.
Apply the same approach to the ligand. Retain its project identifier and check structure, stereochemistry, charge and salt representation before encoding it. Even if Boltz can parse a SMILES string, that string may represent a different chemical species from the one relevant to your assay. Recording how each molecule was prepared helps you later assess whether a difference comes from the compound or its preparation.
Create a Boltz-2 service
Once the system is defined, you need compute capacity to run the model. In the QDivZero deployment wizard, select boltz-community/boltz-2 and review GPU and storage estimates against the size and composition of your molecules. Each instance uses a compatible NVIDIA GPU, but memory requirements also depend on the system you will predict.
To keep the same weights throughout a comparison, select a specific commit revision and record it with the requests. QDivZero prepares dependencies and checkpoints, the files containing those weights, in the runtime. You can therefore submit jobs from Developers or an API client without installing Boltz or CUDA on your computer.
Assign the deployment a serving name, such as fold, so you can direct requests to it. In the examples, this name becomes the value of model, although the instance UUID can also be used. The repository boltz-community/boltz-2 selects the model when creating the deployment, while the serving name selects the instance that will run it within your account.
To use that instance from a terminal, configure the base URL and an account API key with deployment access. Replace the example host with the one in your integration and keep the terminal open, since subsequent commands use these variables.
1export QDIV0_API_URL='https://api.example.com'
2export QDIV0_API_KEY='REPLACE_WITH_YOUR_ACCOUNT_API_KEY'
3When organizing experiments, account for dedicated deployment billing while capacity is active. Queued or running jobs prevent automatic idle shutdown so they can finish. If you need to stop the service manually, wait for jobs to complete, because stopping it can interrupt execution.
Build the molecular input
Separate the molecules from execution parameters
To compare configurations without accidentally changing the molecular system, the request separates what you will study from how you will calculate it. Describe molecules in input using the native Boltz version-1 schema, and retain that section when repeating the calculation with different parameters. Execution settings belong in options, and model directs the job to the service you just created. Supply alignment and reference structure files through attachments when needed.
The document below brings those pieces together for a protein identified as A and a ligand identified as B, whose affinity is requested. Save it as request.json for parameter review and API submission. When comparing your own compounds, retain a copy per system so every result can be connected to the request that produced it.
The example uses the streptavidin segment observed in PDB 1STP and biotin's stereochemical SMILES. Its 121 residues correspond to positions 13 to 133 in the deposition. This single-chain representation teaches the workflow, while the biological complex is a tetramer. Research on that assembly requires explicitly defining the relevant chains and construct, as well as preparing an appropriate MSA.
1{
2 "model": "fold",
3 "input": {
4 "version": 1,
5 "sequences": [
6 {
7 "protein": {
8 "id": "A",
9 "sequence": "AEAGITGTWYNQLGSTFIVTAGADGALTGTYESAVGNAESRYVLTGRYDSAPATDGSGTALGWTVAWKNNYRNAHSATTWSGQYVGGAEARINTQWLLTSGTTEANAWKSTLVGHDTFTKV",
10 "msa": "empty"
11 }
12 },
13 {
14 "ligand": {
15 "id": "B",
16 "smiles": "C1[C@H]2[C@@H]([C@@H](S1)CCCCC(=O)O)NC(=O)N2"
17 }
18 }
19 ],
20 "properties": [
21 {
22 "affinity": {
23 "binder": "B"
24 }
25 }
26 ]
27 },
28 "options": {
29 "sampling_steps": 50,
30 "recycling_steps": 3,
31 "diffusion_samples": 1,
32 "sampling_steps_affinity": 200,
33 "diffusion_samples_affinity": 5,
34 "max_msa_seqs": 8192,
35 "use_msa_server": false,
36 "use_potentials": false,
37 "affinity_mw_correction": false,
38 "msa_pairing_strategy": "greedy",
39 "output_format": "mmcif",
40 "seed": 42
41 }
42}
43Identify chains and repeated copies
To study a complex containing several chains, enter its complete composition so Boltz models the organization you are interested in. Each entity in sequences declares a type and content, using protein, dna or rna for polymers and ligand for non-polymeric molecules. Assign unique chain identifiers, since you will later use them to select the affinity ligand, map templates or define constraints.
When several chains share the same sequence, you can describe the entity once and declare its copies through a list of identifiers. For example, "id": ["A", "B"] introduces two chains of that protein. This changes complex stoichiometry and doubles the residues contributed by the entity, so the decision must match the assembly you intend to study. Use separate entities when sequences differ.
Once the system is complete, check its combined limits. QDivZero accepts up to 32 chain identifiers and 4096 polymer residues per request, counting proteins, DNA, RNA and every repeated copy. Counting one sequence is insufficient when you have declared multiple chains.
Describe the ligand and request affinity
For a comparison to represent the chemical change you are investigating, retain the ligand's identity in the input. If its structure is available as SMILES, enter it in smiles. If you use a Chemical Component Dictionary entry, identify it through ccd. Choose one representation per entity and check that it matches the molecule prepared for your study.
Including the ligand enables protein-ligand structural modelling, but requesting affinity also requires selecting it in properties. In the example, "binder": "B" connects that request to the ligand chain. To examine geometry alone, remove properties and keep the molecule in sequences, so it remains part of the structural system.
Before interpreting affinity, review the module's chemical preparation, which standardizes the ligand before processing. The native limit is 128 atoms, counting heavy atoms and hydrogens retained by RDKit's RemoveHs, and molecules above 56 atoms produce a warning for exceeding the training size limit. This count comes from the prepared molecule rather than SMILES length. For salts and fragment mixtures, retain the original representation and inspect warnings to check which identity was used. Details are in the Boltz 2.2.1 parser.
Choose and provide an MSA
After describing the molecules, you can supply evolutionary information to support protein modelling. Boltz uses a multiple sequence alignment, or MSA, which collects related sequences and reveals patterns of conservation and variation. For this information to be relevant, the alignment must correspond to the submitted sequence. When comparing ligands against the same protein, retaining that MSA keeps the evolutionary input constant throughout the comparison.
Run with a single sequence
For an initial integration check, use "msa": "empty" to run with a single sequence without depending on an external alignment service. The example document makes this choice. QDivZero displays a warning because omitting evolutionary information can reduce accuracy, so a successful job only confirms that the request was processed. Before using its structures for research, assess whether this mode suits your system or whether an MSA is needed.
Request automatic generation
If you need an alignment and do not yet have one, request generation from the service configured by the QDivZero operator. Remove msa from the protein and set options.use_msa_server to true. Both changes are required, because retaining "msa": "empty" still requests single-sequence mode even when the server is enabled.
Before submitting a campaign that depends on this service, check its availability with the following command. Provider authentication and address are configured on the platform, while your request only opts into using it for prediction.
1curl --fail-with-body --get \
2 "$QDIV0_API_URL/v1/biomolecular/capabilities" \
3 -H "Authorization: Bearer $QDIV0_API_KEY" \
4 --data-urlencode 'model=fold'
5Use msa_status together with the checked_at timestamp to decide whether to proceed or investigate the service. A state of reachable means the backend received an accessible response, while not_configured indicates missing configuration. For authentication_failed, unreachable or unavailable, review access before repeating the job. A state of unknown means the check cannot establish availability.
This query does not generate an alignment from the GPU instance, so connectivity or provider failures can still occur during execution. Automatic generation also sends the sequence to the configured MSA service. Consider this when choosing how to prepare project sequences, and supply alignments as attachments if you need to retain their exact content.
Attach a precomputed alignment
To repeat a comparison with the same evolutionary information, provide a precomputed MSA. In Developers, upload it through Attachments and reference its name in the protein's msa field. An .a3m file supplies a protein alignment. To pair sequences across several chains, use CSV with sequence and key columns, where a shared key relates rows from the different chain alignments.
Through the API, send the file content as UTF-8 text in attachments, since the server cannot read a path on your computer. The following code updates request.json to attach an actual file named protein.a3m and reference it from the protein. First check that its alignment corresponds exactly to the document's sequence.
1import json
2from pathlib import Path
3
4request_path = Path("request.json")
5request = json.loads(request_path.read_text(encoding="utf-8"))
6protein = request["input"]["sequences"][0]["protein"]
7protein["msa"] = "protein.a3m"
8request["options"]["use_msa_server"] = False
9request["attachments"] = [{
10 "name": "protein.a3m",
11 "content": Path("protein.a3m").read_text(encoding="utf-8")
12}]
13request_path.write_text(json.dumps(request, indent=2), encoding="utf-8")
14The name protein.a3m connects the molecular reference to the attachment QDivZero resolves. Preserve that correspondence and use lowercase extensions for correct file identification. The integration accepts up to 32 UTF-8 text attachments with a combined maximum size of 8 MiB. If the system contains several proteins, prepare all required alignments consistently, because supplied and automatic MSAs cannot be mixed in one request.
Configure the prediction budget
After fixing molecules and alignments, decide what you need to observe in the result. A reduced configuration can check that a request is processed. To study whether the pose or relative chain orientation varies between proposals, compare several samples. Boltz supports separate structural and affinity sampling settings, so connect each adjustment to that question before increasing compute.
Structural sampling
To explore several proposals for the same system, increase diffusion_samples, which controls how many structures are generated. Comparing them lets you examine whether ligand pose and relative chain placement are maintained. Agreement provides context for review, although it does not establish experimental correctness.
To adjust the work performed within each prediction, use sampling_steps and recycling_steps. The first controls diffusion steps, and the second controls refinement cycles of the model's internal representations. Increasing them adds compute and can change the result, so evaluate their effect on your system. Moving from 50 to 200 steps multiplies diffusion iterations by four, but not necessarily total runtime, which also includes preparation, model loading, MSA and affinity.
Affinity sampling
If your objective includes comparing binding estimates, also review sampling_steps_affinity and diffusion_samples_affinity. These control separate sampling, so reducing structural steps does not automatically reduce the affinity calculation. The example therefore combines 50 structural steps with 200 steps and 5 affinity samples.
During this phase, Boltz 2.2.1 uses 5 refinement cycles and parallelism of 1. QDivZero exposes no option to change that cycle count. Affinity samples contribute to their own calculation and are not additional downloadable structures in the structural prediction set.
The table lets you check API defaults and bounds before submitting the document. All numeric controls above require integers. Defaults apply when an option is omitted, whereas the Developers form starts with 50 structural steps, so review the complete request when switching between the form and API.
Developers profiles and additional options
To try the workflow without entering every parameter manually, choose Quick, which selects 50 structural steps, 3 refinement cycles and 1 sample while preserving each protein's chosen MSA mode. To try 200 structural steps, use Research, which retains 3 cycles and the single sample. Without supplied alignments, Research changes explicitly empty MSAs to automatic generation. When attachments exist, it preserves them and leaves the remaining MSAs empty to avoid an incompatible mixture. In that case, provide missing alignments before carrying out a study that depends on them.
Choose a profile for the specific changes you need and review the resulting document. Neither profile changes the affinity budget by itself, and its name does not certify scientific prediction quality. To compare several poses, adjust the structural sample count as well.
When memory limits the calculation, max_msa_seqs reduces the number of alignment sequences the model can use, at the cost of evolutionary information. With several proteins and automatic MSA, configure pairing through msa_pairing_strategy, whose default is greedy and alternative is complete. Record these choices to keep them constant between compounds.
To evaluate additional guidance for physically plausible poses, enable use_potentials during structural sampling. This option is off by default, adds compute and is disabled by Boltz during affinity prediction. To explore molecular-weight correction of affinity estimates, use affinity_mw_correction, also off by default. Evaluate its effect using project controls and apply the same criterion to every compound in the comparison.
To document the random component of the experiment, set seed to an integer between 0 and 4294967295. If omitted, the runtime generates a seed and retains it in metadata. Saving it helps describe the calculation, but repetition also requires recording weights, versions, molecules and MSAs, as discussed when archiving results.
Add reference structures and constraints
If you already have structural evidence, provide it to guide prediction towards the hypothesis you want to investigate. A template supplies reference coordinates for a protein chain. In Structure templates, select an attached CIF, mmCIF or PDB and review how its chains relate to those in your request before using it.
This mapping uses chain_id for the chain you are predicting and template_id for its counterpart in the template. For example, protein A in your document can correspond to chain X in a CIF. Check sequence and numbering, particularly for domains or truncated constructs, so the reference informs the intended chain. For PDB, also consult subchain conventions in the native template format.
To guide prediction towards the template, enable force and specify the allowed deviation in angstroms through threshold. This introduces a stronger structural hypothesis, so record why the reference is appropriate and which tolerance was chosen. Do not later treat satisfaction of that same condition as independent prediction validation.
If your information concerns contacts or a binding site, describe it in Constraints. For a covalent bond, bond references the participating atoms. To relate two residues or atoms, use contact. To identify binding-site elements for a chain, pocket connects them to the chain specified by binder. In each case, first convert experimental numbering into positions starting at 1 within the submitted sequence.
The contact and pocket constraints express distance through max_distance, accepting 4 to 20 angstroms with 6 as the default. Their force field enables guidance for the condition. Native covalent bonds require valid atom names and support canonical residues and CCD ligands. Choose every constraint from identifiable evidence or a hypothesis and retain it in the protocol so you can interpret what informed the final geometry.
Practise in Developers
With these decisions clear, build the request in Developers by selecting the Biomolecular task and your service. Form lets you define molecules, provide alignments and adjust the calculation. When finished, use Advanced JSON to check that the document retains the intended system and parameters. View code then generates the corresponding integration examples, allowing you to move from visual configuration to an API client.
The following demo lets you practise this workflow with one protein and one SMILES ligand. Change MSA mode, compare profiles and inspect how the request changes before starting the test deployment and submitting a job. Follow simulated queued and running states, and use View code to download request.json and inspect its submission command.
On completion, a viewer helps you practise structure inspection. Rotate the complex, switch between ribbons and bonds, focus on the ligand and download mmCIF or PDB coordinates. You can also use Open your CIF or PDB to examine files from an actual Boltz run. The viewer processes coordinates in the browser without uploading them to a server and does not calculate new structures, confidence or affinity. Both the viewer and example reference are served by this website so they do not depend on another page.
The reference is the experimental streptavidin with biotin complex PDB 1STP, determined by X-ray diffraction at 2.6 Å. The form uses its 121 observed protein residues and biotin so you can recognize the molecules in the viewer. It displays the deposited asymmetric unit, which differs from the tetrameric biological assembly. Request identifiers A and B are input labels, while the file preserves deposition identifiers, which can differ between native mmCIF labels and author identifiers. This molecular correspondence does not make the reference a Boltz prediction or its B factors pLDDT values.
The demo simulates states and displays a fixed reference. Changing molecules or requesting several samples leaves the example coordinates unchanged. To obtain your own structures and metrics, submit the job to your actual deployment. The demo does not validate chemistry or MSA service connectivity either.
For multiple entities, CCD ligands, attachments, templates or constraints, prepare the document on the platform. Its form supports these capabilities, while the demo editor rejects them because it cannot represent them completely.
Dedicated deployment
boltz-community/boltz-2
Stopped
The name fold is an example. Use your serving name or instance UUID in the model field. The Hugging Face repository identifier is used when creating the deployment.
Task Biomolecular
Prepare the inputs and start the deployment to submit a job.
Submit a job through the API
To integrate the calculation into your research workflow, send request.json to POST /v1/biomolecular/jobs using an account API key with service access. Prediction runs asynchronously because preparation and MSA generation may be needed before structures can be produced. The initial response therefore supplies an identifier for following execution and retrieving results afterwards.
Before submission, check that QDIV0_API_URL contains your integration's base URL without a trailing /v1, since the example appends that segment, and that QDIV0_API_KEY contains your key. Also generate an idempotency key for this request. Retaining it lets you retry the same submission if its response is lost and recover the original job.
1export QDIV0_IDEMPOTENCY_KEY="$(python -c 'import uuid; print(uuid.uuid4())')"
2
3curl --fail-with-body --silent --show-error \
4 "$QDIV0_API_URL/v1/biomolecular/jobs" \
5 -H "Authorization: Bearer $QDIV0_API_KEY" \
6 -H 'Content-Type: application/json' \
7 -H "Idempotency-Key: $QDIV0_IDEMPOTENCY_KEY" \
8 --data-binary @request.json \
9 --output job.json \
10 --write-out 'HTTP %{http_code}\n'
11When the API admits the request, the command saves an HTTP 202 response in job.json, whose id field contains the job UUID. Use that UUID to follow progress from this point. Admission confirms that the job can be monitored, but you still need to wait for completion before downloading outputs.
If the submission response is not received, repeat only the cURL command with the same file and key, without regenerating QDIV0_IDEMPOTENCY_KEY. The same normalized request and key recover the original job. Changing the body while retaining the key returns HTTP 409. To test another compound or configuration as a new experiment, use a new key and retain its relationship to that request.
Once the UUID is known, query GET /v1/biomolecular/jobs/{job_id}. The response lets you distinguish whether to keep waiting or inspect a terminal outcome, and includes timestamps and, when available, confidence and affinity summaries, metadata and artifacts, the list of downloadable files.
To interrupt a job, send a request to POST /v1/biomolecular/jobs/{job_id}/cancel and query its state again, because process termination can take time. Retrieve an available output through GET /v1/biomolecular/jobs/{job_id}/artifacts/{artifact_id} using the identifiers in artifacts. The following client automates waiting and downloading so you can incorporate these steps into a campaign.
Retrieve outputs from Python
When comparing multiple systems, retain each run's files with its identifier. The following client reads the UUID from job.json, queries state every two seconds and downloads artifacts when the job completes successfully. It uses only the Python standard library, so save it as collect.py and run python collect.py after receiving HTTP 202 admission.
The client continues the job you already created. If a connection fails, run it again to query the same UUID and retrieve its files without submitting another prediction.
1import json
2import os
3import shutil
4import time
5from pathlib import Path
6from urllib.request import Request, urlopen
7
8base_url = os.environ["QDIV0_API_URL"].rstrip("/")
9api_key = os.environ["QDIV0_API_KEY"]
10admission = json.loads(Path("job.json").read_text(encoding="utf-8"))
11job_id = admission["id"]
12output = Path("results") / job_id
13output.mkdir(parents=True, exist_ok=True)
14
15def get(path):
16 request = Request(
17 base_url + path,
18 headers={"Authorization": f"Bearer {api_key}"}
19 )
20 return urlopen(request, timeout=60)
21
22while True:
23 with get(f"/v1/biomolecular/jobs/{job_id}") as response:
24 job = json.load(response)
25 print(job["status"], flush=True)
26 if job["status"] in {"succeeded", "failed", "cancelled"}:
27 break
28 time.sleep(2)
29
30(output / "job.json").write_text(
31 json.dumps(job, indent=2), encoding="utf-8"
32)
33if job["status"] != "succeeded":
34 raise SystemExit(job.get("error") or job["status"])
35
36for artifact in job["artifacts"]:
37 filename = Path(artifact["filename"]).name
38 destination = output / f"{artifact['id']}_{filename}"
39 path = f"/v1/biomolecular/jobs/{job_id}/artifacts/{artifact['id']}"
40 temporary = destination.with_suffix(destination.suffix + ".part")
41 with get(path) as response, temporary.open("wb") as file:
42 shutil.copyfileobj(response, file)
43 temporary.replace(destination)
44 print(destination)
45Downloads are grouped in a directory named after the job UUID so you can review a run without mixing it with another. Each filename includes its artifact identifier to prevent overwrites between identically named files. Transfers use temporary files renamed only after completion, distinguishing an interrupted download from an available result. The final response is saved in results/<job_id>/job.json to connect each file to its original name and identifier.
Evaluate structures and affinity
With the files retrieved, return to the question that motivated the experiment. To compare two analogues, assess whether the proposed geometries are relevant and then choose the appropriate affinity estimate. Confidence metrics support the first step but do not replace pose inspection or establish experimental activity.
Inspect coordinates and interfaces first
Open each run's mmCIF in a viewer such as PyMOL or UCSF ChimeraX and check that the expected chains and ligand are present. If your tools require PDB, request it by setting options.output_format to pdb before calculation. Focus inspection on the region that answers your question, because a well-defined protein can coexist with an uncertain interface.
Inspect ligand pose, binding-site contacts, distances and possible steric clashes. If several samples were requested, compare ligand placement and chain orientation between proposals. This helps identify geometrical aspects needing further review before using affinity to rank compounds.
Interpret confidence metrics
To put these uncertainties in context, relate metrics to the part of the structure under inspection. Local confidence is summarized by complex_plddt, global organization by ptm, and interfaces by iptm. JSON can also provide metrics per chain and chain pair, useful for examining a particular interface alongside its geometry. Aggregated values in these families are normalized between 0 and 1.
For several proposals, confidence_score provides a ranking using 0.8 × complex_plddt + 0.2 × iptm, replacing iptm with ptm for a single chain. Use it to guide structure inspection without converting that ranking into a compound activity order. To examine predicted distance error, consult complex_pde and complex_ipde, expressed in angstroms. Lower values indicate greater confidence in distances, requiring a different interpretation from normalized metrics where higher values are favourable.
Choose the affinity output for your study
For candidate selection in an initial screen, use affinity_probability_binary to examine predicted binding probability. This output is intended to distinguish candidates from non-binding controls and can help prioritize assays. However, a value of 0.9 does not establish a 90% experimental success probability in your project.
If you are comparing active compounds in a series, such as the two analogues in the working example, use affinity_pred_value. It expresses the estimate on the scale y = log10(IC50 [µM]), which can be converted to concentration through IC50 [µM] = 10^y. This makes the magnitude of differences easier to read, as shown in the table.
For example, moving from -1 to -2 corresponds to a tenfold lower concentration on this scale. This relationship compares model estimates but does not provide an experimental IC50 measurement, a dissociation constant Kd or a binding free energy. IC50 depends on assay conditions, so preserve the distinction between numerical scale and measured activity when deciding which compound to validate.
When automating this comparison, also check which output you are reading. Suffixes 1 and 2 belong to the models in the affinity ensemble, the combination used to obtain final estimates. They identify neither two ligands nor structural sample count. The native output definitions help document this mapping in your pipeline.
Assess performance using project controls
To decide whether prioritization is useful for your target, include compounds with known activity and appropriate controls. Examine how the model ranks them, and retain cases where geometry or affinity disagrees with experimental evidence. This lets you assess the method in your system before applying it to new molecules. Final selection requires activity assays and, depending on the objective, selectivity, toxicity and pharmacokinetic studies.
Published results provide context for that evaluation. The Boltz-2 paper reports a target-averaged Pearson correlation of 0.62 on the OpenFE subset of FEP+, compared with 0.63 for OpenFE. That value describes association between predictions and experimental values in the dataset rather than a percentage of correct predictions. The computational advantage exceeding 1000 times over reference simulations also corresponds to the conditions in the original evaluation results, so it does not by itself determine your campaign's performance.
Troubleshoot jobs and organize campaigns
After reviewing the workflow with a known system, prepare a ligand library using one job per system and the same protocol. Retain the relationship between each compound, its request and the returned UUID, since it is needed both for comparing results and for locating a failure without confusing runs. If a job does not finish successfully, review its diagnostic before changing parameters or repeating it. The table lists checks for common problems.
When scheduling submissions, QDivZero admits up to eight nonterminal jobs per account and runs predictions serially on each instance. More simultaneous requests do not by themselves increase capacity, so use states to monitor pending and completed jobs. Queued jobs expire after six hours, and the default execution limit is one hour from dispatch to the runtime, configurable by the operator. Returned timestamps help follow a campaign, but elapsed time does not establish prediction progress as a percentage.
Retain a reproducible experiment
To later explain why two compounds obtained different results, archive both what was submitted and what was executed. Retain the request, original attachments, final response and all artifacts, together with the downloadable manifest.json included in a successful job. Metadata records effective options, seed, runtime versions, GPU and CUDA, checkpoint revision and input hashes. These digital fingerprints reveal content changes in files or requests when comparing runs.
Organize the archive by UUID and add your experiment and compound identifiers. This lets you recover the molecular composition, protocol and outputs supporting a conclusion. Download results within seven days of completion, the integration's retention period even after instance shutdown.
For automatic MSA, record the service and generation date, since content can change with the provider or its databases. The integration retains provenance but does not export the alignment. To repeat a comparison with exactly the same evolutionary input, use precomputed MSAs that can be archived as attachments.
The seed fixes part of the stochastic process but does not control changes in hardware, dependencies, weights or MSA databases. Reproducibility therefore requires retaining those factors and defining acceptable differences when repeating prediction. This record connects the comparison back to a scientific question and a protocol you can review.
Frequently asked questions
Do I need Boltz or CUDA installed on my computer?
A QDivZero deployment can be used through Developers or an API client. The model and its dependencies run on the GPU instance. The retrieval example in this guide uses only the Python standard library.
Does the MIT license make execution free?
The license permits academic and commercial use of the code and weights. GPU capacity is an infrastructure resource billed separately. Dedicated deployment cost depends on active capacity and how long you keep it allocated.
What should I send in the model field?
To direct the job to your instance, use the deployment serving name, such as fold, or its UUID. QDivZero uses that identifier to locate the service within your account. The repository boltz-community/boltz-2 selects the model when creating the deployment, while model selects where your request will run.
Why can Quick still take time with only 50 steps?
Choosing Quick reduces structural sampling, but the job still needs to prepare the system, load the model and obtain an MSA when applicable. If affinity is requested, that phase retains its own default budget of 200 steps and 5 samples. To assess which computation you reduced, review both structural and affinity parameters in the resulting request.
Is enabling use_msa_server enough for automatic MSA?
For the protein to use the service, enable use_msa_server and remove its msa field. Keeping it set to empty retains the instruction to run with a single sequence even when the server is enabled. Before submitting, check that the operator has configured the service and that the system does not mix supplied and automatic alignments.
Can I submit several ligands together to compare affinity?
To compare compounds, prepare one job per system with the same protein and protocol. Each QDivZero affinity request requires exactly one small-molecule ligand chain and protein targets. Mapping its UUID to the compound identifier lets you later collect the estimates and review the structure associated with each one.
Is a negative affinity_pred_value an error?
Negative values are valid because the scale uses the base-10 logarithm of IC50 expressed in micromolar. For example, -2 corresponds to 0.01 µM on this scale. Lower values correspond to lower predicted concentrations, but conversion does not turn the estimate into an experimental measurement.
What does high structural confidence establish?
Use confidence to guide geometry review, since it describes how certain the model is about particular aspects of its proposal. If your question concerns ligand binding, inspect pose and interface even when global metrics are high. Prioritizing the compound also requires interpreting affinity separately and testing the hypothesis experimentally. Structural confidence does not establish activity or safety.
Can it model systems with DNA or RNA?
You can include DNA and RNA in structural predictions, while the affinity module targets small molecules against proteins. QDivZero rejects affinity requests containing DNA or RNA. Choose the operation according to your system composition.
What if I do not receive the submission response?
Retain the original body and idempotency key. Retry with both to recover the same job. Once its UUID is known, query its state directly. A new key can create another prediction and additional compute cost.
Can I stop the deployment and download results later?
Completed job results remain available for seven days even after the instance stops. Download artifacts and the manifest during that period. Manually stopping an instance with an active job can interrupt execution.
How do I make compound comparisons reproducible?
To relate a difference to the chemical modification, retain the same protein construct, alignments, preparation and sampling protocol. Archive each request, its attachments and manifest.json to check what was executed. To keep evolutionary input exactly the same, use precomputed MSAs, because automatic MSA content is not exported and can change with the service or its databases. A fixed seed does not control those differences.