MODELS
Unlimited Basic subscription: unlimited Qwen 3.8 27B for €14.95/month, with no per-token billing

Industries / 08 / Industry and science

Pharmaceuticals and biotechnology

From a sequence or SMILES to a hypothesis worth investigating.

Protein structure, sequence analysis and molecular properties for drug discovery. Explore ESMFold, ESM-2 and MoLFormer to investigate targets and compounds.

Connect predictions and representations to your research pipeline and your team’s experimental evidence.

Scientific data and documentation

From the target to candidates you want to study.

Structures, sequences, properties and literature: a model for each research stage.

Protein structure prediction

Generate 3D structures from sequences with ESMFold without a multiple sequence alignment search. Review confidence across regions before interpreting your target.

Sequence and variant analysis

Use ESM-2 embeddings to compare sequences and train predictors of properties or variant effects using data from your project.

Molecular properties and QSAR

Use SMILES representations as a basis for adapted property predictors. Define the property and evaluation data.

Literature on targets and compounds

Adapt BiomedBERT to extract entities and classify English biomedical publications. Connect findings with targets and compounds while retaining source references.

With QDivZero / Pharmaceuticals and biotechnology

A drug discovery pipeline with your data.

Combine protein sequences, SMILES molecules and biomedical literature. Evaluate each model’s deployment and compare its results with the evidence from your project.

One possible workflow

  1. Sequences, SMILES and literature
  2. Specialised model
  3. Structure or representation
  4. Experimental evaluation

Open-weight / Pharmaceuticals and biotechnology

Models for your industry

ESMFold for protein structures, ESM-2 for sequences, MoLFormer for molecular representations and BiomedBERT for biomedical literature. Each model serves a specific research task.

Protein structure

ESMFold

3D protein structures from sequences, with Transformers support and no multiple sequence alignment search.

Transformers · self-hosted deploymentView model

Protein sequences

ESM-2 650M

Protein sequence representations as a basis for task-adapted models.

Hugging Face referenceView model

Molecular representations

MoLFormer-XL

Molecular embeddings from SMILES and a basis for adapted property prediction.

Hugging Face referenceView model

Biomedical NLP · English

BiomedBERT

Encoder pretrained on PubMed and PMC for adapted biomedical text extraction or classification.

Hugging Face referenceView model

Pharmaceuticals and biotechnology

AI questions for Pharmaceuticals and biotechnology

How does ESMFold support pharmaceutical research?

ESMFold predicts a protein’s 3D structure from its sequence and can support structural analysis of a target. It needs no multiple sequence alignment search or external databases for inference. Review confidence across regions and compare predictions with available evidence; it does not calculate protein–ligand affinity.

How do ESMFold, ESM-2 and MoLFormer differ?

ESMFold returns a protein structure; ESM-2 produces sequence representations for adapted models; MoLFormer represents molecules from SMILES for tasks such as property prediction. These are complementary tools: choose according to your input data and the output your research needs.

Does MoLFormer generate molecules or predict properties directly?

The linked checkpoint is a molecular representation basis, not a molecule generator. You can use its embeddings or adapt a predictor using data for a specific property. For QSAR tasks, define labels, splits and evaluation appropriate to your project.

What data should I prepare to evaluate a drug discovery pipeline?

Prepare amino acid sequences for ESMFold and ESM-2, SMILES for MoLFormer and English biomedical text for BiomedBERT. Retain target, compound and source identifiers. For adapted models, separate training and evaluation data and hold out experimental evidence to check whether outputs are useful.

What does BiomedBERT offer for target and compound research?

BiomedBERT is an encoder pretrained on English PubMed and PMC text. You can adapt it for entity extraction or biomedical literature classification. Retain paper and passage references; this document task complements molecular prediction rather than predicting affinity.

How do I estimate protein model inference cost?

Sequence length, model, batch size and configuration affect processing time and memory. Test sequences representative of your project and measure both resources. For ESMFold, also check how inference settings affect performance before estimating capacity and cost for a larger batch.

Can I connect molecular results to a LIMS or electronic notebook?

Your pipeline can save predictions and representations through those tools’ available interfaces. Retain compound, target, experiment and model-version identifiers. Your application validates format and records outputs without confusing them with assay measurements.

Does ESMFold have Hugging Face support, and can I deploy it with QDivZero?

The facebook/esmfold_v1 checkpoint has an MIT licence and Transformers support through EsmForProteinFolding. It is not currently served by a Hugging Face Inference Provider: it requires self-hosted deployment. To discuss deployment with QDivZero, share dependencies, resources and input/output formats; compatibility is reviewed before defining a solution.

How do I validate structures and property predictions before scaling their use?

For structures, review confidence across regions and compare with available experimental structures. For molecular properties, evaluate on held-out compounds and check performance outside the training domain. Retain versions and configurations to compare runs; predictions guide research and require experimental comparison.

Choose a drug discovery task. Evaluate its model.

Share your target, input format and inference volume to discuss deployment compatibility.