> ## Documentation Index
> Fetch the complete documentation index at: https://docs.biohub.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# ESMFold2

> ESMFold2 is the successor to ESMFold that sets a new state of the art for single-sequence structure prediction.

ESMFold2 is the successor to ESMFold that sets a new state of the art for single-sequence structure
prediction and enables the generation of new functional proteins through searching the ESMC model's
latent space. The model predicts high-resolution, all-atom 3D structures of biomolecular complexes
directly from sequence, with optional multiple sequence alignment (MSA) input for enhanced accuracy
on challenging targets.

<Columns cols={2}>
  <Card title="Hugging Face" icon="https://mintlify.s3.us-west-1.amazonaws.com/biohub-8940d923/icons/hugging-face.svg" href="https://huggingface.co/collections/biohub/esmfold2-model-family" horizontal />

  <Card title="GitHub" icon="github" href="https://github.com/Biohub/esm" horizontal />

  <Card title="Paper" icon="https://mintlify.s3.us-west-1.amazonaws.com/biohub-8940d923/icons/arxiv.svg" href="https://www.biorxiv.org/content/10.64898/2026.06.03.729735v1" horizontal />

  <Card title="API Reference" icon="code" href="https://docs.biohub.ai/api/protein/models/logits" horizontal />
</Columns>

<h2 id="get-started">
  Get Started
</h2>

Get started with ESMFold2 with our Quickstart and Tutorials.

### Quickstart Guide

<Steps>
  <Step title="Install the `esm` Python package">
    ```python theme={null}
    pip install esm
    ```
  </Step>

  <Step title="Create an API key">
    [Generate an API key](https://biohub.ai/developer-console/api-keys) from your Biohub account.
    This API key manages your access to credits and tokens, and the term API key/token is often used
    interchangeably within documentation.
  </Step>

  <Step title="Connect to the Biohub Platform API">
    Call the inference client with the selected model of choice and replace `<your API token>` with
    your token name.

    ```python theme={null}
    from esm.sdk.forge import SequenceStructureForgeInferenceClient

    client = SequenceStructureForgeInferenceClient(model="esmfold2-fast-2026-05", url="https://biohub.ai", token="<your API token>")
    ```
  </Step>

  <Step title="Run your inference">
    Now you are ready to use your model. For examples of specific use cases, check out our
    [Tutorials](#tutorials).
  </Step>
</Steps>

<h3 id="tutorials">
  Model Tutorials
</h3>

<Columns cols={2}>
  <Card title="Folding with ESMFold2" icon="https://mintlify.s3.us-west-1.amazonaws.com/biohub-8940d923/icons/colab.svg" href="https://colab.research.google.com/github/Biohub/esm/blob/main/cookbook/tutorials/esmfold2.ipynb">
    Fold proteins in combination with DNA, RNA, and small-molecule ligands.
  </Card>

  <Card title="Binder design" icon="github" href="https://github.com/Biohub/esm/blob/main/cookbook/tutorials/binder_design.ipynb">
    Design antibodies and minibinders with high hit rates. Implements the protocol featured in our
    paper, which produced binders exhibiting nanomolar affinity, target specificity, and functional
    activity in laboratory assays.
  </Card>
</Columns>

<h2 id="model-details">
  Model Details
</h2>

For additional information, see the Hugging Face link.

### Model Card

<Accordion title="Cite this Model">
  [Language Modeling Materializes a World Model of Protein Biology](https://www.biorxiv.org/content/10.64898/2026.06.03.729735)

  ```bibtex theme={null}
  @misc{candido2026language,
    title  = {Language Modeling Materializes a World Model of Protein Biology},
    author = {Candido, Salvatore and Hayes, Thomas and Derry, Alexander and Rao, Roshan
              and Lin, Zeming and Verkuil, Robert and Wu, Bryan and Lee, Jin Sub
              and Bruguera, Elise S. and Keval, Jehan A. and Kopylov, Mykhailo
              and Pak, John E. and Wu, Wesley and Thomas, Neil and Mataraso, Samson
              and Hsu, Alvin and Trotman-Grant, Ashton C. and Fatras, Kilian
              and dos Santos Costa, Allan and Badkundri, Rohil and Ak{\i}n, Halil
              and Oktay, Deniz and Deaton, Jonathan and Montabana, Elizabeth
              and Sitwala, Hrishita and Yu, Yue and Wiggert, Marius
              and Carlin, Dylan Alexander and Goering, Anthony W. and Blazejewski, Tomasz
              and Sandora, McCullen and Hla, Michael and Jia, Tina Z.
              and Kloker, Leon H. and Sofroniew, Nicholas J. and Uehara, Masatoshi
              and Pannu, Jassi and Bachas, Sharrol and Liu, Daniel S.
              and Sercu, Tom and Rives, Alexander},
    year   = {2026},
    url    = {https://www.biorxiv.org/content/10.64898/2026.06.03.729735},
    note   = {Preprint}
  }
  ```
</Accordion>

<Tabs>
  <Tab title="Overview">
    | | |
    | - | - |
    | **Version** | 2026-05 |
    | **Architecture** | ESM representations that power a series of looped folding layers. A diffusion model projects pairwise representations to atomic resolution predictions. |
    | **Supported Modalities** | Sequence and structure |
    | **Training Data** | ESMFold2 was trained on sequences from the Protein Data Bank (PDB) and the AlphaFold DB (AFDB). |

    #### Intended Use

    ESMFold2 predicts high-resolution, all-atom 3D protein structures directly from amino acid
    sequences, with optional multiple sequence alignment (MSA) input. The outputs include
    comprehensive structural information including all-atom coordinates (backbone and side chains),
    confidence metrics, and optional distogram predictions for detailed analysis of predicted
    structures.

    #### Limitations & Risks

    The model predicts single static conformations and is not designed for modeling protein
    dynamics, conformational flexibility, or multiple conformations of the same protein. Outputs
    should be validated experimentally. Not intended for clinical or therapeutic applications
    without further validation.

    This model is released under the
    [MIT License](https://github.com/Biohub/esm/blob/main/LICENSE.md).
  </Tab>

  <Tab title="Details">
    | | |
    | - | - |
    | **Type of model architecture** | Hybrid architecture that combines the representational power of large-scale protein language models with an efficient structure prediction module. ESM representations that power a series of looped folding layers. A diffusion model projects pairwise representations to atomic resolution predictions. The architecture stages include a language model embedding stage which uses a frozen ESMC 6B parameter language model to generate rich single-sequence embeddings, a trunk of looped folding layers that refine pairwise and single representations, and a diffusion-based structure module that employs a diffusion process to generate all-atom 3D coordinates (backbone and sidechains) and uses an ODE (Ordinary Differential Equation) solver for efficient sampling. |
    | **Training datasets** | ESMFold2 is trained on all biomolecules in the Protein Data Bank (PDB) and 8.8M synthetic structures from the AlphaFold DB (AFDB). |
  </Tab>

  <Tab title="Variants">
    | Model | MSA Conditioning | Description | Release Date |
    | - | - | - | - |
    | esmfold2-fast-2026-05 | No | Inference optimized single-sequence structure prediction model | 2026-05 |
    | esmfold2-2026-05 | Yes | Flagship model, capable of either single-sequence or MSA conditioned structure prediction for improved accuracy on difficult targets | 2026-05 |
  </Tab>

  <Tab title="Usage">
    | | |
    | - | - |
    | **Primary Use Case** | ESMFold2 produces high-resolution, all-atom 3D protein folding structures directly from input amino acid sequences, with optional multiple sequence alignment (MSA) input for enhanced accuracy on challenging targets. The model outputs comprehensive structural information including all-atom coordinates (backbone and side chains), confidence metrics (pLDDT, pAE, pTM, iPTM), and optional distogram predictions. |
    | **Supported Input Modalities** | Protein sequences |
    | **Access** | Through [Biohub](#get-started) or [Hugging Face](https://huggingface.co/collections/biohub/esmfold2-model-family) |
    | **License** | This model is released under the [MIT License](https://github.com/Biohub/esm/blob/main/LICENSE.md). |
    | **Not recommended for** | Clinical diagnosis or treatment recommendations. Computational metrics do not replace wet-lab validation. Treat model outputs as machine-generated hypotheses that require further experimental validation, not as established biological facts. |
  </Tab>
</Tabs>

<h2 id="model-data">
  ESM Atlas Data
</h2>

The ESM Atlas is a map of 6.8 billion proteins covering the full breadth of life's biodiversity and
more than one billion predicted structures. ESM Atlas organizes protein space by the shared
representation space learned by ESMC, with high-resolution structures predicted by ESMFold2. Users
can download the dataset used to generate the Atlas from AWS S3 at no cost. The table below lists
the available datasets with their approximate sizes and the AWS CLI command to download each one.

If you are unsure where to start, **SAE Clusters** is the most manageable entry point for exploring
cluster-level organization. Learn more information about SAEs
[here](https://huggingface.co/biohub/esmc-SAE-overview). Download **All Data** only if you need the
complete set of sequences, structures, and features across all 6.8 billion proteins.

| Dataset | Description | Package Size | Download command |
| - | - | - | - |
| Sequences | Protein sequences (6.8B proteins) | 2.2 TB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/sequences/ /mydrive` |
| Structures | Protein structures (1B proteins) | 68.9 TB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/folds/ /mydrive` |
| SAE features | Per protein and per-residue feature vectors (6.8B proteins) | 306 TB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/sae/data_shards/ /mydrive` |
| SAE Clusters | Cluster-level organization based on SAE features (7.5M clusters) | 26 GB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/clusters/indexes/secondary/cluster_members/ /mydrive` |
| HMM Results | Predicted pfam and taxonomy (6.8B proteins) | 653 MB | `aws s3 cp --no-sign-request s3://esm-protein-atlas/v1/clusters/data/representative_proteins.parquet /mydrive/` |
| Protein\_to\_accession | Mapping of protein IDs to accession numbers (6.8B proteins) | 162 GB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/shared_indexes/ /mydrive` |
| Normalization | SAE feature normalization | 192 KB | `aws s3 cp --no-sign-request s3://esm-protein-atlas/v1/normalization/max_idf_log10.pkl /mydrive/` |
| All Data | Complete set of sequences, structures, features, and clusters | 377 TB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/ /mydrive` |
