> ## Documentation Index
> Fetch the complete documentation index at: https://docs.biohub.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# ESMC

> ESMC is the latest in the ESM family of protein language models, establishing a new frontier in representation learning for protein biology.

ESMC is the latest in the ESM family of protein language models, establishing a new frontier in
representation learning for protein biology. Trained on billions of evolutionary sequences, it
learns representations that reflect a mechanistic reduction of protein structure and function.

<Columns cols={2}>
  <Card title="Hugging Face" icon="https://mintlify.s3.us-west-1.amazonaws.com/biohub-8940d923/icons/hugging-face.svg" href="https://huggingface.co/collections/biohub/esmc-model-family" horizontal />

  <Card title="GitHub" icon="github" href="https://github.com/Biohub/esm/tree/main#esm-c-" horizontal />

  <Card title="Paper" icon="https://mintlify.s3.us-west-1.amazonaws.com/biohub-8940d923/icons/arxiv.svg" href="https://www.biorxiv.org/content/10.64898/2026.06.03.729735v1" horizontal />

  <Card title="API Reference" icon="code" href="https://docs.biohub.ai/api/protein/models/logits" horizontal />
</Columns>

<h2 id="get-started">
  Get Started
</h2>

Get started using ESMC with our Quickstart and Tutorials.

### Quickstart Guide

<Steps>
  <Step title="Install the `esm` Python package">
    ```python theme={null}
    pip install esm
    ```
  </Step>

  <Step title="Create an API key">
    [Generate an API key](https://biohub.ai/developer-console/api-keys) from your Biohub account.
    This API key manages your access to credits and tokens, and the term API key/token is often used
    interchangeably within documentation.
  </Step>

  <Step title="Connect to the Biohub Platform API">
    Call the ESM client with the selected model of choice and replace `<your API token>` with your
    token name.

    ```python theme={null}
    from esm.sdk.forge import ESMCForgeInferenceClient

    client = ESMCForgeInferenceClient(model="esmc-6b-2024-12", url="https://biohub.ai", token="<your API token>")
    ```
  </Step>

  <Step title="Run your inference">
    Now you are ready to use your model. For examples of specific use cases or to append a SAE onto
    your ESMC model, check out our [Tutorials](#tutorials).
  </Step>
</Steps>

<h3 id="tutorials">
  Model Tutorials
</h3>

<Columns cols={2}>
  <Card title="Embedding sequences with ESMC" icon="https://mintlify.s3.us-west-1.amazonaws.com/biohub-8940d923/icons/colab.svg" href="https://colab.research.google.com/github/Biohub/esm/blob/main/cookbook/tutorials/embed.ipynb">
    Embed protein sequences and explore how different transformer layers encode structural and
    functional information.
  </Card>

  <Card title="Zero-shot entropy and mutation analysis" icon="https://mintlify.s3.us-west-1.amazonaws.com/biohub-8940d923/icons/colab.svg" href="https://colab.research.google.com/github/Biohub/esm/blob/main/cookbook/tutorials/esmc_mutation_scoring.ipynb">
    Compute per-position entropy and log-likelihood ratios to identify constrained vs.
    mutation-tolerant sites.
  </Card>

  <Card title="Layer sweep for enzyme function classification" icon="https://mintlify.s3.us-west-1.amazonaws.com/biohub-8940d923/icons/colab.svg" href="https://colab.research.google.com/github/Biohub/esm/blob/main/cookbook/tutorials/esmc_layer_sweep.ipynb">
    Learn how to sweep all layers to find which one is best using enzyme classification as a task.
  </Card>

  <Card title="Understanding proteins with SAE features" icon="https://mintlify.s3.us-west-1.amazonaws.com/biohub-8940d923/icons/colab.svg" href="https://colab.research.google.com/github/Biohub/esm/blob/main/cookbook/tutorials/esmc_sae_feature_interpretation.ipynb">
    Extract and visualize sparse autoencoder features, rank by peak activation, and map activations
    onto 3D structures.
  </Card>

  <Card title="Fine-tuning ESMC" icon="https://mintlify.s3.us-west-1.amazonaws.com/biohub-8940d923/icons/colab.svg" href="https://colab.research.google.com/github/Biohub/esm/blob/main/cookbook/tutorials/esmc_finetune.ipynb">
    Fine-tune a classification or regression head for your dataset on top of ESMC using Parameter
    Efficient Fine-tuning (PEFT).
  </Card>
</Columns>

<h2 id="model-details">
  Model Details
</h2>

For additional information, see the Hugging Face link.

### Model Card

<Accordion title="Cite this Model">
  [Language Modeling Materializes a World Model of Protein Biology](https://www.biorxiv.org/content/10.64898/2026.06.03.729735)

  ```bibtex theme={null}
  @misc{candido2026language,
    title  = {Language Modeling Materializes a World Model of Protein Biology},
    author = {Candido, Salvatore and Hayes, Thomas and Derry, Alexander and Rao, Roshan
              and Lin, Zeming and Verkuil, Robert and Wu, Bryan and Lee, Jin Sub
              and Bruguera, Elise S. and Keval, Jehan A. and Kopylov, Mykhailo
              and Pak, John E. and Wu, Wesley and Thomas, Neil and Mataraso, Samson
              and Hsu, Alvin and Trotman-Grant, Ashton C. and Fatras, Kilian
              and dos Santos Costa, Allan and Badkundri, Rohil and Ak{\i}n, Halil
              and Oktay, Deniz and Deaton, Jonathan and Montabana, Elizabeth
              and Sitwala, Hrishita and Yu, Yue and Wiggert, Marius
              and Carlin, Dylan Alexander and Goering, Anthony W. and Blazejewski, Tomasz
              and Sandora, McCullen and Hla, Michael and Jia, Tina Z.
              and Kloker, Leon H. and Sofroniew, Nicholas J. and Uehara, Masatoshi
              and Pannu, Jassi and Bachas, Sharrol and Liu, Daniel S.
              and Sercu, Tom and Rives, Alexander},
    year   = {2026},
    url    = {https://www.biorxiv.org/content/10.64898/2026.06.03.729735},
    note   = {Preprint}
  }
  ```
</Accordion>

<Tabs>
  <Tab title="Overview">
    | | |
    | - | - |
    | **Version** | 2026-04 |
    | **Architecture** | Transformer |
    | **Supported Modalities** | Sequence |
    | **Training Data** | Up to 6 billion proteins |

    #### Intended Use

    ESMC is designed for protein science research including structure prediction, function
    annotation, protein design, and understanding evolutionary relationships between proteins. It
    can generate novel proteins given partial sequence, structure, or functional constraints.

    #### Limitations & Risks

    Outputs should be validated experimentally. The model may generate proteins that are not
    synthesizable or functional. Not intended for clinical or therapeutic applications without
    further validation.

    This model is released under the
    [MIT License](https://github.com/Biohub/esm/blob/main/LICENSE.md).
  </Tab>

  <Tab title="Details">
    | | |
    | - | - |
    | **Parameters** | Approximately 300 million, 600 million, and 6 billion parameters for ESMC 300M, ESMC 600M, and ESMC 6B respectively. |
    | **Transformer Layers** | 30, 36, and 80 layers for ESMC 300M, ESMC 600M, and ESMC 6B respectively. |
    | **Type of Model Architecture** | Transformer architecture where each model consists of a token embedding layer, a stack of transformer blocks, and a regression output head. It features Pre-LN, rotary embeddings, and SwiGLU activations. No biases are used in linear layers or layer norms. |
    | **Training FLOPs** | 1×10²², 1×10²², and 2×10²³ for ESMC 300M, ESMC 600M, and ESMC 6B respectively. |
    | **Training Datasets** | ESMC was trained on data pulled from UniRef, MGnify, and the Joint Genome Institute (JGI) Metagenome database. |
  </Tab>

  <Tab title="Variants">
    | Model | Model Size | Number of Layers | Max Context Length | Release Date |
    | - | - | - | - | - |
    | esmc-6b-2024-12 | 6B | 80 | 16,384 | 2024-12 |
    | esmc-600m-2024-12 | 600M | 36 | 16,384 | 2024-12 |
    | esmc-300m-2024-12 | 300M | 30 | 16,384 | 2024-12 |
  </Tab>

  <Tab title="Usage">
    | | |
    | - | - |
    | **Primary Use Case** | Protein representation learning and embeddings, generating protein embeddings that can be fine-tuned for various downstream prediction tasks including functional annotation, mutational effect analysis, and the design of novel proteins and peptides. Predicting the functional impact of mutations and amino acid substitutions on protein function. |
    | **Supported Input Modalities** | Protein sequences |
    | **Access** | Through [Biohub](#get-started) or [Hugging Face](https://huggingface.co/collections/biohub/esmc-model-family) |
    | **License** | This model is released under the [MIT License](https://github.com/Biohub/esm/blob/main/LICENSE.md). |
    | **Not Recommended For** | Clinical diagnosis or treatment recommendations. Computational metrics do not replace wet-lab validation. Treat model outputs as machine-generated hypotheses that require further experimental validation, not as established biological facts. |
  </Tab>
</Tabs>

<h2 id="model-data">
  ESM Atlas Data
</h2>

The ESM Atlas is a map of 6.8 billion proteins covering the full breadth of life's biodiversity and
more than one billion predicted structures. ESM Atlas organizes protein space by the shared
representation space learned by ESMC, with high-resolution structures predicted by ESMFold2. Users
can download the dataset used to generate the Atlas from AWS S3 at no cost. The table below lists
the available datasets with their approximate sizes and the AWS CLI command to download each one.

If you are unsure where to start, **SAE Clusters** is the most manageable entry point for exploring
cluster-level organization. Learn more information about SAEs
[here](https://huggingface.co/biohub/esmc-SAE-overview). Download **All Data** only if you need the
complete set of sequences, structures, and features across all 6.8 billion proteins.

| Dataset | Description | Package Size | Download command |
| - | - | - | - |
| Sequences | Protein sequences (6.8B proteins) | 2.2 TB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/sequences/ /mydrive` |
| Structures | Protein structures (1B proteins) | 68.9 TB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/folds/ /mydrive` |
| SAE features | Per protein and per-residue feature vectors (6.8B proteins) | 306 TB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/sae/data_shards/ /mydrive` |
| SAE Clusters | Cluster-level organization based on SAE features (7.5M clusters) | 26 GB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/clusters/indexes/secondary/cluster_members/ /mydrive` |
| HMM Results | Predicted pfam and taxonomy (6.8B proteins) | 653 MB | `aws s3 cp --no-sign-request s3://esm-protein-atlas/v1/clusters/data/representative_proteins.parquet /mydrive/` |
| Protein\_to\_accession | Mapping of protein IDs to accession numbers (6.8B proteins) | 162 GB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/shared_indexes/ /mydrive` |
| Normalization | SAE feature normalization | 192 KB | `aws s3 cp --no-sign-request s3://esm-protein-atlas/v1/normalization/max_idf_log10.pkl /mydrive/` |
| All Data | Complete set of sequences, structures, features, and clusters | 377 TB | `aws s3 sync --no-sign-request s3://esm-protein-atlas/v1/ /mydrive` |
