Skip to main content
ESMC is the latest in the ESM family of protein language models, establishing a new frontier in representation learning for protein biology. Trained on billions of evolutionary sequences, it learns representations that reflect a mechanistic reduction of protein structure and function.

Hugging Face

GitHub

Paper

API Reference

Get Started

Get started using ESMC with our Quickstart and Tutorials.

Quickstart Guide

1

Install the esm Python package

2

Create an API key

Generate an API key from your Biohub account. This API key manages your access to credits and tokens, and the term API key/token is often used interchangeably within documentation.
3

Connect to the Biohub Platform API

Call the ESM client with the selected model of choice and replace <your API token> with your token name.
4

Run your inference

Now you are ready to use your model. For examples of specific use cases or to append a SAE onto your ESMC model, check out our Tutorials.

Model Tutorials

Embedding sequences with ESMC

Embed protein sequences and explore how different transformer layers encode structural and functional information.

Zero-shot entropy and mutation analysis

Compute per-position entropy and log-likelihood ratios to identify constrained vs. mutation-tolerant sites.

Layer sweep for enzyme function classification

Learn how to sweep all layers to find which one is best using enzyme classification as a task.

Understanding proteins with SAE features

Extract and visualize sparse autoencoder features, rank by peak activation, and map activations onto 3D structures.

Fine-tuning ESMC

Fine-tune a classification or regression head for your dataset on top of ESMC using Parameter Efficient Fine-tuning (PEFT).

Model Details

For additional information, see the Hugging Face link.

Model Card

Language Modeling Materializes a World Model of Protein Biology

Intended Use

ESMC is designed for protein science research including structure prediction, function annotation, protein design, and understanding evolutionary relationships between proteins. It can generate novel proteins given partial sequence, structure, or functional constraints.

Limitations & Risks

Outputs should be validated experimentally. The model may generate proteins that are not synthesizable or functional. Not intended for clinical or therapeutic applications without further validation.This model is released under the MIT License.

ESM Atlas Data

The ESM Atlas is a map of 6.8 billion proteins covering the full breadth of life’s biodiversity and more than one billion predicted structures. ESM Atlas organizes protein space by the shared representation space learned by ESMC, with high-resolution structures predicted by ESMFold2. Users can download the dataset used to generate the Atlas from AWS S3 at no cost. The table below lists the available datasets with their approximate sizes and the AWS CLI command to download each one. If you are unsure where to start, SAE Clusters is the most manageable entry point for exploring cluster-level organization. Learn more information about SAEs here. Download All Data only if you need the complete set of sequences, structures, and features across all 6.8 billion proteins.