Skip to main content
ESMFold2 is the successor to ESMFold that sets a new state of the art for single-sequence structure prediction and enables the generation of new functional proteins through searching the ESMC model’s latent space. The model predicts high-resolution, all-atom 3D structures of biomolecular complexes directly from sequence, with optional multiple sequence alignment (MSA) input for enhanced accuracy on challenging targets.

Hugging Face

GitHub

Paper

API Reference

Get Started

Get started with ESMFold2 with our Quickstart and Tutorials.

Quickstart Guide

1

Install the esm Python package

2

Create an API key

Generate an API key from your Biohub account. This API key manages your access to credits and tokens, and the term API key/token is often used interchangeably within documentation.
3

Connect to the Biohub Platform API

Call the inference client with the selected model of choice and replace <your API token> with your token name.
4

Run your inference

Now you are ready to use your model. For examples of specific use cases, check out our Tutorials.

Model Tutorials

Folding with ESMFold2

Fold proteins in combination with DNA, RNA, and small-molecule ligands.

Binder design

Design antibodies and minibinders with high hit rates. Implements the protocol featured in our paper, which produced binders exhibiting nanomolar affinity, target specificity, and functional activity in laboratory assays.

Model Details

For additional information, see the Hugging Face link.

Model Card

Language Modeling Materializes a World Model of Protein Biology

Intended Use

ESMFold2 predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input. The outputs include comprehensive structural information including all-atom coordinates (backbone and side chains), confidence metrics, and optional distogram predictions for detailed analysis of predicted structures.

Limitations & Risks

The model predicts single static conformations and is not designed for modeling protein dynamics, conformational flexibility, or multiple conformations of the same protein. Outputs should be validated experimentally. Not intended for clinical or therapeutic applications without further validation.This model is released under the MIT License.

ESM Atlas Data

The ESM Atlas is a map of 6.8 billion proteins covering the full breadth of life’s biodiversity and more than one billion predicted structures. ESM Atlas organizes protein space by the shared representation space learned by ESMC, with high-resolution structures predicted by ESMFold2. Users can download the dataset used to generate the Atlas from AWS S3 at no cost. The table below lists the available datasets with their approximate sizes and the AWS CLI command to download each one. If you are unsure where to start, SAE Clusters is the most manageable entry point for exploring cluster-level organization. Learn more information about SAEs here. Download All Data only if you need the complete set of sequences, structures, and features across all 6.8 billion proteins.