Skip to main content
The Biohub Platform is developed by Biohub to accelerate biological research by providing open access to frontier AI models, curated datasets, and interactive tools. The platform hosts the ESM family of protein language models (ESMC, ESMFold2, and ESM3), the ESM Atlas consisting of 1B+ predicted protein structures, and interactive applications for structure prediction and protein design.
No account is needed to browse models, datasets, documentation, tools, or the ESM Atlas. However, if you would like to bypass daily request limits or access the API, please create a free account.
Biohub.ai provides access to the ESM family of protein AI models and associated datasets.Models include:
  • ESMC (ESM Cambrian): A next-generation protein representation learning model available at 300M, 600M, and 6B parameter scales. ESMC is a drop-in replacement for ESM2 with major performance and efficiency gains.
  • ESMFold2: Structure prediction from a single DNA, RNA or protein sequence.
  • ESM3: A frontier multimodal generative model that reasons jointly across protein sequence, structure, and function. Available at 1.4B, 7B, and 98B parameter scales.
Data includes:
  • The ESM Atlas: A comprehensive protein database built from over 6 billion protein sequences extracted from metagenomic sources. The Atlas includes predicted structures for 1 billion proteins, sparse autoencoder (SAE) features for all 6 billion proteins, and cluster assignments with metadata for 10–100 million cluster representatives. The Atlas Explorer lets you browse and search cluster representatives interactively, and bulk downloads are available via S3.
Refer to the Model pages for full details.
There are several ways to get started:
  • Open-weight models: Download from Hugging Face to run locally.
  • GitHub: Access model code and tutorials.
  • API access: Use the Python SDK to run inference programmatically. Install the esm package via pip and authenticate with your API token.
  • Fold app: Interactive web application for structure prediction. Requires a free account.
  • Tutorials and quickstart guides: Available for each model (ESMC, ESMFold2, ESM3) to walk you through common workflows.
Visit the model specific pages (ESMC, ESMFold2, ESM3) to explore what is available and find the right starting point for your research.
ESM models can be accessed through the Biohub Platform API. In addition:
  • Many models—including ESMC, ESMFold2, and ESM3-open—are also available as open-source releases on GitHub and Hugging Face under the MIT License.
  • Each model card identifies the available access methods and applicable terms.
  • To obtain API access to variants of ESM3 apart from ESM3-open, fill out this form with your project details and desired model.
You can access datasets through AWS S3 via Model pages (ESMC, ESMFold2, ESM3) or the ESM Atlas. For some datasets, processed data or data processing scripts are also available for download, making it easier to train and validate your own models.
Please reach out to us via the “Contact Us” or “Slack” buttons on the lower left-hand sidebar.
Yes, we use cookies and similar technologies to help improve the website, personalize content, and provide a better experience while using the platform. View our cookie policy for more information.
We implement guardrails that detect and restrict the use of keywords and sequences corresponding to controlled pathogens and toxins on our openly accessible platform. We recognize there are many legitimate reasons to use AI models to understand and model these sequences. If you are a researcher who would like to request access to guardrails-free models and data for research purposes, please complete this short form.
Yes, we may make minor updates to the models behind the APIs. We will also release new models under new identifiers. At that point, older models may become unavailable. Currently, we do not make guarantees about how long specific models will be available. Over time, we may introduce availability guarantees per model. We will do our best to communicate clearly and well ahead of any model deprecations.
Use of biohub.ai is subject to our Terms of Use and Privacy Policy. By submitting inputs or otherwise using Biohub.ai, you agree to those terms.
Biohub.ai is the AI platform built and maintained by Biohub, a nonprofit research institute created by the Chan Zuckerberg Initiative (CZI). Biohub is building the first large-scale scientific initiative combining frontier AI with frontier biology to solve disease. In 2025, EvolutionaryScale joined Biohub to accelerate this mission, and the ESM family of protein language models can now be accessed through Biohub.ai.
EvolutionaryScale Forge (forge.evolutionaryscale.ai) have been unified into Biohub.ai. Models, datasets, and apps from both platforms are now available in one place. Your existing Forge account data and API tokens have been transferred over, so you can sign in and pick up where you left off. If you experience any issues accessing your account, please reach out via the “Contact Us” button on the lower left-hand sidebar.
The ESM Atlas is a database of 1 billion predicted protein structures generated from metagenomic sequences using ESMFold2. It covers a vast range of proteins discovered through environmental sequencing that have no experimentally determined structure. You can search, browse, and download structures from the Atlas directly on Biohub.ai. The Atlas is freely available for academic and commercial use.
The Biohub Platform offers interactive tools that let you use ESM models without writing code:
  • Fold: Predict the 3D structure of a protein from its amino acid sequence, DNA or RNA using ESMFold2.
  • ESM Atlas: Explore over 1 billion predicted protein structures from metagenomic sequences, with searchable cluster representatives, SAE features, and bulk downloads via S3.