> ## Documentation Index
> Fetch the complete documentation index at: https://docs.biohub.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# FAQ

> Answers to common questions about the Biohub Platform, including open-source models, datasets, and research resources.

<AccordionGroup>
  <Accordion title="What is Biohub.ai?">
    The Biohub Platform is developed by Biohub to accelerate biological research by providing open
    access to frontier AI models, curated datasets, and interactive tools. The platform hosts the ESM
    family of protein language models (ESMC, ESMFold2, and ESM3), the ESM Atlas consisting of 1B+
    predicted protein structures, and interactive applications for structure prediction and protein
    design.
  </Accordion>

  <Accordion title="How do I register for the platform?">
    No account is needed to browse models, datasets, documentation, tools, or the ESM Atlas. However,
    if you would like to bypass daily request limits or access the API, please
    [create a free account](https://biohub.ai/sign-up).
  </Accordion>

  <Accordion title="What models and data are available on Biohub.ai?">
    Biohub.ai provides access to the ESM family of protein AI models and associated datasets.

    Models include:

    * [ESMC](https://biohub.ai/models/esmc) (ESM Cambrian): A next-generation protein representation learning model available
      at 300M, 600M, and 6B parameter scales. ESMC is a drop-in replacement for ESM2 with major
      performance and efficiency gains.
    * [ESMFold2](https://biohub.ai/models/esmfold2): Structure prediction from a single DNA, RNA or protein sequence.
    * [ESM3](https://biohub.ai/models/esm3): A frontier multimodal generative model that reasons jointly across protein
      sequence, structure, and function. Available at 1.4B, 7B, and 98B parameter scales.

    Data includes:

    * The [ESM Atlas](https://biohub.ai/esm/protein/atlas): A comprehensive protein database built
      from over 6 billion protein sequences extracted from metagenomic sources. The Atlas includes
      predicted structures for 1 billion proteins, sparse autoencoder (SAE) features for all 6 billion
      proteins, and cluster assignments with metadata for 10–100 million cluster representatives. The
      Atlas Explorer lets you browse and search cluster representatives interactively, and bulk
      downloads are available via S3.

    Refer to the Model pages for full details.
  </Accordion>

  <Accordion title="How do I get started with the models?">
    There are several ways to get started:

    * Open-weight models: Download from [Hugging Face](https://huggingface.co/biohub) to run locally.
    * [GitHub](https://github.com/Biohub): Access model code and tutorials.
    * [API](/api/protein/models/logits) access: Use the Python SDK to run inference programmatically.
      Install the esm package via pip and authenticate with your API token.
    * [Fold](https://biohub.ai/tools/fold) app: Interactive web application for structure prediction.
      Requires a free account.
    * Tutorials and quickstart guides: Available for each model ([ESMC](/learn/tutorials/esmc),
      [ESMFold2](/learn/tutorials/esmfold2), [ESM3](/learn/tutorials/esm3)) to walk you through
      common workflows.

    Visit the model specific pages ([ESMC](https://biohub.ai/models/esmc), [ESMFold2](https://biohub.ai/models/esmfold2), [ESM3](https://biohub.ai/models/esm3)) to explore
    what is available and find the right starting point for your research.
  </Accordion>

  <Accordion title="How can I access current ESM models and their variants?">
    ESM models can be accessed through the Biohub Platform API. In addition:

    * Many models—including ESMC, ESMFold2, and ESM3-open—are also available as open-source releases
      on GitHub and Hugging Face under the MIT License.
    * Each model card identifies the available access methods and applicable terms.
    * To obtain API access to variants of ESM3 apart from ESM3-open, fill out
      [this form](https://info.biohub.org/biohub-additional-compute-credits) with your project details
      and desired model.
  </Accordion>

  <Accordion title="How can I access the datasets?">
    You can access datasets through AWS S3 via Model pages ([ESMC](https://biohub.ai/models/esmc), [ESMFold2](https://biohub.ai/models/esmfold2),
    [ESM3](https://biohub.ai/models/esm3)) or the [ESM Atlas](https://biohub.ai/esm/protein/atlas). For some datasets,
    processed data or data processing scripts are also available for download, making it easier to
    train and validate your own models.
  </Accordion>

  <Accordion title="How can I provide feedback or get help on the platform and its resources?">
    Please reach out to us via the "Contact Us" or "Slack" buttons on the lower left-hand sidebar.
  </Accordion>

  <Accordion title="Do you use Cookies on the platform?">
    Yes, we use cookies and similar technologies to help improve the website, personalize content, and
    provide a better experience while using the platform. View our cookie policy for more information.
  </Accordion>

  <Accordion title="Does Biohub.ai restrict certain keywords and sequences?">
    We implement guardrails that detect and restrict the use of keywords and sequences corresponding
    to controlled pathogens and toxins on our openly accessible platform. We recognize there are many
    legitimate reasons to use AI models to understand and model these sequences. If you are a
    researcher who would like to request access to guardrails-free models and data for research
    purposes, please complete this short
    [form](https://info.biohub.org/biohub-request-for-elevated-access).
  </Accordion>

  <Accordion title="Will models on Biohub.ai be updated or deprecated?">
    Yes, we may make minor updates to the models behind the APIs. We will also release new models
    under new identifiers. At that point, older models may become unavailable. Currently, we do not
    make guarantees about how long specific models will be available. Over time, we may introduce
    availability guarantees per model. We will do our best to communicate clearly and well ahead of
    any model deprecations.
  </Accordion>

  <Accordion title="What terms apply to my use of Biohub.ai?">
    Use of biohub.ai is subject to our Terms of Use and Privacy Policy. By submitting inputs or
    otherwise using Biohub.ai, you agree to those terms.
  </Accordion>

  <Accordion title="What is the relationship between Biohub, Chan Zuckerberg Initiative, Biohub.ai and EvolutionaryScale?">
    Biohub.ai is the AI platform built and maintained by Biohub, a nonprofit research institute
    created by the Chan Zuckerberg Initiative (CZI). Biohub is building the first large-scale
    scientific initiative combining frontier AI with frontier biology to solve disease. In 2025,
    EvolutionaryScale joined Biohub to accelerate this mission, and the ESM family of protein language
    models can now be accessed through Biohub.ai.
  </Accordion>

  <Accordion title="What happened to Forge?">
    EvolutionaryScale Forge (forge.evolutionaryscale.ai) have been unified into Biohub.ai. Models,
    datasets, and apps from both platforms are now available in one place. Your existing Forge account
    data and API tokens have been transferred over, so you can sign in and pick up where you left off.
    If you experience any issues accessing your account, please reach out via the "Contact Us" button
    on the lower left-hand sidebar.
  </Accordion>

  <Accordion title="What is the ESM Atlas?">
    The ESM Atlas is a database of 1 billion predicted protein structures generated from metagenomic
    sequences using ESMFold2. It covers a vast range of proteins discovered through environmental
    sequencing that have no experimentally determined structure. You can search, browse, and download
    structures from the [Atlas](https://biohub.ai/esm/protein/atlas) directly on Biohub.ai.
    **The Atlas is freely available for academic and commercial use.**
  </Accordion>

  <Accordion title="What tools are available on the Biohub Platform?">
    The Biohub Platform offers interactive tools that let you use ESM models without writing code:

    * Fold: Predict the 3D structure of a protein from its amino acid sequence, DNA or RNA using
      ESMFold2.
    * ESM Atlas: Explore over 1 billion predicted protein structures from metagenomic sequences, with
      searchable cluster representatives, SAE features, and bulk downloads via S3.
  </Accordion>
</AccordionGroup>
