Skip to main content
The ESM Atlas is a free, open-access resource for exploring the global protein universe. It provides predicted 3D structures from ESMFold2 and interpretable feature annotations, such as function, from the ESMC language model over a billion proteins, spanning organisms from bacteria and archaea to fungi, viruses, and animals, many of which have never been characterized. The Atlas is built on top of ESMC, a protein language model trained on billions of protein sequences from across the tree of life. Rather than organizing proteins by sequence similarity or curated database annotations, the Atlas organizes them by learned biological signals through patterns the model detected directly from protein sequence data.

Open the Atlas

Atlas API Reference

Download the Data

What You Can Do

With the Atlas

  • Search the Atlas using the agent for any protein or protein function
  • View a protein’s predicted 3D structure and top activated biological features
  • Find structurally and functionally similar proteins, even ones with no obvious sequence similarity to your query
  • Explore clusters of related proteins and understand their biological context
  • Download structures, sequences, and feature data for use in your own analyses

With the API

  • Retrieve an Atlas protein: fetch the full stored record for a protein by its MD5 hash — metadata (source, accession), sequence, predicted structure (PDB with per-residue pLDDT and pTM), SAE features and per-residue activations, and a pointer to its cluster representative
  • Resolve a UniProt accession: look up a UniProtKB accession or entry name (e.g. P24941, P53_HUMAN) and get back its name, gene, organism, function, sequence, and Atlas protein_hash
  • Predict SAE Features: compute a sequence’s SAE activations using ESMC and the trained SAE — returns protein-level activations, per-residue activations, and top-K enriched features with biological labels
  • Predict 3D structure: predict an ESMFold2 structure for an amino acid sequence (<700aa)
  • Search for similar proteins: find proteins in the Atlas similar to a query sequence using SAE feature embedding similarity
  • Explore SAE features: retrieve descriptions of the 16,384-dimensional feature space that characterizes proteins in the Atlas
  • Inspect a protein cluster: for any protein, retrieve information about the cluster it belongs to — member proteins, top Pfam domains, top SAE features, and taxonomy distribution
  • Batch protein download: submit many proteins as an asynchronous job and download their sequences, SAE features, structures, and cluster membership from S3
See the API Reference for full endpoint details and request/response schemas. When you first open the Atlas, you will see a map of the protein universe. You can click on a protein in the UMAP to view its details or you can ask the agent to help you search for a protein of interest. We recommend starting with a protein you are familiar with to familiarize yourself with the Atlas.

UMAP Visualization

This is a way of taking an enormous, complex dataset and compressing it into two dimensions so you can see patterns visually. Each dot represents a cluster of proteins, positioned using a 2D projection of their biological feature profiles. Clusters that appear close together share similar biological signals. Think of it less like a precise coordinate system and more like a neighborhood map: proximity is meaningful, but exact distance isn’t. Two clusters sitting next to each other share biological signals. Two clusters on opposite ends of the map are functionally very different. Note that the UMAP is showing only cluster centers for clusters with at least 50 protein members. What the colors mean: Clusters are colored by the proportion of member proteins that match a known, annotated protein family. Brightly colored clusters on the yellow end of the spectrum tend to contain unknown proteins; while darker clusters contain more known protein families.

Searching the Atlas

You can explore the Atlas by using the agent. Describe what you’re looking for (e.g., “mitochondrial membrane transporter involved in calcium signaling”) and the Atlas agent will interpret your query, search UniProt, and return candidate proteins for you to explore. When you search for a protein, you are searching against the clustered dataset (those with at least 50 members). This is designed to capture most large functional groups and families, but note there may be rare functions or edge cases that are not shown in the UI. Once a protein is retrieved, you’ll see:
  • Its predicted 3D structure, colored by per-residue confidence (pLDDT)
  • Its top activated SAE features with interpretable labels (e.g., “Folate cofactor-binding pocket”)
  • Its cluster membership and a link to explore similar proteins
  • The option to view its neighborhood of similar SAE features

Understanding SAE Features

What Are SAE Features?

When ESMC processes a protein sequence, it encodes everything it learns about that protein into a dense numerical representation that represents structural, functional, and evolutionary information all at once. The problem is that this representation is difficult to interpret since each number reflects a mix of many biological concepts tangled together. Sparse Autoencoders (SAEs) are simple neural networks trained to untangle that signal. Think of it like separating the instruments in a piece of music. Instead of hearing everything blended together, you can isolate the violin, the cello, the piano. SAEs decompose ESMC’s representation into thousands of individual, interpretable features with each one corresponding to a specific biological concept. The Atlas uses an SAE with ~16,000 features. Each feature has been mapped to a biological concept by examining which proteins activate it and what those proteins have in common. SAE features capture a wide range of biological concepts, from broad patterns like aromatic residues to specific functional motifs like folate cofactor-binding pockets, and their interpretability varies accordingly. Examples include:
  • Folate cofactor-binding pocket
  • Hydrophobic transmembrane helices
  • Acidic juxtamembrane segments
  • Extended polar low-complexity IDRs

What Does “Activation” Mean?

For any given protein, only a small number of features will activate (typically fewer than 1% of the full set). A feature activates when ESMC detects that biological signal in your protein. Higher activation scores indicate stronger, more confident signals. When you view a protein in the Atlas, the top activated features are shown ranked by activation score. You can click on any feature to see its description and view which residues on the structure are driving the activation. Features are ranked by a normalized activation score where each feature’s activation is scaled by its maximum observed activation across 208 million UniRef90 proteins, then weighted by how rarely and selectively the feature appears. Features that are both strongly activated and rare across the broader protein space are surfaced as the most informative features for interpreting a given protein.

Are SAE Features the Same as Database Annotations?

No, and this distinction is important. Standard database annotations (like UniProt function entries or Pfam domain labels) reflect what scientists have experimentally observed and curated over decades. SAE features, by contrast, emerge from the model learning patterns across billions of sequences. They often align closely with known biology, but they can also detect functional signals that haven’t been formally annotated.

What Do We Mean by “Similarity”?

The Difference Between Homology and Feature Similarity

Traditional protein similarity relies on sequence homology where two proteins are considered related if their amino acid sequences are similar enough. This works well for well-studied protein families, but it breaks down for proteins that have arrived at the same function through different evolutionary paths (convergent evolution), for the vast regions of protein space that have simply never been studied, and for remote homologs, sequences that share a common ancestor, but over time have diverged beyond the detection limit for sequence homology. The Atlas uses a different approach: SAE feature similarity. Two proteins are considered similar if they share a similar set of activated biological features according to the ESMC world model, regardless of whether their sequences or even their overall structures look alike.

How Is Similarity Calculated?

Similarity is measured using cosine similarity. Two proteins score as similar if they activate the same biological features in similar proportions.

What Is a Cluster?

The Atlas groups proteins into clusters based on SAE feature similarity. Each cluster is a group of proteins that share a highly similar set of activated biological features — meaning they likely share functional characteristics, even if they don’t share obvious sequence identity.

How Are Clusters Calculated?

Clusters are built using a specialized linear-time hash-based algorithm inspired by Linclust, adapted to operate directly in SAE feature space rather than sequence space. The algorithm uses MinHash signatures and Locality-Sensitive Hashing to efficiently identify candidate protein pairs, which are then verified using exact Jaccard similarity. Proteins are grouped into clusters using a greedy algorithm where each cluster member is guaranteed to share at least 60% feature overlap (Jaccard similarity ≥ 0.6) with its cluster representative, meaning the number of SAE features active in both the member and representative is at least 60% of the total unique features active. Each cluster is automatically labeled by a language model, which reads the functional annotations of cluster members and generates a concise 2–5 word description of the shared theme (e.g., “Zinc finger proteins,” “Mitochondrial transporters”).

What Can I Learn From a Cluster?

The Cluster Report for each cluster includes:
  • Number of proteins: how many proteins belong to this cluster globally
  • Top PFAM domains: known domain annotations found in cluster members, giving a sense of what’s characterized
  • Taxonomy distribution: which organisms and lineages are represented, useful for understanding evolutionary conservation
  • Top SAE features: the biological signals most strongly shared across cluster members

What Does “Partially Characterized” Mean?

  • Characterized: the protein matches a known Pfam domain with a defined function
  • Partially characterized: the protein doesn’t match a characterized domain itself, but shares a cluster with proteins that do, suggesting it may share functional characteristics via convergent evolution
  • Uncharacterized: the entire cluster contains only proteins with no known functional annotation

Interpreting Results

pLDDT (Per-Residue Confidence Score)

pLDDT stands for predicted Local Distance Difference Test. It’s a per-residue confidence score that tells you how confident ESMFold2 is in the predicted position of each amino acid in the 3D structure. Scores range from 0 to 100: In the structure viewer, residues are colored by pLDDT. High-confidence regions are shown in blue; low-confidence regions shade toward yellow and orange. Low pLDDT doesn’t always mean the prediction is wrong and many disordered regions are functionally important and are simply not expected to fold into a fixed structure.

pTM (Predicted TM-Score)

pTM is a single global confidence score for the entire predicted structure, ranging from 0 to 1. A pTM above ~0.5 is generally considered a confident prediction. It reflects the model’s estimate of how well the predicted structure would superimpose with the true structure if it were known.

Mean pLDDT

The average pLDDT across all residues in the protein, giving a quick overall sense of structural confidence. Proteins with very low mean pLDDT are likely intrinsically disordered across most of their length.

SAE Feature Activation Score

Activation scores reflect how strongly a given feature was detected in the protein. Higher scores mean the biological signal associated with that feature is more prominent. The specific numeric range isn’t directly comparable across different features, focus on which features are activated and their relative ranking for your protein rather than comparing raw scores across proteins.

Atlas Data

The full Atlas data is openly available for download.
  • Explorable dataset (~1.1 billion proteins). All source databases were concatenated and deduplicated, then clustered at 70% sequence identity to reduce redundancy. The clustered dataset is used as input for ESMC and resulting embeddings are used to compute SAE features. ESMFold2 is used to predict 3D structures.
  • Full dataset (~6.5 billion proteins). All source databases are concatenated and deduplicated without clustering. ESMC is used to compute SAE features across the full set.

Data Sources

What Data Is Available for Download?

The Atlas dataset is available for download from AWS S3 at no cost. There are several download options depending on what you need:
For the S3 bucket paths and the aws s3 commands to download each dataset, see ESM Atlas Data on the ESMC model page.

Guardrails

The Atlas has multiple guardrails that detect and restrict the use of queries related to controlled pathogens and toxins. If you query the Atlas agent using keywords (such as name, accession, function) or sequences corresponding to these, you will encounter our guardrails and the agent will refuse to continue.
If this occurs, you should refresh the Atlas and begin a new conversation, otherwise the prior refusal may impact how the agent answers future queries.
We recognize that there are many legitimate reasons to use AI models to understand and model these sequences and proteins. If you are a researcher whose work is impacted by our guardrails, you can request elevated access to our platform here. Elevated access has no additional costs. If you have feedback on how our guardrails function, please share it with us by filling out our Feedback Form.

Frequently Asked Questions

If your sequence isn’t in the pre-computed dataset, the Atlas will fold it on the fly using ESMFold2 and compute its SAE features in real time. This takes a bit longer than retrieving a pre-computed result but works for any valid protein sequence.
Some proteins, particularly very short sequences, highly disordered proteins, or proteins from poorly represented lineages, may show weaker feature signals. Low activation doesn’t necessarily mean the protein is unimportant; it may simply reflect that the model has less evolutionary context to work with.
AlphaFold DB provides predicted structures but organizes proteins by species and sequence identity, not by learned functional signals. UniProt provides curated functional annotations, but coverage is uneven with many proteins having no annotation at all. The Atlas provides a complementary, model-derived view that surfaces functional relationships between proteins that may share no sequence similarity or have no annotations, based purely on the biological signals ESMC detected in their sequences.
Yes. The Atlas API is fully public and requires no account or authentication. See the API Reference for endpoint details.
Yes, and it’s often the most interesting result. Feature similarity is not constrained by taxonomy. A bacterial protein and a human protein can share strong feature similarity if they’ve evolved to perform the same molecular function. This is especially useful for placing uncharacterized proteins in a biological context when no close relatives exist in model organisms.

Next Steps

Learn how to analyze your own data using ESMC and ESMFold2. A great place to start is our tutorials.

Tutorials

Getting Started

Atlas API guides

Use the ESM Atlas from an AI assistant