Open the Atlas
Atlas API Reference
Download the Data
What You Can Do
With the Atlas
- Search the Atlas using the agent for any protein or protein function
- View a protein’s predicted 3D structure and top activated biological features
- Find structurally and functionally similar proteins, even ones with no obvious sequence similarity to your query
- Explore clusters of related proteins and understand their biological context
- Download structures, sequences, and feature data for use in your own analyses
With the API
- Retrieve an Atlas protein: fetch the full stored record for a protein by its MD5 hash — metadata (source, accession), sequence, predicted structure (PDB with per-residue pLDDT and pTM), SAE features and per-residue activations, and a pointer to its cluster representative
- Resolve a UniProt accession: look up a UniProtKB accession or entry name (e.g.
P24941,P53_HUMAN) and get back its name, gene, organism, function, sequence, and Atlasprotein_hash - Predict SAE Features: compute a sequence’s SAE activations using ESMC and the trained SAE — returns protein-level activations, per-residue activations, and top-K enriched features with biological labels
- Predict 3D structure: predict an ESMFold2 structure for an amino acid sequence (<700aa)
- Search for similar proteins: find proteins in the Atlas similar to a query sequence using SAE feature embedding similarity
- Explore SAE features: retrieve descriptions of the 16,384-dimensional feature space that characterizes proteins in the Atlas
- Inspect a protein cluster: for any protein, retrieve information about the cluster it belongs to — member proteins, top Pfam domains, top SAE features, and taxonomy distribution
- Batch protein download: submit many proteins as an asynchronous job and download their sequences, SAE features, structures, and cluster membership from S3
Navigating the Atlas
When you first open the Atlas, you will see a map of the protein universe. You can click on a protein in the UMAP to view its details or you can ask the agent to help you search for a protein of interest. We recommend starting with a protein you are familiar with to familiarize yourself with the Atlas.UMAP Visualization
This is a way of taking an enormous, complex dataset and compressing it into two dimensions so you can see patterns visually. Each dot represents a cluster of proteins, positioned using a 2D projection of their biological feature profiles. Clusters that appear close together share similar biological signals. Think of it less like a precise coordinate system and more like a neighborhood map: proximity is meaningful, but exact distance isn’t. Two clusters sitting next to each other share biological signals. Two clusters on opposite ends of the map are functionally very different. Note that the UMAP is showing only cluster centers for clusters with at least 50 protein members. What the colors mean: Clusters are colored by the proportion of member proteins that match a known, annotated protein family. Brightly colored clusters on the yellow end of the spectrum tend to contain unknown proteins; while darker clusters contain more known protein families.Searching the Atlas
You can explore the Atlas by using the agent. Describe what you’re looking for (e.g., “mitochondrial membrane transporter involved in calcium signaling”) and the Atlas agent will interpret your query, search UniProt, and return candidate proteins for you to explore. When you search for a protein, you are searching against the clustered dataset (those with at least 50 members). This is designed to capture most large functional groups and families, but note there may be rare functions or edge cases that are not shown in the UI. Once a protein is retrieved, you’ll see:- Its predicted 3D structure, colored by per-residue confidence (pLDDT)
- Its top activated SAE features with interpretable labels (e.g., “Folate cofactor-binding pocket”)
- Its cluster membership and a link to explore similar proteins
- The option to view its neighborhood of similar SAE features
Understanding SAE Features
What Are SAE Features?
When ESMC processes a protein sequence, it encodes everything it learns about that protein into a dense numerical representation that represents structural, functional, and evolutionary information all at once. The problem is that this representation is difficult to interpret since each number reflects a mix of many biological concepts tangled together. Sparse Autoencoders (SAEs) are simple neural networks trained to untangle that signal. Think of it like separating the instruments in a piece of music. Instead of hearing everything blended together, you can isolate the violin, the cello, the piano. SAEs decompose ESMC’s representation into thousands of individual, interpretable features with each one corresponding to a specific biological concept. The Atlas uses an SAE with ~16,000 features. Each feature has been mapped to a biological concept by examining which proteins activate it and what those proteins have in common. SAE features capture a wide range of biological concepts, from broad patterns like aromatic residues to specific functional motifs like folate cofactor-binding pockets, and their interpretability varies accordingly. Examples include:- Folate cofactor-binding pocket
- Hydrophobic transmembrane helices
- Acidic juxtamembrane segments
- Extended polar low-complexity IDRs
What Does “Activation” Mean?
For any given protein, only a small number of features will activate (typically fewer than 1% of the full set). A feature activates when ESMC detects that biological signal in your protein. Higher activation scores indicate stronger, more confident signals. When you view a protein in the Atlas, the top activated features are shown ranked by activation score. You can click on any feature to see its description and view which residues on the structure are driving the activation. Features are ranked by a normalized activation score where each feature’s activation is scaled by its maximum observed activation across 208 million UniRef90 proteins, then weighted by how rarely and selectively the feature appears. Features that are both strongly activated and rare across the broader protein space are surfaced as the most informative features for interpreting a given protein.Are SAE Features the Same as Database Annotations?
No, and this distinction is important. Standard database annotations (like UniProt function entries or Pfam domain labels) reflect what scientists have experimentally observed and curated over decades. SAE features, by contrast, emerge from the model learning patterns across billions of sequences. They often align closely with known biology, but they can also detect functional signals that haven’t been formally annotated.What Do We Mean by “Similarity”?
The Difference Between Homology and Feature Similarity
Traditional protein similarity relies on sequence homology where two proteins are considered related if their amino acid sequences are similar enough. This works well for well-studied protein families, but it breaks down for proteins that have arrived at the same function through different evolutionary paths (convergent evolution), for the vast regions of protein space that have simply never been studied, and for remote homologs, sequences that share a common ancestor, but over time have diverged beyond the detection limit for sequence homology. The Atlas uses a different approach: SAE feature similarity. Two proteins are considered similar if they share a similar set of activated biological features according to the ESMC world model, regardless of whether their sequences or even their overall structures look alike.How Is Similarity Calculated?
Similarity is measured using cosine similarity. Two proteins score as similar if they activate the same biological features in similar proportions.What Is a Cluster?
The Atlas groups proteins into clusters based on SAE feature similarity. Each cluster is a group of proteins that share a highly similar set of activated biological features — meaning they likely share functional characteristics, even if they don’t share obvious sequence identity.How Are Clusters Calculated?
Clusters are built using a specialized linear-time hash-based algorithm inspired by Linclust, adapted to operate directly in SAE feature space rather than sequence space. The algorithm uses MinHash signatures and Locality-Sensitive Hashing to efficiently identify candidate protein pairs, which are then verified using exact Jaccard similarity. Proteins are grouped into clusters using a greedy algorithm where each cluster member is guaranteed to share at least 60% feature overlap (Jaccard similarity ≥ 0.6) with its cluster representative, meaning the number of SAE features active in both the member and representative is at least 60% of the total unique features active. Each cluster is automatically labeled by a language model, which reads the functional annotations of cluster members and generates a concise 2–5 word description of the shared theme (e.g., “Zinc finger proteins,” “Mitochondrial transporters”).What Can I Learn From a Cluster?
The Cluster Report for each cluster includes:- Number of proteins: how many proteins belong to this cluster globally
- Top PFAM domains: known domain annotations found in cluster members, giving a sense of what’s characterized
- Taxonomy distribution: which organisms and lineages are represented, useful for understanding evolutionary conservation
- Top SAE features: the biological signals most strongly shared across cluster members
What Does “Partially Characterized” Mean?
- Characterized: the protein matches a known Pfam domain with a defined function
- Partially characterized: the protein doesn’t match a characterized domain itself, but shares a cluster with proteins that do, suggesting it may share functional characteristics via convergent evolution
- Uncharacterized: the entire cluster contains only proteins with no known functional annotation
Interpreting Results
pLDDT (Per-Residue Confidence Score)
pLDDT stands for predicted Local Distance Difference Test. It’s a per-residue confidence score that tells you how confident ESMFold2 is in the predicted position of each amino acid in the 3D structure. Scores range from 0 to 100:
In the structure viewer, residues are colored by pLDDT. High-confidence regions are shown in blue;
low-confidence regions shade toward yellow and orange. Low pLDDT doesn’t always mean the prediction
is wrong and many disordered regions are functionally important and are simply not expected to fold
into a fixed structure.
pTM (Predicted TM-Score)
pTM is a single global confidence score for the entire predicted structure, ranging from 0 to 1. A pTM above ~0.5 is generally considered a confident prediction. It reflects the model’s estimate of how well the predicted structure would superimpose with the true structure if it were known.Mean pLDDT
The average pLDDT across all residues in the protein, giving a quick overall sense of structural confidence. Proteins with very low mean pLDDT are likely intrinsically disordered across most of their length.SAE Feature Activation Score
Activation scores reflect how strongly a given feature was detected in the protein. Higher scores mean the biological signal associated with that feature is more prominent. The specific numeric range isn’t directly comparable across different features, focus on which features are activated and their relative ranking for your protein rather than comparing raw scores across proteins.Atlas Data
The full Atlas data is openly available for download.- Explorable dataset (~1.1 billion proteins). All source databases were concatenated and deduplicated, then clustered at 70% sequence identity to reduce redundancy. The clustered dataset is used as input for ESMC and resulting embeddings are used to compute SAE features. ESMFold2 is used to predict 3D structures.
- Full dataset (~6.5 billion proteins). All source databases are concatenated and deduplicated without clustering. ESMC is used to compute SAE features across the full set.
Data Sources
What Data Is Available for Download?
The Atlas dataset is available for download from AWS S3 at no cost. There are several download options depending on what you need:For the S3 bucket paths and the
aws s3 commands to download each dataset, see
ESM Atlas Data on the ESMC model page.Guardrails
The Atlas has multiple guardrails that detect and restrict the use of queries related to controlled pathogens and toxins. If you query the Atlas agent using keywords (such as name, accession, function) or sequences corresponding to these, you will encounter our guardrails and the agent will refuse to continue.If this occurs, you should refresh the Atlas and begin a new conversation, otherwise the prior
refusal may impact how the agent answers future queries.
Frequently Asked Questions
My protein isn't in the Atlas — what happens when I search for it?
My protein isn't in the Atlas — what happens when I search for it?
If your sequence isn’t in the pre-computed dataset, the Atlas will fold it on the fly using
ESMFold2 and compute its SAE features in real time. This takes a bit longer than retrieving a
pre-computed result but works for any valid protein sequence.
Why does my protein have low feature activation scores across the board?
Why does my protein have low feature activation scores across the board?
Some proteins, particularly very short sequences, highly disordered proteins, or proteins from
poorly represented lineages, may show weaker feature signals. Low activation doesn’t necessarily
mean the protein is unimportant; it may simply reflect that the model has less evolutionary
context to work with.
How is this different from searching AlphaFold DB or UniProt?
How is this different from searching AlphaFold DB or UniProt?
AlphaFold DB provides predicted structures but organizes proteins by species and sequence identity,
not by learned functional signals. UniProt provides curated functional annotations, but coverage is
uneven with many proteins having no annotation at all. The Atlas provides a complementary,
model-derived view that surfaces functional relationships between proteins that may share no
sequence similarity or have no annotations, based purely on the biological signals ESMC detected in
their sequences.
Can I use the Atlas API without a login?
Can I use the Atlas API without a login?
Yes. The Atlas API is fully public and requires no account or authentication. See the
API Reference for endpoint details.
The similarity search returned proteins from a completely different organism than I expected, is that right?
The similarity search returned proteins from a completely different organism than I expected, is that right?
Yes, and it’s often the most interesting result. Feature similarity is not constrained by taxonomy.
A bacterial protein and a human protein can share strong feature similarity if they’ve evolved to
perform the same molecular function. This is especially useful for placing uncharacterized proteins
in a biological context when no close relatives exist in model organisms.