ui_show_protein_structure, is covered in Structure viewer.
Sequence inputs
Tools that take asequence expect a raw amino-acid sequence in one-letter codes.
- Use the 20 standard residues or
X,B,U,Z, orO. - Letters are case-insensitive, and spaces before or after the sequence are removed.
- Don’t include a FASTA header, spaces, or line breaks inside the sequence.
- The ESM Atlas matches the exact sequence, so a variant or trimmed sequence counts as a different protein.
invalid_input.
The server also accepts : and | and passes them to the ESM Atlas unchanged.
Search UniProt
esm_atlas_search_uniprot searches UniProt and returns protein metadata with sequences you can pass to the other tools.
Inputs
A query shaped like a UniProt accession, such as
P42212, is looked up directly.
Anything else runs as a UniProt search, so fielded syntax works, for example gene:CDK2 AND organism_id:9606 AND reviewed:true.
Returns
Resolve an accession
esm_atlas_lookup_accession fetches the sequence for an accession from UniParc, MGnify, or IMG.
Inputs
The tool recognizes these formats:
For example,
UPI0000002FB4 is the UniParc entry for GFP.
A UniProt accession such as P42212 returns invalid_input.
Use esm_atlas_search_uniprot for those.
Returns
An accession the database doesn’t have, or one without a usable amino-acid sequence, returns
not_found.
A sequence longer than 4,000 residues returns invalid_input.
Get protein details
esm_atlas_get_protein_details returns a protein’s ESM Atlas record and its ten strongest SAE features.
Inputs
Compute on miss
If the ESM Atlas has the sequence, the tool returns the stored record. If it doesn’t, the ESM Atlas computes SAE features for sequences up to 2,048 residues and setsfeatures_computed_on_miss to true.
A longer sequence that misses returns sequence_too_long_for_features.
This tool never predicts a structure.
To view one, call ui_show_protein_structure with the same sequence.
Returns
Positions in
residue_regions start at 0, while the structure viewer numbers residues from 1, so add 1 to match them.
Find similar protein clusters
esm_atlas_search_similar_protein_clusters compares a sequence’s SAE feature vector with ESM Atlas cluster representatives and returns the closest ones.
Inputs
Set
uncharacterized_only to true to return only clusters with no characterized Pfam annotations.
Behavior
- You can get fewer than
top_khits, because the ESM Atlas leaves out matches with a similarity score below 0.5. - The ESM Atlas computes the query’s features if it doesn’t have them.
- As a side effect, the ESM Atlas can predict structures for up to six hits that don’t have one yet. The result contains no structures.
- Every hit includes its sequence, so the assistant can pass it to the other tools.
Returns
The REST equivalent is similarity search.
Get cluster details
esm_atlas_get_cluster_info summarizes the cluster that a representative sequence belongs to.
Inputs
The sequence must be a cluster representative, such as a hit from
esm_atlas_search_similar_protein_clusters.
Other sequences, including many well-known UniProt proteins, return not_found.
This tool reads stored data only and never computes.
Returns
The REST equivalent is get cluster.
Get SAE feature details
esm_atlas_get_sae_feature_detail explains one of the 16,384 SAE features learned from ESMC.
Inputs
Returns
The REST equivalent is get SAE feature.
Errors
A failed call returns a JSON object with a stablecode and a message, for example:
Troubleshooting explains what to do about each code.