Skip to main content
Biohub MCP has six ESM Atlas tools. Each returns structured JSON that your assistant can pass straight to the next tool, and each is marked read-only. The seventh tool, ui_show_protein_structure, is covered in Structure viewer.

Sequence inputs

Tools that take a sequence expect a raw amino-acid sequence in one-letter codes.
  • Use the 20 standard residues or X, B, U, Z, or O.
  • Letters are case-insensitive, and spaces before or after the sequence are removed.
  • Don’t include a FASTA header, spaces, or line breaks inside the sequence.
  • The ESM Atlas matches the exact sequence, so a variant or trimmed sequence counts as a different protein.
Any other character returns invalid_input. The server also accepts : and | and passes them to the ESM Atlas unchanged.

Search UniProt

esm_atlas_search_uniprot searches UniProt and returns protein metadata with sequences you can pass to the other tools.

Inputs

A query shaped like a UniProt accession, such as P42212, is looked up directly. Anything else runs as a UniProt search, so fielded syntax works, for example gene:CDK2 AND organism_id:9606 AND reviewed:true.

Returns

Resolve an accession

esm_atlas_lookup_accession fetches the sequence for an accession from UniParc, MGnify, or IMG.

Inputs

The tool recognizes these formats: For example, UPI0000002FB4 is the UniParc entry for GFP. A UniProt accession such as P42212 returns invalid_input. Use esm_atlas_search_uniprot for those.

Returns

An accession the database doesn’t have, or one without a usable amino-acid sequence, returns not_found. A sequence longer than 4,000 residues returns invalid_input.

Get protein details

esm_atlas_get_protein_details returns a protein’s ESM Atlas record and its ten strongest SAE features.

Inputs

Compute on miss

If the ESM Atlas has the sequence, the tool returns the stored record. If it doesn’t, the ESM Atlas computes SAE features for sequences up to 2,048 residues and sets features_computed_on_miss to true. A longer sequence that misses returns sequence_too_long_for_features. This tool never predicts a structure. To view one, call ui_show_protein_structure with the same sequence.

Returns

Positions in residue_regions start at 0, while the structure viewer numbers residues from 1, so add 1 to match them.

Find similar protein clusters

esm_atlas_search_similar_protein_clusters compares a sequence’s SAE feature vector with ESM Atlas cluster representatives and returns the closest ones.

Inputs

Set uncharacterized_only to true to return only clusters with no characterized Pfam annotations.

Behavior

  • You can get fewer than top_k hits, because the ESM Atlas leaves out matches with a similarity score below 0.5.
  • The ESM Atlas computes the query’s features if it doesn’t have them.
  • As a side effect, the ESM Atlas can predict structures for up to six hits that don’t have one yet. The result contains no structures.
  • Every hit includes its sequence, so the assistant can pass it to the other tools.

Returns

The REST equivalent is similarity search.

Get cluster details

esm_atlas_get_cluster_info summarizes the cluster that a representative sequence belongs to.

Inputs

The sequence must be a cluster representative, such as a hit from esm_atlas_search_similar_protein_clusters. Other sequences, including many well-known UniProt proteins, return not_found. This tool reads stored data only and never computes.

Returns

The REST equivalent is get cluster.

Get SAE feature details

esm_atlas_get_sae_feature_detail explains one of the 16,384 SAE features learned from ESMC.

Inputs

Returns

The REST equivalent is get SAE feature.

Errors

A failed call returns a JSON object with a stable code and a message, for example:
Troubleshooting explains what to do about each code.