> ## Documentation Index
> Fetch the complete documentation index at: https://docs.biohub.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Biohub MCP ESM Atlas tools reference

> Reference for the six Biohub MCP ESM Atlas tools: UniProt search, accession lookup, protein details, similarity search, cluster info, and SAE features.

Biohub MCP has six ESM Atlas tools.
Each returns structured JSON that your assistant can pass straight to the next tool, and each is marked read-only.
The seventh tool, `ui_show_protein_structure`, is covered in [Structure viewer](/biohub-mcp/structure-viewer).

| Tool | Use it to |
| - | - |
| [`esm_atlas_search_uniprot`](#search-uniprot) | Find UniProt proteins and their sequences. |
| [`esm_atlas_lookup_accession`](#resolve-an-accession) | Get the sequence for a UniParc, MGnify, or IMG accession. |
| [`esm_atlas_get_protein_details`](#get-protein-details) | See a protein's ESM Atlas record and strongest SAE features. |
| [`esm_atlas_search_similar_protein_clusters`](#find-similar-protein-clusters) | Find cluster representatives with similar SAE features. |
| [`esm_atlas_get_cluster_info`](#get-cluster-details) | Summarize a cluster representative's cluster. |
| [`esm_atlas_get_sae_feature_detail`](#get-sae-feature-details) | Explain one SAE feature. |

## Sequence inputs

Tools that take a `sequence` expect a raw amino-acid sequence in one-letter codes.

* Use the 20 standard residues or `X`, `B`, `U`, `Z`, or `O`.
* Letters are case-insensitive, and spaces before or after the sequence are removed.
* Don't include a FASTA header, spaces, or line breaks inside the sequence.
* The ESM Atlas matches the exact sequence, so a variant or trimmed sequence counts as a different protein.

Any other character returns `invalid_input`.
The server also accepts `:` and `|` and passes them to the ESM Atlas unchanged.

## Search UniProt

`esm_atlas_search_uniprot` searches UniProt and returns protein metadata with sequences you can pass to the other tools.

### Inputs

| Name | Type | Required | Limits | Default |
| - | - | - | - | - |
| `query` | string | Yes | 1 to 500 characters | |
| `size` | integer | No | 1 to 6 | 6 |

A query shaped like a UniProt accession, such as `P42212`, is looked up directly.
Anything else runs as a UniProt search, so fielded syntax works, for example `gene:CDK2 AND organism_id:9606 AND reviewed:true`.

### Returns

| Field | Description |
| - | - |
| `query_label` | A short label for the query. |
| `proteins[]` | Matching records, at most `size`. |
| `proteins[].accession` | UniProt accession. |
| `proteins[].protein_name`, `gene`, `organism` | Recommended name, first gene name, and scientific name of the organism. |
| `proteins[].function` | UniProt function text, which can be empty. |
| `proteins[].sequence`, `sequence_length` | Amino-acid sequence and length. The sequence is `null` when it is longer than 4,000 residues. |
| `num_results` | Number of records returned. |
| `total_results` | Total matches UniProt reported. |

## Resolve an accession

`esm_atlas_lookup_accession` fetches the sequence for an accession from UniParc, MGnify, or IMG.

### Inputs

| Name | Type | Required | Limits | Default |
| - | - | - | - | - |
| `accession` | string | Yes | 1 to 128 characters | |

The tool recognizes these formats:

| Database | Format |
| - | - |
| UniParc | `UPI` followed by 10 hexadecimal characters |
| MGnify | `MGYP` followed by 12 digits |
| IMG | A 9 or 10 digit gene ID, or a `Ga0` or `JGI` metagenome gene ID |

For example, `UPI0000002FB4` is the UniParc entry for GFP.

A UniProt accession such as `P42212` returns `invalid_input`.
Use [`esm_atlas_search_uniprot`](#search-uniprot) for those.

### Returns

| Field | Description |
| - | - |
| `accession` | The accession you sent. |
| `database` | `uniparc`, `mgnify`, or `img`. |
| `sequence`, `sequence_length` | Amino-acid sequence and length. |

An accession the database doesn't have, or one without a usable amino-acid sequence, returns `not_found`.
A sequence longer than 4,000 residues returns `invalid_input`.

## Get protein details

`esm_atlas_get_protein_details` returns a protein's ESM Atlas record and its ten strongest SAE features.

### Inputs

| Name | Type | Required | Limits | Default |
| - | - | - | - | - |
| `sequence` | string | Yes | 1 to 4,000 residues | |

### Compute on miss

If the ESM Atlas has the sequence, the tool returns the stored record.
If it doesn't, the ESM Atlas computes SAE features for sequences up to 2,048 residues and sets `features_computed_on_miss` to `true`.
A longer sequence that misses returns `sequence_too_long_for_features`.

This tool never predicts a structure.
To view one, call [`ui_show_protein_structure`](/biohub-mcp/structure-viewer) with the same sequence.

### Returns

| Field | Description |
| - | - |
| `sequence`, `sequence_length` | The normalized sequence and its length. |
| `accession`, `header`, `source` | ESM Atlas identity, when the sequence is in the ESM Atlas. `source` names the collection, such as `uniparc`. |
| `sae_features[]` | Up to ten features, strongest first. |
| `sae_features[].feature_index` | Feature number, 0 to 16,383. Pass it to [`esm_atlas_get_sae_feature_detail`](#get-sae-feature-details). |
| `sae_features[].label`, `description` | What the feature is thought to represent. |
| `sae_features[].value` | Activation strength. |
| `sae_features[].label_reliability` | `high` when `value` is at least 1.0, `moderate` from 0.5, and `low` below 0.5. |
| `sae_features[].residue_regions[]` | Stretches where the feature is active: `start`, `end`, `peak_residue`, and `mean_activation`. |
| `ptm`, `mean_plddt` | Confidence of the stored structure, from 0 to 1, when there is one. |
| `structure_available` | Whether the ESM Atlas has a stored structure. |
| `structure_can_be_computed` | `true` when the sequence is 700 residues or fewer, so the viewer can predict a structure on a miss. |
| `features_computed_on_miss` | Whether the features were computed for this request. |
| `structure_source`, `structure_model` | Where the stored structure came from, when available. |
| `structure_size_bytes`, `residue_confidence_count` | Size of the stored PDB file and number of per-residue confidence values. |
| `folded_on_miss` | Always `false` for this tool. |

Positions in `residue_regions` start at 0, while the structure viewer numbers residues from 1, so add 1 to match them.

## Find similar protein clusters

`esm_atlas_search_similar_protein_clusters` compares a sequence's SAE feature vector with ESM Atlas cluster representatives and returns the closest ones.

### Inputs

| Name | Type | Required | Limits | Default |
| - | - | - | - | - |
| `sequence` | string | Yes | 1 to 2,048 residues | |
| `top_k` | integer | No | 1 to 12 | 10 |
| `uncharacterized_only` | boolean | No | | `false` |

Set `uncharacterized_only` to `true` to return only clusters with no characterized Pfam annotations.

### Behavior

* You can get fewer than `top_k` hits, because the ESM Atlas leaves out matches with a similarity score below 0.5.
* The ESM Atlas computes the query's features if it doesn't have them.
* As a side effect, the ESM Atlas can predict structures for up to six hits that don't have one yet.
  The result contains no structures.
* Every hit includes its sequence, so the assistant can pass it to the other tools.

### Returns

| Field | Description |
| - | - |
| `query_sequence_length` | Length of your sequence. |
| `hits[]` | Cluster representatives, most similar first. |
| `hits[].protein_accession` | Representative's accession. |
| `hits[].sequence`, `sequence_length` | Representative's sequence and length. |
| `hits[].similarity_score` | Cosine similarity of the SAE feature vectors, from 0 to 1. |
| `hits[].cluster_size`, `protein_name` | Cluster size and name, when available. |
| `top_features_across_results[]` | Features common across the hits: `feature_index`, `occurrence_count`, and `mean_activation`. |
| `restricted_count` | Number of hits withheld by ESM Atlas [guardrails](/learn/guides/esm-atlas#guardrails). |

The REST equivalent is [similarity search](/api/protein/atlas/similarity-search).

## Get cluster details

`esm_atlas_get_cluster_info` summarizes the cluster that a representative sequence belongs to.

### Inputs

| Name | Type | Required | Limits | Default |
| - | - | - | - | - |
| `sequence` | string | Yes | 1 to 4,000 residues | |

The sequence must be a cluster representative, such as a hit from [`esm_atlas_search_similar_protein_clusters`](#find-similar-protein-clusters).
Other sequences, including many well-known UniProt proteins, return `not_found`.
This tool reads stored data only and never computes.

### Returns

| Field | Description |
| - | - |
| `accession`, `source`, `protein_name` | Representative's identity. |
| `cluster_size` | Number of proteins in the cluster. |
| `cluster_pct_characterized` | Percentage of members with characterized Pfam annotations, 0 to 100. |
| `cluster_mean_domain_coverage` | Mean Pfam domain coverage across members, from 0 to 1. |
| `top_pfam_domains` | Map of Pfam ID to `count` and `name`, or `null`. |
| `representative_features[]` | Up to ten SAE features with `feature_index`, `label`, `description`, and `value`. |
| `taxonomy_info` | Lowest common ancestor `rank` and `name`, or `null`. |
| `top_phyla` | Map of phylum to member count, or `null`. |

The REST equivalent is [get cluster](/api/protein/atlas/clusters).

## Get SAE feature details

`esm_atlas_get_sae_feature_detail` explains one of the 16,384 SAE features learned from [ESMC](/models/esmc).

### Inputs

| Name | Type | Required | Limits | Default |
| - | - | - | - | - |
| `feature_index` | integer | Yes | 0 to 16,383 | |

### Returns

| Field | Description |
| - | - |
| `feature_index` | Feature number. |
| `label`, `summary`, `description` | Short label and longer interpretations. |
| `category` | Broad category, such as `Compositional bias`, when available. |
| `activation_pattern` | Where and how the feature tends to fire, when available. |
| `exemplar_protein_families` | Example families, when available. |
| `uniref90_frequency` | How often the feature appears across UniRef90. |
| `top_swissprot_activations[]` | Up to ten SwissProt proteins with the strongest activation: `uniprot_id` and `activation`. |
| `decoder_nearest_neighbors[]` | Up to five related feature numbers. |

The REST equivalent is [get SAE feature](/api/protein/atlas/features/detail).

## Errors

A failed call returns a JSON object with a stable `code` and a `message`, for example:

```json theme={null}
{"code":"invalid_input","message":"Unrecognized accession format. Supported databases are UniParc, MGnify, and IMG."}
```

| Code | Returned by | Typical cause |
| - | - | - |
| `invalid_input` | All | An input breaks a type, format, or limit. |
| `not_found` | Lookup, cluster, and feature tools | The accession, cluster, or feature doesn't exist. |
| `restricted` | Tools that call the ESM Atlas | ESM Atlas guardrails blocked the request. |
| `rate_limited` | Tools that call the ESM Atlas | The ESM Atlas asked callers to slow down. The error includes `retry_after_seconds`. |
| `dependency_unavailable` | All | The ESM Atlas or an external database is unavailable or sent an unusable response. |
| `indeterminate` | All | The 120-second deadline passed, or the outcome of a compute step is unknown. |
| `sequence_too_long_for_features` | `esm_atlas_get_protein_details` | A miss longer than 2,048 residues. |
| `resource_too_large` | All | The complete result is over 6 MiB. |
| `internal_error` | All | An unexpected server failure. |

[Troubleshooting](/biohub-mcp/troubleshooting#error-codes) explains what to do about each code.
