> ## Documentation Index
> Fetch the complete documentation index at: https://docs.biohub.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Retrieve protein and cluster context

> Resolve a UniProt accession, retrieve its Atlas protein record, and inspect its feature-based cluster.

This guide moves from a UniProt accession to its Atlas protein record and then to the report for the nearest feature-based cluster.

<Steps>
  <Step title="Resolve the UniProt accession">
    ```python theme={null}
    import httpx

    base_url = "https://biohub.ai/esm/protein/api/v1alpha1"

    with httpx.Client(timeout=60) as client:
        response = client.get(f"{base_url}/uniprot/P24941")
        response.raise_for_status()
        resolved = response.json()

    protein_hash = resolved["protein_hash"]
    print(resolved["accession"], resolved["protein_name"], protein_hash)
    ```

    Confirm the identity before using the returned hash.
  </Step>

  <Step title="Retrieve the protein record">
    ```python theme={null}
    with httpx.Client(timeout=60) as client:
        response = client.get(
            f"{base_url}/proteins/{protein_hash}",
            params={"topk_features": 5, "fold_on_miss": False},
        )
        response.raise_for_status()
        protein = response.json()

    print(protein["accession"], protein["sequence_length"])
    for feature in protein["sae_features"]:
        print(feature["feature_index"], feature["label"])
    ```

    The record carries the strongest SAE features for the sequence.
    `ptm` and `mean_plddt` are `None` when the Atlas has no stored structure for this exact sequence.
    Read the [protein endpoint reference](/api/protein/atlas/proteins/get) before changing feature selection or on-demand folding options.
  </Step>

  <Step title="Find the nearest cluster representative">
    ```python theme={null}
    cluster_hash = protein["cluster_rep_protein_hash"]
    if cluster_hash is None:
        with httpx.Client(timeout=120) as client:
            response = client.get(
                f"{base_url}/similarity-search",
                params={
                    "sequence": resolved["sequence"],
                    "topk_results": 1,
                    "include_cluster_info": True,
                },
            )
            response.raise_for_status()
            nearest = response.json()["similar_proteins"][0]

        cluster_hash = nearest["protein_hash"]
        print(nearest["protein_name"], "similarity", round(nearest["similarity_score"], 3))
    ```

    Only proteins inside an Atlas cluster carry `cluster_rep_protein_hash`.
    Most UniProt proteins do not, so the nearest cluster representative by SAE feature similarity stands in.
  </Step>

  <Step title="Retrieve the cluster report">
    ```python theme={null}
    with httpx.Client(timeout=60) as client:
        response = client.get(f"{base_url}/clusters/{cluster_hash}")
        response.raise_for_status()
        cluster = response.json()

    print("Cluster size:", cluster["cluster_size"])
    print("Characterized percent:", cluster["cluster_pct_characterized"])
    print("Top Pfam domains:", cluster["cluster_top_pfam_domains"])
    print("Taxonomy:", cluster["cluster_taxonomy_info"])
    ```

    Use the [cluster endpoint reference](/api/protein/atlas/clusters) for the complete response contract.
  </Step>
</Steps>

Treat the protein and cluster records as Atlas evidence.
Feature labels and cluster similarity do not independently prove protein function.
