Before you begin
Prepare a bounded list of valid amino-acid sequences and a Biohub Platform API key.1
Create the client and embedding function
2
Submit the sequence collection
3
Verify every output
Documentation Index
Fetch the complete documentation index at: /llms.txt
Use this file to discover all available pages before exploring further.
Use the esm SDK parallel executor to generate ESMC representations for multiple protein sequences.
Create the client and embedding function
from getpass import getpass
from esm.sdk import esmc_client, parallel_executor
from esm.sdk.api import ESMCInferenceClient, ESMProtein, LogitsConfig, LogitsOutput
token = getpass("Biohub API key: ")
model = esmc_client(
model="esmc-300m-2024-12",
url="https://biohub.ai",
token=token,
)
embedding_config = LogitsConfig(
sequence=True,
return_hidden_states=False,
return_mean_hidden_states=True,
)
def embed_sequence(model: ESMCInferenceClient, sequence: str) -> LogitsOutput:
protein_tensor = model.encode(ESMProtein(sequence=sequence))
return model.logits(protein_tensor, embedding_config)
Submit the sequence collection
sequences = [
"MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFAGKQLEDGRTLSDYNIQKESTLHLVLRLRGG",
"MKTIIALSYIFCLVFADYKDDDDK",
]
with parallel_executor() as executor:
outputs = executor.execute_batch(
user_func=embed_sequence,
model=model,
sequence=sequences,
)
Verify every output
for index, output in enumerate(outputs, start=1):
embedding = output.mean_hidden_state.float().squeeze()
print(index, embedding.shape)