> ## Documentation Index
> Fetch the complete documentation index at: https://docs.biohub.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Batch API

> Submit inference jobs, track their progress, and retrieve results asynchronously.

The Batch API supports larger and longer-running asynchronous workloads.
Submit inputs, poll the batch, then retrieve each completed item's result.

| Operation | Purpose |
| - | - |
| [Submit](/api/protein/models/batch_submit) | Queue inputs and return a batch ID. |
| [List](/api/protein/models/batch_list) | Read batch status, counts, and pages of items. |
| [Status](/api/protein/models/batch_status) | Read one task and retrieve its result and artifact URLs. |
| [Cancel](/api/protein/models/batch_cancel) | Cancel queued tasks. |

## Submit inputs

Choose the inference operation with `endpoint` and put its inputs in `payload`.
Each payload entry becomes a task.

```bash theme={null}
curl https://biohub.ai/api/v1/batch/submit \
  -H "Authorization: Bearer $BIOHUB_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "endpoint": "fold_max_accuracy",
    "payload": [
      {
        "model": "esmfold2-2026-05-cutoff-2025",
        "all_atom_input": {
          "sequences": [
            {
              "id": "A",
              "type": "protein",
              "sequence": "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQ"
            }
          ]
        },
        "include_pae": true
      }
    ]
  }'
```

Save `data.batch_id` from the response.
A one-item batch also returns `data.task_id`.
Task IDs append a zero-padded index, such as `<batch_id>-00000`.
For repeated payloads, ordering is by repeat first, then position within `payload`.

## Track progress

Poll [List](/api/protein/models/batch_list) with the saved `batch_id`:

```bash theme={null}
curl https://biohub.ai/api/v1/batch/list \
  -H "Authorization: Bearer $BIOHUB_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"batch_id": "<batch_id>"}'
```

The aggregate batch state is `in_progress`, `done`, `failed`, or `cancelled`.
An aggregate `done` means all tasks are terminal; some or all tasks may have failed.
A cancelled batch can still contain running tasks.

| Task state | Meaning |
| - | - |
| `queued` | Waiting to start. |
| `in_progress` | Processing has started. |
| `done` | A result is available. |
| `failed` | Processing failed. Read the task's error. |
| `cancelled` | Cancelled before processing started, by you or by the 24-hour queue limit. |

`counts` covers the whole batch and omits states with zero items.
Set `include_responses: true` to retrieve an item page, then follow the returned `cursor` until it is null.
The default page size is 100 and the maximum is 1000.
A full final page can require one additional empty-page request.
A `status` filter narrows the item page, while counts remain batch-wide.

Tasks that have not started 24 hours after the batch was submitted are cancelled and never charged; tasks already running finish normally.
Status and List return this deadline as `data.expires_at`.

## Retrieve results

Call [Status](/api/protein/models/batch_status) with a `task_id` from the submission response or an item page:

```bash theme={null}
curl https://biohub.ai/api/v1/batch/status \
  -H "Authorization: Bearer $BIOHUB_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"task_id": "<task_id>"}'
```

When the task is `done`, `data.response` contains an inline result object, inline text, or a signed URL to download it, depending on the endpoint.
For `fold_max_accuracy`, it is a signed URL for the folding-result JSON.
Other outputs, when present, are signed URLs in `data.artifacts`, keyed by artifact name.
List responses can include inline results but omit artifact URLs; externally stored results are indicated by `response_available: true`.

Download signed URLs without the Biohub bearer token.
Read status again to refresh expired URLs.
When present, `data.artifacts_expire_in` gives the shortest artifact URL lifetime in seconds.

## Result retention

Results are kept for 31 days after the batch is submitted, after which they are deleted.
Status and List return the deadline as `data.results_expire_at`. After that, Status returns `410` and List still reports each item's `status` but omits `response`, `response_available`, and `error`.

## Responses and failures

A successful HTTP operation returns this envelope:

```json theme={null}
{"status": "success", "data": {}, "request_id": "<request-id>"}
```

Route errors carry `message` instead:

```json theme={null}
{"status": "error", "message": "Batch not found", "request_id": "<request-id>"}
```

Authentication failures can use `error` instead of `message`; handle both shapes.
Retain `request_id` when reporting a failed call.
The envelope's `status` describes the HTTP operation, while `data.status` describes the job.
An HTTP `200` can report a failed task; inspect `data.error` on the status response.
Failure to read status does not mean the task failed.

Passing a batch ID to Status returns `404` with instructions to use List or an item ID.

## Cancel

[Cancel](/api/protein/models/batch_cancel) accepts exactly one of `task_id` or `batch_id`.
It affects queued tasks only. Running tasks continue and retain their usage charges.
A batch with no cancellable work returns `409`; cancelling an already cancelled batch succeeds.
