Skip to content

Batch APIs Explained

Cheatsheet of LLM batch APIs: what they trade (latency for cost), the request anatomy, and the limits that shape a real batch job.

Batch APIs answer one question: do you need the answer now? If not, you trade minutes-to-hours of latency for roughly half the price and much higher throughput caps.

Reference table · 14 entries
14 of 14 rows
Batch vs realtime
CostBatch jobs bill at roughly 50% of realtime prices.$1.00 / 1M → $0.50 / 1M
LatencyResults in minutes to 24h, not milliseconds.complete within 24h
ThroughputFar higher caps — designed for backlogs, not conversations.100k+ requests per file
Rate limitsSeparate (higher) quotas; TPM still applies per file.batch_tpm, not chat_tpm
Anatomy of a batch job
JSONL fileOne request per line, each with a custom id you choose.{"custom_id": "job-1", "body": {…}}
uploadSend the file, get a file id back.files.create(purpose='batch')
create batchPoint the batch at the file + endpoint.batches.create(input_file_id, endpoint)
poll statusvalidating → in_progress → completed/failed/expired.batches.retrieve(id)
download resultsJSONL out — match rows by custom_id.files.content(output_file_id)
Limits & gotchas
File size capInput files cap around 100–200 MB depending on provider.split large jobs into files
ExpirationOutput files expire (days–months); download early.download within 30 days
No streamingBatch is all-or-nothing per request — no partial reads.n/a
Per-row errorsOne bad request fails that row, not the file — check per-row status.row.status === 'failed'
Same pricing quirksCached input tokens and reasoning tokens bill by the same rules as realtime.check usage per row