| Batch vs realtime |
|---|
| Cost | Batch jobs bill at roughly 50% of realtime prices. | $1.00 / 1M → $0.50 / 1M |
| Latency | Results in minutes to 24h, not milliseconds. | complete within 24h |
| Throughput | Far higher caps — designed for backlogs, not conversations. | 100k+ requests per file |
| Rate limits | Separate (higher) quotas; TPM still applies per file. | batch_tpm, not chat_tpm |
| Anatomy of a batch job |
|---|
| JSONL file | One request per line, each with a custom id you choose. | {"custom_id": "job-1", "body": {…}} |
| upload | Send the file, get a file id back. | files.create(purpose='batch') |
| create batch | Point the batch at the file + endpoint. | batches.create(input_file_id, endpoint) |
| poll status | validating → in_progress → completed/failed/expired. | batches.retrieve(id) |
| download results | JSONL out — match rows by custom_id. | files.content(output_file_id) |
| Limits & gotchas |
|---|
| File size cap | Input files cap around 100–200 MB depending on provider. | split large jobs into files |
| Expiration | Output files expire (days–months); download early. | download within 30 days |
| No streaming | Batch is all-or-nothing per request — no partial reads. | n/a |
| Per-row errors | One bad request fails that row, not the file — check per-row status. | row.status === 'failed' |
| Same pricing quirks | Cached input tokens and reasoning tokens bill by the same rules as realtime. | check usage per row |