Scrape thousands of URLs with a batch job
Submit up to 10,000 URLs per call to POST /v1/batch, poll the job until it is terminal, and page the JSONL results before they expire after 72 hours.
Use this when you have a list of URLs to fetch and do not need each answer immediately: a nightly catalogue refresh, a site archive, a price sweep. You submit the list once, get a job id back with 202, and the platform works through it with its own retries while you poll. Requires an API key with the batch scope.
Batch workers currently apply only js_render, proxy and block_resources, send every request as GET, and return the raw page (HTML, or the decoded body for non-rendered fetches). Every other /v1/scrape field is validated and priced but has no effect: no markdown, no extract, no ai_extract, no screenshots, no actions, no sessions. If you need markdown or extraction per URL, run /v1/scrape with client-side concurrency instead:
spicrawl scrape - --concurrency 8 --format markdown --jsonl < urls.txt > pages.jsonlengine: "chromium" is rejected on batch with 400 ERR::REQUEST::INVALID_PARAMETER; use /v1/scrape for Chromium.
Minimal request
curl https://api.spicrawl.com/v1/batch \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "catalog-2026-09-22",
"urls": ["https://example.com/products/1", "https://example.com/products/2"],
"js_render": true
}'urls gives every URL the job-level settings. items gives each URL its own object with overrides and an external_id you can correlate on: {"url": "https://example.com/products/2", "external_id": "sku-42", "js_render": true}. Send one or the other, never both (400 ERR::REQUEST::INCOMPATIBLE_FLAGS). An item's value wins over the job-level value, including an explicit false. The CLI reads a URL-per-line file, a .json array of items, or a .jsonl file of items.
What comes back
Submission returns 202 with the job object: id, status: "queued", status_url, results_url, total_items, estimated_credits (the hold) and progress. Read warnings: it reports a clamped concurrency, or items not yet handed to the queue. Do not resubmit in that case; a recovery sweep dispatches them, and a resubmission is a second job billed again.
Poll GET /v1/batch/{id} every 2–5 seconds for jobs under about 1,000 items, every 10–30 seconds for larger ones. Progress counters are cheap to read.
| Status | Meaning | Keep polling? |
|---|---|---|
queued | Accepted, not started. | Yes |
running | Items in flight. | Yes |
paused | Operator state; no public endpoint pauses a job. | Yes |
cancelling | Cancel accepted; in-flight items are finishing. | Yes |
completed | Every item finished. | No |
failed | Aborted, e.g. failure_threshold reached. See error_code. | No |
cancelled | Cancelled and drained. | No |
GET /v1/batch/{id}/results returns finished items as JSON Lines (application/x-ndjson), one object per line, in seq order:
{"seq":1,"external_id":"sku-42","url":"https://example.com/products/2","status":"succeeded","attempts":1,"http_status":200,"request_id":"01J9Z7A1B2C3D4E5F6G7H8J9KA","credits_micro":4000000,"bytes":326000,"duration_ms":1840,"result":{"content":"\"<!doctype html>…\"","bytes":326002},"finished_at":"2026-09-22T06:41:14Z"}statusissucceeded,failed(readerror.codeanderror.retryable),cancelledorskipped. Unfinished items are never listed.http_statusis the target's status. A 404 here is a successful scrape of a missing page.result.contentis a JSON string literal holding the page. Decode it once to get the HTML.resultis absent while the job is still running: bodies are inlined from a file written when the job becomes terminal.- Paging is in headers. If
X-Next-Cursoris present, request again withcursor=<value>(theLinkheader has the full URL).limitis 1–5,000, default 500. Filter withstatus=failedto list only failures.
Body caps: 512 KiB per item and 2 MiB per page, spent in seq order. Exactly one of these applies to each result:
| Field | Meaning | What to do |
|---|---|---|
content | Full body. | Use it. |
truncated: true | content is a prefix and not valid JSON. | Request that seq alone (cursor just before it, limit=1) for up to 512 KiB. |
omitted: true | The page's 2 MiB budget was spent. | Lower limit, or resume with a cursor just before this seq. |
unavailable | The result file could not be read. | Retry later. |
Results are readable until results_expire_at, 72 hours after submission by default. After that the results endpoint answers 410 ERR::REQUEST::BEYOND_RETENTION; the job object stays readable.
Limits and options
- Body up to 1 MiB (
413), 10,000 items per call, job-level parameters up to 64 KiB serialised, each item's overrides up to 8 KiB. - A bad item fails the whole submission;
detailstarts withItem N:. open: truekeeps the job accepting items viaPOST /v1/batch/{id}/items(same body shape, 10,000 per call, 100,000 per job). An open job never completes on its own: callPOST /v1/batch/{id}/closewhen done. Appended items use the job's original settings plus their own overrides.
| Field | Default | Effect |
|---|---|---|
concurrency | 10 | Items in flight at once. Above 50 is clamped to 50 with a warning. |
priority | 100 | 0–1000, relative to your other jobs. |
max_attempts | 3 | 1–10 attempts per item; only retryable failures are re-attempted. |
failure_threshold | none | Abort the job as failed after this many failed items. |
credit_budget | projected cost (none for an open job) | Ceiling on the whole run, in credits. Items beyond it are skipped. |
max_cost | none | Per-item ceiling, checked at submission. |
cache | false | Batch never reads from the result cache, unlike /v1/scrape. |
webhook_endpoint_id | none | A registered webhook endpoint notified with the job object when it finishes. Fetch results yourself. |
Retry and cancel
POST /v1/batch/{id}/retryresets everyfaileditem to queued. Succeeded, cancelled and skipped items never re-run, andattemptsis not reset. A terminal job goes back toqueued. With nothing failed, it returnsitems_reset: 0.POST /v1/batch/{id}/cancelmoves a live job tocancelling; in-flight items finish (and are charged if they succeed), then the job becomescancelled. Cancelling a terminal job is409.
CLI: spicrawl batch retry <id>, spicrawl batch cancel <id>, spicrawl batch append <id> more.txt, spicrawl batch close <id>, spicrawl batch wait <id> --max-wait 30m (exit 11 on timeout).
Failure modes
| Code | HTTP | What to do |
|---|---|---|
ERR::REQUEST::INVALID_PARAMETER | 400 | A field or an item is invalid, or engine: "chromium". detail names the item. |
ERR::REQUEST::MISSING_PARAMETER | 400 | Neither urls nor items. |
ERR::REQUEST::PAYLOAD_TOO_LARGE | 413 | Body over 1 MiB. Split the list, or use an open job with appends. |
ERR::LIMIT::QUOTA_EXCEEDED | 402 | The job's dearest cost does not fit under your monthly credit allowance. Submit fewer items or set max_cost. |
ERR::REQUEST::CONFLICT | 409 | Appending to a closed job, or cancelling a terminal one. |
ERR::REQUEST::BEYOND_RETENTION | 410 | Results expired. Resubmit if you still need them. |
Per-item failures are on the result line in error.code, from the same catalogue as /v1/scrape (ERR::UPSTREAM::TIMEOUT, ERR::PROXY::EXHAUSTED, ...). Re-run them with retry.
Cost
Submitting is free (X-Credits-Charged: 0) but reserves the dearest possible cost of every item against your monthly allowance; that hold is estimated_credits, and the response's X-Credits-Remaining is the allowance left after it. Append and retry responses carry X-Credits-Remaining after their holds too. Credits held for items that fail, are cancelled or skipped are returned to the monthly allowance within about a minute of the item finishing. Appending to an open job reserves the appended items the same way, and is refused with 402 if the allowance cannot cover them. Each item is charged on success only, at the /v1/scrape price for its engine (1 credit on fetch, 3 with js_render). Billable target statuses are 200, 404 and 410. Retrying a failed item (retry) holds its price again first and is refused with 402 if the allowance cannot cover it. Actual spend is progress.credits_charged. See Credits.
Related
- CLI batch commands
- Markdown for LLMs for per-URL markdown with
/v1/scrape - Proxies and geo
- Rate limits