spicrawlspicrawlDocs

Scrape thousands of URLs with a batch job

Submit up to 10,000 URLs per call to POST /v1/batch, poll the job until it is terminal, and page the JSONL results before they expire after 72 hours.

Use this when you have a list of URLs to fetch and do not need each answer immediately: a nightly catalogue refresh, a site archive, a price sweep. You submit the list once, get a job id back with 202, and the platform works through it with its own retries while you poll. Requires an API key with the batch scope.

Batch workers currently apply only js_render, proxy and block_resources, send every request as GET, and return the raw page (HTML, or the decoded body for non-rendered fetches). Every other /v1/scrape field is validated and priced but has no effect: no markdown, no extract, no ai_extract, no screenshots, no actions, no sessions. If you need markdown or extraction per URL, run /v1/scrape with client-side concurrency instead:

spicrawl scrape - --concurrency 8 --format markdown --jsonl < urls.txt > pages.jsonl

engine: "chromium" is rejected on batch with 400 ERR::REQUEST::INVALID_PARAMETER; use /v1/scrape for Chromium.

Minimal request

curl https://api.spicrawl.com/v1/batch \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "catalog-2026-09-22",
    "urls": ["https://example.com/products/1", "https://example.com/products/2"],
    "js_render": true
  }'

urls gives every URL the job-level settings. items gives each URL its own object with overrides and an external_id you can correlate on: {"url": "https://example.com/products/2", "external_id": "sku-42", "js_render": true}. Send one or the other, never both (400 ERR::REQUEST::INCOMPATIBLE_FLAGS). An item's value wins over the job-level value, including an explicit false. The CLI reads a URL-per-line file, a .json array of items, or a .jsonl file of items.

What comes back

Submission returns 202 with the job object: id, status: "queued", status_url, results_url, total_items, estimated_credits (the hold) and progress. Read warnings: it reports a clamped concurrency, or items not yet handed to the queue. Do not resubmit in that case; a recovery sweep dispatches them, and a resubmission is a second job billed again.

Poll GET /v1/batch/{id} every 2–5 seconds for jobs under about 1,000 items, every 10–30 seconds for larger ones. Progress counters are cheap to read.

StatusMeaningKeep polling?
queuedAccepted, not started.Yes
runningItems in flight.Yes
pausedOperator state; no public endpoint pauses a job.Yes
cancellingCancel accepted; in-flight items are finishing.Yes
completedEvery item finished.No
failedAborted, e.g. failure_threshold reached. See error_code.No
cancelledCancelled and drained.No

GET /v1/batch/{id}/results returns finished items as JSON Lines (application/x-ndjson), one object per line, in seq order:

{"seq":1,"external_id":"sku-42","url":"https://example.com/products/2","status":"succeeded","attempts":1,"http_status":200,"request_id":"01J9Z7A1B2C3D4E5F6G7H8J9KA","credits_micro":4000000,"bytes":326000,"duration_ms":1840,"result":{"content":"\"<!doctype html>…\"","bytes":326002},"finished_at":"2026-09-22T06:41:14Z"}
  • status is succeeded, failed (read error.code and error.retryable), cancelled or skipped. Unfinished items are never listed.
  • http_status is the target's status. A 404 here is a successful scrape of a missing page.
  • result.content is a JSON string literal holding the page. Decode it once to get the HTML.
  • result is absent while the job is still running: bodies are inlined from a file written when the job becomes terminal.
  • Paging is in headers. If X-Next-Cursor is present, request again with cursor=<value> (the Link header has the full URL). limit is 1–5,000, default 500. Filter with status=failed to list only failures.

Body caps: 512 KiB per item and 2 MiB per page, spent in seq order. Exactly one of these applies to each result:

FieldMeaningWhat to do
contentFull body.Use it.
truncated: truecontent is a prefix and not valid JSON.Request that seq alone (cursor just before it, limit=1) for up to 512 KiB.
omitted: trueThe page's 2 MiB budget was spent.Lower limit, or resume with a cursor just before this seq.
unavailableThe result file could not be read.Retry later.

Results are readable until results_expire_at, 72 hours after submission by default. After that the results endpoint answers 410 ERR::REQUEST::BEYOND_RETENTION; the job object stays readable.

Limits and options

  • Body up to 1 MiB (413), 10,000 items per call, job-level parameters up to 64 KiB serialised, each item's overrides up to 8 KiB.
  • A bad item fails the whole submission; detail starts with Item N:.
  • open: true keeps the job accepting items via POST /v1/batch/{id}/items (same body shape, 10,000 per call, 100,000 per job). An open job never completes on its own: call POST /v1/batch/{id}/close when done. Appended items use the job's original settings plus their own overrides.
FieldDefaultEffect
concurrency10Items in flight at once. Above 50 is clamped to 50 with a warning.
priority1000–1000, relative to your other jobs.
max_attempts31–10 attempts per item; only retryable failures are re-attempted.
failure_thresholdnoneAbort the job as failed after this many failed items.
credit_budgetprojected cost (none for an open job)Ceiling on the whole run, in credits. Items beyond it are skipped.
max_costnonePer-item ceiling, checked at submission.
cachefalseBatch never reads from the result cache, unlike /v1/scrape.
webhook_endpoint_idnoneA registered webhook endpoint notified with the job object when it finishes. Fetch results yourself.

Retry and cancel

  • POST /v1/batch/{id}/retry resets every failed item to queued. Succeeded, cancelled and skipped items never re-run, and attempts is not reset. A terminal job goes back to queued. With nothing failed, it returns items_reset: 0.
  • POST /v1/batch/{id}/cancel moves a live job to cancelling; in-flight items finish (and are charged if they succeed), then the job becomes cancelled. Cancelling a terminal job is 409.

CLI: spicrawl batch retry <id>, spicrawl batch cancel <id>, spicrawl batch append <id> more.txt, spicrawl batch close <id>, spicrawl batch wait <id> --max-wait 30m (exit 11 on timeout).

Failure modes

CodeHTTPWhat to do
ERR::REQUEST::INVALID_PARAMETER400A field or an item is invalid, or engine: "chromium". detail names the item.
ERR::REQUEST::MISSING_PARAMETER400Neither urls nor items.
ERR::REQUEST::PAYLOAD_TOO_LARGE413Body over 1 MiB. Split the list, or use an open job with appends.
ERR::LIMIT::QUOTA_EXCEEDED402The job's dearest cost does not fit under your monthly credit allowance. Submit fewer items or set max_cost.
ERR::REQUEST::CONFLICT409Appending to a closed job, or cancelling a terminal one.
ERR::REQUEST::BEYOND_RETENTION410Results expired. Resubmit if you still need them.

Per-item failures are on the result line in error.code, from the same catalogue as /v1/scrape (ERR::UPSTREAM::TIMEOUT, ERR::PROXY::EXHAUSTED, ...). Re-run them with retry.

Cost

Submitting is free (X-Credits-Charged: 0) but reserves the dearest possible cost of every item against your monthly allowance; that hold is estimated_credits, and the response's X-Credits-Remaining is the allowance left after it. Append and retry responses carry X-Credits-Remaining after their holds too. Credits held for items that fail, are cancelled or skipped are returned to the monthly allowance within about a minute of the item finishing. Appending to an open job reserves the appended items the same way, and is refused with 402 if the allowance cannot cover them. Each item is charged on success only, at the /v1/scrape price for its engine (1 credit on fetch, 3 with js_render). Billable target statuses are 200, 404 and 410. Retrying a failed item (retry) holds its price again first and is refused with 402 if the allowance cannot cover it. Actual spend is progress.credits_charged. See Credits.

On this page