spicrawlspicrawlDocs

spicrawl batch

Submit up to 10,000 URLs as one asynchronous batch job, wait for it, and download the results as JSON Lines with spicrawl batch.

spicrawl batch runs many URLs as one server-side job (POST /v1/batch). You submit a file of URLs, poll until the job finishes, then download one JSON line per item.

spicrawl batch submit urls.txt --render --name nightly --wait --max-wait 20m
spicrawl batch results 01J9Z6V0Q8M4K2T7R3N5B1C9XA --all -o results.jsonl

Batch workers currently apply only js_render (--render), proxy (--proxy) and block_resources (--block), and every item returns the raw page HTML. Other scrape flags (--format markdown, --extract, --ai, --screenshot, --actions, ...) are validated by the API but not applied. batch submit prints a note: on stderr naming the fields it will not apply. For markdown or extraction per URL, use spicrawl scrape - with --concurrency.

When to use batch instead of scrape -

spicrawl scrape - --concurrency Nspicrawl batch submit
RunsOn your machine, N requests at a timeServer-side, up to 50 at a time
Survives your process exitingNoYes; resume with batch wait
OutputMarkdown, extraction, screenshots: every scrape flagRaw page HTML only (today)
SizeAnyUp to 10,000 targets and 1 MiB per submit

Submit a job

spicrawl batch submit <file|-> [scrape flags] [job flags]

Input formats

InputRead as
A file ending .jsonA JSON array of URL strings, or of item objects.
A file ending .jsonl or .ndjsonOne item object per line.
Any other fileOne URL per line. Blank lines and # comments are skipped.
- (stdin)Sniffed: a leading [ is a JSON array, a leading { is JSONL, anything else is URL lines.

An item is {"url": "...", "external_id": "..."} plus any per-item scrape overrides. external_id is your correlation id; it comes back on the item's result line. An item without url exits 2.

items.jsonl
{"url":"https://example.com/products/1","external_id":"sku-1"}
{"url":"https://example.com/products/2","external_id":"sku-2","js_render":false}
spicrawl batch submit items.jsonl --render --proxy "$MY_PROXY_URL"

Scrape flags set job-wide defaults; item fields override them. The scrape flags are the same as spicrawl scrape, except --wait (see below).

Job flags

FlagAPI fieldDefaultMeaning
--name NnamenoneLabel shown on the job and in the dashboard.
--concurrency Nconcurrency10Items in flight at once. Values above 50 are clamped to 50 with a warning.
--priority Npriority1000-1000, relative to your other jobs.
--max-attempts Nmax_attempts3Attempts per item, 1-10. Only retryable failures are re-attempted.
--failure-threshold Nfailure_thresholdnoneAbort the job as failed once N items have failed.
--credit-budget Ncredit_budgetprojected costCeiling on the whole run, in credits.
--openopenoffKeep accepting items with batch append until batch close.
--waitnoneoffPoll until the job finishes, then print the final job.
--poll Dnone5sWith --wait: poll interval.
--max-wait Dnone30mWith --wait: give up after this long and exit 11. 0 waits forever.

On batch submit and batch wait, --wait means "wait for the job" (not the scrape wait delay in milliseconds), and --max-wait is the wait deadline; the global --timeout stays the per-call HTTP timeout. --wait --open exits 2: an open job completes only after batch close.

Output

Without --wait, submit prints the job as soon as it is accepted. In JSON mode that is the API's BatchJob object unchanged:

id=$(spicrawl batch submit urls.txt --render | jq -r .id)

In human mode it prints id, status, progress, credits, submitted and a next: spicrawl batch wait <id> hint on stderr.

Wait for a job

spicrawl batch wait 01J9Z6V0Q8M4K2T7R3N5B1C9XA --max-wait 10m
spicrawl batch wait "$id" --poll 10s --json | jq '.progress'

Polls GET /v1/batch/{id} every --poll (default 5s) until the status is completed, failed or cancelled, then prints the job.

  • In human mode a progress line on stderr shows batch <id>: running 412/1000 (41.2%) succeeded 405 failed 7. JSON mode and --quiet print no progress.
  • A 429, or a retryable 503, does not stop the wait: the next poll waits for the server's Retry-After.
  • If --max-wait (default 30m) elapses first, it prints the last job seen and exits 11 with batch <id> is still running after 10m0s; resume with: spicrawl batch wait <id>. Run the same command again to keep waiting.
spicrawl batch wait "$id" --max-wait 5m
case $? in
  0)  echo "finished" ;;
  11) echo "still running; try again later" ;;
  *)  echo "failed" >&2 ;;
esac

A job whose status is failed still exits 0 from wait: the wait itself succeeded. Check .status and .error_code in the printed job.

Inspect jobs

spicrawl batch get 01J9Z6V0Q8M4K2T7R3N5B1C9XA --json | jq '{status, progress}'
spicrawl batch list --status running
spicrawl batch list --all --json | jq -r '.batches[] | select(.status=="failed") | .id'
CommandFlags
batch get <id>none. Prints the job and its progress (total, completed, succeeded, failed, cancelled, skipped, percent_complete, credits_charged).
batch list--status S (queued, running, paused, cancelling, completed, failed, cancelled), --limit N (1-200, default 50), --cursor C, --all (follow next_cursor to the oldest job).

batch list prints {"batches": [...], "next_cursor": "..."} in JSON mode, and an ID STATUS NAME DONE SUBMITTED table otherwise. next_cursor is present only when more pages exist and you did not pass --all.

Download results

spicrawl batch results <id> [--status S] [--all] [--limit N] [--cursor C] [-o FILE]

Fetches GET /v1/batch/{id}/results: one JSON object per finished item, in seq order. In JSON mode, or with -o, the lines are written unchanged as JSONL; in human mode a SEQ STATUS HTTP URL BYTES table is printed.

FlagDefaultMeaning
--status SallOnly succeeded, failed, cancelled or skipped items.
--alloffFollow X-Next-Cursor to the last finished item. Without it you get one page and a more results: --cursor ... hint on stderr.
--limit N500 (server)Lines per page, 1-5000. Lower it if bodies come back omitted.
--cursor CnoneResume from a previous page's cursor.
-o FILEstdoutWrite the JSONL to FILE.

Each line has seq, external_id, url, status, attempts, http_status, request_id, bytes, duration_ms, result and error. result.content is the stored page as JSON text (a JSON string holding the HTML), so decode it once:

# Everything, to a file
spicrawl batch results "$id" --all -o results.jsonl

# Only failures, with their error codes
spicrawl batch results "$id" --status failed --all | jq -r '[.seq, .url, .error.code] | @tsv'

# The HTML of every succeeded item
spicrawl batch results "$id" --status succeeded --all | jq -r '.result.content | fromjson'

Bodies are inlined only once the job is terminal. Results expire at the job's results_expire_at (72 hours after submission by default); after that the API answers ERR::REQUEST::BEYOND_RETENTION (HTTP 410, exit 4).

Change a running job

CommandAPIDoes
batch cancel <id>POST /v1/batch/{id}/cancelMoves the job to cancelling; in-flight items drain, the rest are cancelled, then the job is cancelled.
batch retry <id>POST /v1/batch/{id}/retryResets every failed item to queued and dispatches it again. Prints items_reset and items_dispatched.
batch close <id>POST /v1/batch/{id}/closeStops an --open job accepting items so it can complete.
batch append <id> <file|->POST /v1/batch/{id}/itemsAdds items to an --open job. Same input formats as submit. Prints items_added and items_dispatched.
spicrawl batch retry "$id" && spicrawl batch wait "$id" --max-wait 15m
spicrawl batch cancel "$id" && spicrawl batch wait "$id" --max-wait 2m

After retry or append, items not dispatched immediately are picked up by the worker's recovery sweep. Do not call retry again for them, and never re-append them: they would run twice.

Open jobs for crawls

id=$(spicrawl batch submit seeds.txt --open --name crawl | jq -r .id)
echo https://example.com/discovered/17 | spicrawl batch append "$id" -
spicrawl batch append "$id" more.jsonl
spicrawl batch close "$id" && spicrawl batch wait "$id"

Errors

You seeExitDo next
item 3: missing "url"2Fix the input file; nothing was sent.
ERR::REQUEST::INCOMPATIBLE_FLAGS on submit4Both urls and items, or conflicting scrape flags. Read detail.
ERR::REQUEST::PAYLOAD_TOO_LARGE4Split into submits of at most 10,000 targets and 1 MiB.
ERR::LIMIT::QUOTA_EXCEEDED5Out of credits for the period; see spicrawl usage summary.
ERR::REQUEST::BEYOND_RETENTION on results4The results expired; resubmit.
Wait timed out11Run spicrawl batch wait <id> again.

See the batch guide for the job lifecycle and webhooks.

On this page