spicrawl batch
Submit up to 10,000 URLs as one asynchronous batch job, wait for it, and download the results as JSON Lines with spicrawl batch.
spicrawl batch runs many URLs as one server-side job (POST /v1/batch). You submit a file of URLs, poll until the job finishes, then download one JSON line per item.
spicrawl batch submit urls.txt --render --name nightly --wait --max-wait 20m
spicrawl batch results 01J9Z6V0Q8M4K2T7R3N5B1C9XA --all -o results.jsonlBatch workers currently apply only js_render (--render), proxy (--proxy) and block_resources (--block), and every item returns the raw page HTML. Other scrape flags (--format markdown, --extract, --ai, --screenshot, --actions, ...) are validated by the API but not applied. batch submit prints a note: on stderr naming the fields it will not apply. For markdown or extraction per URL, use spicrawl scrape - with --concurrency.
When to use batch instead of scrape -
spicrawl scrape - --concurrency N | spicrawl batch submit | |
|---|---|---|
| Runs | On your machine, N requests at a time | Server-side, up to 50 at a time |
| Survives your process exiting | No | Yes; resume with batch wait |
| Output | Markdown, extraction, screenshots: every scrape flag | Raw page HTML only (today) |
| Size | Any | Up to 10,000 targets and 1 MiB per submit |
Submit a job
spicrawl batch submit <file|-> [scrape flags] [job flags]Input formats
| Input | Read as |
|---|---|
A file ending .json | A JSON array of URL strings, or of item objects. |
A file ending .jsonl or .ndjson | One item object per line. |
| Any other file | One URL per line. Blank lines and # comments are skipped. |
- (stdin) | Sniffed: a leading [ is a JSON array, a leading { is JSONL, anything else is URL lines. |
An item is {"url": "...", "external_id": "..."} plus any per-item scrape overrides. external_id is your correlation id; it comes back on the item's result line. An item without url exits 2.
{"url":"https://example.com/products/1","external_id":"sku-1"}
{"url":"https://example.com/products/2","external_id":"sku-2","js_render":false}spicrawl batch submit items.jsonl --render --proxy "$MY_PROXY_URL"Scrape flags set job-wide defaults; item fields override them. The scrape flags are the same as spicrawl scrape, except --wait (see below).
Job flags
| Flag | API field | Default | Meaning |
|---|---|---|---|
--name N | name | none | Label shown on the job and in the dashboard. |
--concurrency N | concurrency | 10 | Items in flight at once. Values above 50 are clamped to 50 with a warning. |
--priority N | priority | 100 | 0-1000, relative to your other jobs. |
--max-attempts N | max_attempts | 3 | Attempts per item, 1-10. Only retryable failures are re-attempted. |
--failure-threshold N | failure_threshold | none | Abort the job as failed once N items have failed. |
--credit-budget N | credit_budget | projected cost | Ceiling on the whole run, in credits. |
--open | open | off | Keep accepting items with batch append until batch close. |
--wait | none | off | Poll until the job finishes, then print the final job. |
--poll D | none | 5s | With --wait: poll interval. |
--max-wait D | none | 30m | With --wait: give up after this long and exit 11. 0 waits forever. |
On batch submit and batch wait, --wait means "wait for the job" (not the scrape wait delay in milliseconds), and --max-wait is the wait deadline; the global --timeout stays the per-call HTTP timeout. --wait --open exits 2: an open job completes only after batch close.
Output
Without --wait, submit prints the job as soon as it is accepted. In JSON mode that is the API's BatchJob object unchanged:
id=$(spicrawl batch submit urls.txt --render | jq -r .id)In human mode it prints id, status, progress, credits, submitted and a next: spicrawl batch wait <id> hint on stderr.
Wait for a job
spicrawl batch wait 01J9Z6V0Q8M4K2T7R3N5B1C9XA --max-wait 10m
spicrawl batch wait "$id" --poll 10s --json | jq '.progress'Polls GET /v1/batch/{id} every --poll (default 5s) until the status is completed, failed or cancelled, then prints the job.
- In human mode a progress line on stderr shows
batch <id>: running 412/1000 (41.2%) succeeded 405 failed 7. JSON mode and--quietprint no progress. - A
429, or a retryable503, does not stop the wait: the next poll waits for the server'sRetry-After. - If
--max-wait(default30m) elapses first, it prints the last job seen and exits11withbatch <id> is still running after 10m0s; resume with: spicrawl batch wait <id>. Run the same command again to keep waiting.
spicrawl batch wait "$id" --max-wait 5m
case $? in
0) echo "finished" ;;
11) echo "still running; try again later" ;;
*) echo "failed" >&2 ;;
esacA job whose status is failed still exits 0 from wait: the wait itself succeeded. Check .status and .error_code in the printed job.
Inspect jobs
spicrawl batch get 01J9Z6V0Q8M4K2T7R3N5B1C9XA --json | jq '{status, progress}'
spicrawl batch list --status running
spicrawl batch list --all --json | jq -r '.batches[] | select(.status=="failed") | .id'| Command | Flags |
|---|---|
batch get <id> | none. Prints the job and its progress (total, completed, succeeded, failed, cancelled, skipped, percent_complete, credits_charged). |
batch list | --status S (queued, running, paused, cancelling, completed, failed, cancelled), --limit N (1-200, default 50), --cursor C, --all (follow next_cursor to the oldest job). |
batch list prints {"batches": [...], "next_cursor": "..."} in JSON mode, and an ID STATUS NAME DONE SUBMITTED table otherwise. next_cursor is present only when more pages exist and you did not pass --all.
Download results
spicrawl batch results <id> [--status S] [--all] [--limit N] [--cursor C] [-o FILE]Fetches GET /v1/batch/{id}/results: one JSON object per finished item, in seq order. In JSON mode, or with -o, the lines are written unchanged as JSONL; in human mode a SEQ STATUS HTTP URL BYTES table is printed.
| Flag | Default | Meaning |
|---|---|---|
--status S | all | Only succeeded, failed, cancelled or skipped items. |
--all | off | Follow X-Next-Cursor to the last finished item. Without it you get one page and a more results: --cursor ... hint on stderr. |
--limit N | 500 (server) | Lines per page, 1-5000. Lower it if bodies come back omitted. |
--cursor C | none | Resume from a previous page's cursor. |
-o FILE | stdout | Write the JSONL to FILE. |
Each line has seq, external_id, url, status, attempts, http_status, request_id, bytes, duration_ms, result and error. result.content is the stored page as JSON text (a JSON string holding the HTML), so decode it once:
# Everything, to a file
spicrawl batch results "$id" --all -o results.jsonl
# Only failures, with their error codes
spicrawl batch results "$id" --status failed --all | jq -r '[.seq, .url, .error.code] | @tsv'
# The HTML of every succeeded item
spicrawl batch results "$id" --status succeeded --all | jq -r '.result.content | fromjson'Bodies are inlined only once the job is terminal. Results expire at the job's results_expire_at (72 hours after submission by default); after that the API answers ERR::REQUEST::BEYOND_RETENTION (HTTP 410, exit 4).
Change a running job
| Command | API | Does |
|---|---|---|
batch cancel <id> | POST /v1/batch/{id}/cancel | Moves the job to cancelling; in-flight items drain, the rest are cancelled, then the job is cancelled. |
batch retry <id> | POST /v1/batch/{id}/retry | Resets every failed item to queued and dispatches it again. Prints items_reset and items_dispatched. |
batch close <id> | POST /v1/batch/{id}/close | Stops an --open job accepting items so it can complete. |
batch append <id> <file|-> | POST /v1/batch/{id}/items | Adds items to an --open job. Same input formats as submit. Prints items_added and items_dispatched. |
spicrawl batch retry "$id" && spicrawl batch wait "$id" --max-wait 15m
spicrawl batch cancel "$id" && spicrawl batch wait "$id" --max-wait 2mAfter retry or append, items not dispatched immediately are picked up by the worker's recovery sweep. Do not call retry again for them, and never re-append them: they would run twice.
Open jobs for crawls
id=$(spicrawl batch submit seeds.txt --open --name crawl | jq -r .id)
echo https://example.com/discovered/17 | spicrawl batch append "$id" -
spicrawl batch append "$id" more.jsonl
spicrawl batch close "$id" && spicrawl batch wait "$id"Errors
| You see | Exit | Do next |
|---|---|---|
item 3: missing "url" | 2 | Fix the input file; nothing was sent. |
ERR::REQUEST::INCOMPATIBLE_FLAGS on submit | 4 | Both urls and items, or conflicting scrape flags. Read detail. |
ERR::REQUEST::PAYLOAD_TOO_LARGE | 4 | Split into submits of at most 10,000 targets and 1 MiB. |
ERR::LIMIT::QUOTA_EXCEEDED | 5 | Out of credits for the period; see spicrawl usage summary. |
ERR::REQUEST::BEYOND_RETENTION on results | 4 | The results expired; resubmit. |
| Wait timed out | 11 | Run spicrawl batch wait <id> again. |
See the batch guide for the job lifecycle and webhooks.