# spicrawl batch

> Submit up to 10,000 URLs as one asynchronous batch job, wait for it, and download the results as JSON Lines with spicrawl batch.

Source: https://docs.spicrawl.com/cli/batch

`spicrawl batch` runs many URLs as one server-side job (`POST /v1/batch`). You submit a file of URLs, poll until the job finishes, then download one JSON line per item.

```bash
spicrawl batch submit urls.txt --render --name nightly --wait --max-wait 20m
spicrawl batch results 01J9Z6V0Q8M4K2T7R3N5B1C9XA --all -o results.jsonl
```

> **Warning:** Batch workers currently apply only `js_render` (`--render`), `proxy` (`--proxy`) and `block_resources` (`--block`), and every item returns the **raw page HTML**. Other scrape flags (`--format markdown`, `--extract`, `--ai`, `--screenshot`, `--actions`, ...) are validated by the API but not applied. `batch submit` prints a `note:` on stderr naming the fields it will not apply. For markdown or extraction per URL, use [`spicrawl scrape -`](https://docs.spicrawl.com/cli/scrape.md#many-urls-from-stdin) with `--concurrency`.

## When to use batch instead of `scrape -`

|                               | `spicrawl scrape - --concurrency N`                  | `spicrawl batch submit`                   |
| ----------------------------- | ---------------------------------------------------- | ----------------------------------------- |
| Runs                          | On your machine, N requests at a time                | Server-side, up to 50 at a time           |
| Survives your process exiting | No                                                   | Yes; resume with `batch wait`             |
| Output                        | Markdown, extraction, screenshots: every scrape flag | Raw page HTML only (today)                |
| Size                          | Any                                                  | Up to 10,000 targets and 1 MiB per submit |

## Submit a job

```bash
spicrawl batch submit <file|-> [scrape flags] [job flags]
```

### Input formats

| Input                               | Read as                                                                                     |
| ----------------------------------- | ------------------------------------------------------------------------------------------- |
| A file ending `.json`               | A JSON array of URL strings, or of item objects.                                            |
| A file ending `.jsonl` or `.ndjson` | One item object per line.                                                                   |
| Any other file                      | One URL per line. Blank lines and `#` comments are skipped.                                 |
| `-` (stdin)                         | Sniffed: a leading `[` is a JSON array, a leading `{` is JSONL, anything else is URL lines. |

An item is `{"url": "...", "external_id": "..."}` plus any per-item scrape overrides. `external_id` is your correlation id; it comes back on the item's result line. An item without `url` exits `2`.

```jsonl title="items.jsonl"
{"url":"https://example.com/products/1","external_id":"sku-1"}
{"url":"https://example.com/products/2","external_id":"sku-2","js_render":false}
```

```bash
spicrawl batch submit items.jsonl --render --proxy "$MY_PROXY_URL"
```

Scrape flags set job-wide defaults; item fields override them. The scrape flags are the same as [`spicrawl scrape`](https://docs.spicrawl.com/cli/scrape.md#flags), except `--wait` (see below).

### Job flags

| Flag                    | API field           | Default        | Meaning                                                                    |
| ----------------------- | ------------------- | -------------- | -------------------------------------------------------------------------- |
| `--name N`              | `name`              | none           | Label shown on the job and in the dashboard.                               |
| `--concurrency N`       | `concurrency`       | `10`           | Items in flight at once. Values above 50 are clamped to 50 with a warning. |
| `--priority N`          | `priority`          | `100`          | 0-1000, relative to your other jobs.                                       |
| `--max-attempts N`      | `max_attempts`      | `3`            | Attempts per item, 1-10. Only retryable failures are re-attempted.         |
| `--failure-threshold N` | `failure_threshold` | none           | Abort the job as `failed` once N items have failed.                        |
| `--credit-budget N`     | `credit_budget`     | projected cost | Ceiling on the whole run, in credits.                                      |
| `--open`                | `open`              | off            | Keep accepting items with `batch append` until `batch close`.              |
| `--wait`                | none                | off            | Poll until the job finishes, then print the final job.                     |
| `--poll D`              | none                | `5s`           | With `--wait`: poll interval.                                              |
| `--max-wait D`          | none                | `30m`          | With `--wait`: give up after this long and exit `11`. `0` waits forever.   |

> **Note:** On `batch submit` and `batch wait`, `--wait` means "wait for the job" (not the scrape `wait` delay in milliseconds), and `--max-wait` is the wait deadline; the global `--timeout` stays the per-call HTTP timeout. `--wait --open` exits `2`: an open job completes only after `batch close`.

### Output

Without `--wait`, `submit` prints the job as soon as it is accepted. In JSON mode that is the API's `BatchJob` object unchanged:

```bash
id=$(spicrawl batch submit urls.txt --render | jq -r .id)
```

In human mode it prints `id`, `status`, `progress`, `credits`, `submitted` and a `next: spicrawl batch wait <id>` hint on stderr.

## Wait for a job

```bash
spicrawl batch wait 01J9Z6V0Q8M4K2T7R3N5B1C9XA --max-wait 10m
spicrawl batch wait "$id" --poll 10s --json | jq '.progress'
```

Polls `GET /v1/batch/{id}` every `--poll` (default `5s`) until the status is `completed`, `failed` or `cancelled`, then prints the job.

* In human mode a progress line on stderr shows `batch <id>: running 412/1000 (41.2%)  succeeded 405  failed 7`. JSON mode and `--quiet` print no progress.
* A `429`, or a retryable `503`, does not stop the wait: the next poll waits for the server's `Retry-After`.
* If `--max-wait` (default `30m&#x60;) elapses first, it prints the last job seen and exits &#x2A;*`11`** with `batch <id> is still running after 10m0s; resume with: spicrawl batch wait <id>`. Run the same command again to keep waiting.

```bash
spicrawl batch wait "$id" --max-wait 5m
case $? in
  0)  echo "finished" ;;
  11) echo "still running; try again later" ;;
  *)  echo "failed" >&2 ;;
esac
```

A job whose status is `failed` still exits `0` from `wait`: the wait itself succeeded. Check `.status` and `.error_code` in the printed job.

## Inspect jobs

```bash
spicrawl batch get 01J9Z6V0Q8M4K2T7R3N5B1C9XA --json | jq '{status, progress}'
spicrawl batch list --status running
spicrawl batch list --all --json | jq -r '.batches[] | select(.status=="failed") | .id'
```

| Command          | Flags                                                                                                                                                                                            |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `batch get <id>` | none. Prints the job and its `progress` (`total`, `completed`, `succeeded`, `failed`, `cancelled`, `skipped`, `percent_complete`, `credits_charged`).                                            |
| `batch list`     | `--status S` (`queued`, `running`, `paused`, `cancelling`, `completed`, `failed`, `cancelled`), `--limit N` (1-200, default 50), `--cursor C`, `--all` (follow `next_cursor` to the oldest job). |

`batch list` prints `{"batches": [...], "next_cursor": "..."}` in JSON mode, and an `ID STATUS NAME DONE SUBMITTED` table otherwise. `next_cursor` is present only when more pages exist and you did not pass `--all`.

## Download results

```bash
spicrawl batch results <id> [--status S] [--all] [--limit N] [--cursor C] [-o FILE]
```

Fetches `GET /v1/batch/{id}/results`: one JSON object per finished item, in `seq` order. In JSON mode, or with `-o`, the lines are written unchanged as JSONL; in human mode a `SEQ STATUS HTTP URL BYTES` table is printed.

| Flag         | Default      | Meaning                                                                                                                          |
| ------------ | ------------ | -------------------------------------------------------------------------------------------------------------------------------- |
| `--status S` | all          | Only `succeeded`, `failed`, `cancelled` or `skipped` items.                                                                      |
| `--all`      | off          | Follow `X-Next-Cursor` to the last finished item. Without it you get one page and a `more results: --cursor ...` hint on stderr. |
| `--limit N`  | 500 (server) | Lines per page, 1-5000. Lower it if bodies come back `omitted`.                                                                  |
| `--cursor C` | none         | Resume from a previous page's cursor.                                                                                            |
| `-o FILE`    | stdout       | Write the JSONL to `FILE`.                                                                                                       |

Each line has `seq`, `external_id`, `url`, `status`, `attempts`, `http_status`, `request_id`, `bytes`, `duration_ms`, `result` and `error`. `result.content` is the stored page as JSON text (a JSON string holding the HTML), so decode it once:

```bash
# Everything, to a file
spicrawl batch results "$id" --all -o results.jsonl

# Only failures, with their error codes
spicrawl batch results "$id" --status failed --all | jq -r '[.seq, .url, .error.code] | @tsv'

# The HTML of every succeeded item
spicrawl batch results "$id" --status succeeded --all | jq -r '.result.content | fromjson'
```

Bodies are inlined only once the job is terminal. Results expire at the job's `results_expire_at` (72 hours after submission by default); after that the API answers `ERR::REQUEST::BEYOND_RETENTION` (HTTP 410, exit `4`).

## Change a running job

| Command                       | API                          | Does                                                                                                        |
| ----------------------------- | ---------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `batch cancel <id>`           | `POST /v1/batch/{id}/cancel` | Moves the job to `cancelling`; in-flight items drain, the rest are cancelled, then the job is `cancelled`.  |
| `batch retry <id>`            | `POST /v1/batch/{id}/retry`  | Resets every failed item to queued and dispatches it again. Prints `items_reset` and `items_dispatched`.    |
| `batch close <id>`            | `POST /v1/batch/{id}/close`  | Stops an `--open` job accepting items so it can complete.                                                   |
| `batch append <id> <file\|->` | `POST /v1/batch/{id}/items`  | Adds items to an `--open` job. Same input formats as `submit`. Prints `items_added` and `items_dispatched`. |

```bash
spicrawl batch retry "$id" && spicrawl batch wait "$id" --max-wait 15m
spicrawl batch cancel "$id" && spicrawl batch wait "$id" --max-wait 2m
```

> **Warning:** After `retry` or `append`, items not dispatched immediately are picked up by the worker's recovery sweep. Do not call `retry` again for them, and never re-append them: they would run twice.

### Open jobs for crawls

```bash
id=$(spicrawl batch submit seeds.txt --open --name crawl | jq -r .id)
echo https://example.com/discovered/17 | spicrawl batch append "$id" -
spicrawl batch append "$id" more.jsonl
spicrawl batch close "$id" && spicrawl batch wait "$id"
```

## Errors

| You see                                      | Exit | Do next                                                              |
| -------------------------------------------- | ---- | -------------------------------------------------------------------- |
| `item 3: missing "url"`                      | `2`  | Fix the input file; nothing was sent.                                |
| `ERR::REQUEST::INCOMPATIBLE_FLAGS` on submit | `4`  | Both `urls` and `items`, or conflicting scrape flags. Read `detail`. |
| `ERR::REQUEST::PAYLOAD_TOO_LARGE`            | `4`  | Split into submits of at most 10,000 targets and 1 MiB.              |
| `ERR::LIMIT::QUOTA_EXCEEDED`                 | `5`  | Out of credits for the period; see `spicrawl usage summary`.         |
| `ERR::REQUEST::BEYOND_RETENTION` on results  | `4`  | The results expired; resubmit.                                       |
| Wait timed out                               | `11` | Run `spicrawl batch wait <id>` again.                                |

See the [batch guide](https://docs.spicrawl.com/guides/batch.md) for the job lifecycle and webhooks.
