# Scrape thousands of URLs with a batch job

> Submit up to 10,000 URLs per call to POST /v1/batch, poll the job until it is terminal, and page the JSONL results before they expire after 72 hours.

Source: https://docs.spicrawl.com/guides/batch

Use this when you have a list of URLs to fetch and do not need each answer immediately: a nightly catalogue refresh, a site archive, a price sweep. You submit the list once, get a job id back with `202`, and the platform works through it with its own retries while you poll. Requires an API key with the `batch` scope.

> **Warning:** Batch workers currently apply only `js_render`, `proxy` and `block_resources`, send every request as GET, and return the **raw page** (HTML, or the decoded body for non-rendered fetches). Every other `/v1/scrape` field is validated and priced but has no effect: no markdown, no `extract`, no `ai_extract`, no screenshots, no actions, no sessions. If you need markdown or extraction per URL, run `/v1/scrape` with client-side concurrency instead:
>
> ```bash
> spicrawl scrape - --concurrency 8 --format markdown --jsonl < urls.txt > pages.jsonl
> ```
>
> `engine: "chromium"` is rejected on batch with `400 ERR::REQUEST::INVALID_PARAMETER`; use `/v1/scrape` for Chromium.

## Minimal request

```bash title="curl"
curl https://api.spicrawl.com/v1/batch \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "catalog-2026-09-22",
    "urls": ["https://example.com/products/1", "https://example.com/products/2"],
    "js_render": true
  }'
```

```python title="Python"
import os, time, json, requests

API = "https://api.spicrawl.com"
H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}

job = requests.post(f"{API}/v1/batch", headers=H, json={
    "name": "catalog-2026-09-22",
    "urls": ["https://example.com/products/1", "https://example.com/products/2"],
    "js_render": True,
}, timeout=60).json()

while job["status"] not in ("completed", "failed", "cancelled"):
    time.sleep(5)
    job = requests.get(f"{API}/v1/batch/{job['id']}", headers=H, timeout=30).json()

cursor = None
while True:
    params = {"cursor": cursor} if cursor else {}
    r = requests.get(f"{API}/v1/batch/{job['id']}/results", headers=H, params=params, timeout=120)
    r.raise_for_status()
    for line in r.text.splitlines():
        item = json.loads(line)
        if item["status"] == "succeeded" and "content" in item.get("result", {}):
            html = json.loads(item["result"]["content"])  # content is a JSON string literal
    cursor = r.headers.get("X-Next-Cursor")
    if not cursor:
        break
```

```typescript title="TypeScript"
const API = "https://api.spicrawl.com";
const H = { Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`, "Content-Type": "application/json" };

let job = await (await fetch(`${API}/v1/batch`, {
  method: "POST",
  headers: H,
  body: JSON.stringify({
    name: "catalog-2026-09-22",
    urls: ["https://example.com/products/1", "https://example.com/products/2"],
    js_render: true,
  }),
})).json();

while (!["completed", "failed", "cancelled"].includes(job.status)) {
  await new Promise((r) => setTimeout(r, 5000));
  job = await (await fetch(`${API}/v1/batch/${job.id}`, { headers: H })).json();
}

let cursor: string | null = null;
do {
  const q = cursor ? `?cursor=${cursor}` : "";
  const r = await fetch(`${API}/v1/batch/${job.id}/results${q}`, { headers: H });
  for (const line of (await r.text()).split("\n").filter(Boolean)) {
    const item = JSON.parse(line);
    if (item.status === "succeeded" && item.result?.content) {
      const html: string = JSON.parse(item.result.content);
    }
  }
  cursor = r.headers.get("X-Next-Cursor");
} while (cursor);
```

```bash title="CLI"
spicrawl batch submit urls.txt --render --name catalog-2026-09-22 --wait --max-wait 20m
spicrawl batch results 01J9Z6T3W9E21T5TZARVJRVN5C --all -o results.jsonl
```

`urls` gives every URL the job-level settings. `items` gives each URL its own object with overrides and an `external_id` you can correlate on: `{"url": "https://example.com/products/2", "external_id": "sku-42", "js_render": true}`. Send one or the other, never both (`400 ERR::REQUEST::INCOMPATIBLE_FLAGS`). An item's value wins over the job-level value, including an explicit `false`. The CLI reads a URL-per-line file, a `.json` array of items, or a `.jsonl` file of items.

## What comes back

Submission returns `202` with the job object: `id`, `status: "queued"`, `status_url`, `results_url`, `total_items`, `estimated_credits` (the hold) and `progress`. Read `warnings`: it reports a clamped `concurrency`, or items not yet handed to the queue. **Do not resubmit** in that case; a recovery sweep dispatches them, and a resubmission is a second job billed again.

Poll `GET /v1/batch/{id}` every 2–5 seconds for jobs under about 1,000 items, every 10–30 seconds for larger ones. Progress counters are cheap to read.

| Status       | Meaning                                                      | Keep polling? |
| ------------ | ------------------------------------------------------------ | ------------- |
| `queued`     | Accepted, not started.                                       | Yes           |
| `running`    | Items in flight.                                             | Yes           |
| `paused`     | Operator state; no public endpoint pauses a job.             | Yes           |
| `cancelling` | Cancel accepted; in-flight items are finishing.              | Yes           |
| `completed`  | Every item finished.                                         | No            |
| `failed`     | Aborted, e.g. `failure_threshold` reached. See `error_code`. | No            |
| `cancelled`  | Cancelled and drained.                                       | No            |

`GET /v1/batch/{id}/results` returns finished items as JSON Lines (`application/x-ndjson`), one object per line, in `seq` order:

```json
{"seq":1,"external_id":"sku-42","url":"https://example.com/products/2","status":"succeeded","attempts":1,"http_status":200,"request_id":"01J9Z7A1B2C3D4E5F6G7H8J9KA","credits_micro":4000000,"bytes":326000,"duration_ms":1840,"result":{"content":"\"<!doctype html>…\"","bytes":326002},"finished_at":"2026-09-22T06:41:14Z"}
```

* `status` is `succeeded`, `failed` (read `error.code` and `error.retryable`), `cancelled` or `skipped`. Unfinished items are never listed.
* `http_status` is the target's status. A 404 here is a successful scrape of a missing page.
* `result.content` is a JSON string literal holding the page. Decode it once to get the HTML.
* `result` is absent while the job is still running: bodies are inlined from a file written when the job becomes terminal.
* Paging is in headers. If `X-Next-Cursor` is present, request again with `cursor=<value>` (the `Link` header has the full URL). `limit` is 1–5,000, default 500. Filter with `status=failed` to list only failures.

Body caps: 512 KiB per item and 2 MiB per page, spent in `seq` order. Exactly one of these applies to each `result`:

| Field             | Meaning                                   | What to do                                                                     |
| ----------------- | ----------------------------------------- | ------------------------------------------------------------------------------ |
| `content`         | Full body.                                | Use it.                                                                        |
| `truncated: true` | `content` is a prefix and not valid JSON. | Request that `seq` alone (cursor just before it, `limit=1`) for up to 512 KiB. |
| `omitted: true`   | The page's 2 MiB budget was spent.        | Lower `limit`, or resume with a cursor just before this `seq`.                 |
| `unavailable`     | The result file could not be read.        | Retry later.                                                                   |

Results are readable until `results_expire_at`, 72 hours after submission by default. After that the results endpoint answers `410 ERR::REQUEST::BEYOND_RETENTION`; the job object stays readable.

## Limits and options

* Body up to 1 MiB (`413`), 10,000 items per call, job-level parameters up to 64 KiB serialised, each item's overrides up to 8 KiB.
* A bad item fails the whole submission; `detail` starts with `Item N:`.
* `open: true` keeps the job accepting items via `POST /v1/batch/{id}/items` (same body shape, 10,000 per call, 100,000 per job). An open job never completes on its own: call `POST /v1/batch/{id}/close` when done. Appended items use the job's original settings plus their own overrides.

| Field                 | Default                                 | Effect                                                                                               |
| --------------------- | --------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `concurrency`         | `10`                                    | Items in flight at once. Above 50 is clamped to 50 with a warning.                                   |
| `priority`            | `100`                                   | 0–1000, relative to your other jobs.                                                                 |
| `max_attempts`        | `3`                                     | 1–10 attempts per item; only retryable failures are re-attempted.                                    |
| `failure_threshold`   | none                                    | Abort the job as `failed` after this many failed items.                                              |
| `credit_budget`       | projected cost (none for an `open` job) | Ceiling on the whole run, in credits. Items beyond it are `skipped`.                                 |
| `max_cost`            | none                                    | Per-item ceiling, checked at submission.                                                             |
| `cache`               | `false`                                 | Batch never reads from the result cache, unlike `/v1/scrape`.                                        |
| `webhook_endpoint_id` | none                                    | A registered webhook endpoint notified with the job object when it finishes. Fetch results yourself. |

## Retry and cancel

* `POST /v1/batch/{id}/retry` resets every `failed` item to queued. Succeeded, cancelled and skipped items never re-run, and `attempts` is not reset. A terminal job goes back to `queued`. With nothing failed, it returns `items_reset: 0`.
* `POST /v1/batch/{id}/cancel` moves a live job to `cancelling`; in-flight items finish (and are charged if they succeed), then the job becomes `cancelled`. Cancelling a terminal job is `409`.

CLI: `spicrawl batch retry <id>`, `spicrawl batch cancel <id>`, `spicrawl batch append <id> more.txt`, `spicrawl batch close <id>`, `spicrawl batch wait <id> --max-wait 30m` (exit 11 on timeout).

## Failure modes

| Code                              | HTTP | What to do                                                                                                     |
| --------------------------------- | ---- | -------------------------------------------------------------------------------------------------------------- |
| `ERR::REQUEST::INVALID_PARAMETER` | 400  | A field or an item is invalid, or `engine: "chromium"`. `detail` names the item.                               |
| `ERR::REQUEST::MISSING_PARAMETER` | 400  | Neither `urls` nor `items`.                                                                                    |
| `ERR::REQUEST::PAYLOAD_TOO_LARGE` | 413  | Body over 1 MiB. Split the list, or use an open job with appends.                                              |
| `ERR::LIMIT::QUOTA_EXCEEDED`      | 402  | The job's dearest cost does not fit under your monthly credit allowance. Submit fewer items or set `max_cost`. |
| `ERR::REQUEST::CONFLICT`          | 409  | Appending to a closed job, or cancelling a terminal one.                                                       |
| `ERR::REQUEST::BEYOND_RETENTION`  | 410  | Results expired. Resubmit if you still need them.                                                              |

Per-item failures are on the result line in `error.code`, from the same catalogue as `/v1/scrape` (`ERR::UPSTREAM::TIMEOUT`, `ERR::PROXY::EXHAUSTED`, ...). Re-run them with `retry`.

## Cost

Submitting is free (`X-Credits-Charged: 0`) but reserves the dearest possible cost of every item against your monthly allowance; that hold is `estimated_credits`, and the response's `X-Credits-Remaining` is the allowance left after it. Append and retry responses carry `X-Credits-Remaining` after their holds too. Credits held for items that fail, are cancelled or skipped are returned to the monthly allowance within about a minute of the item finishing. Appending to an open job reserves the appended items the same way, and is refused with `402` if the allowance cannot cover them. Each item is charged on success only, at the `/v1/scrape` price for its engine (1 credit on fetch, 3 with `js_render`). Billable target statuses are 200, 404 and 410. Retrying a failed item (`retry`) holds its price again first and is refused with `402` if the allowance cannot cover it. Actual spend is `progress.credits_charged`. See [Credits](https://docs.spicrawl.com/credits.md).

## Related

* [CLI batch commands](https://docs.spicrawl.com/cli/batch.md)
* [Markdown for LLMs](https://docs.spicrawl.com/guides/markdown.md) for per-URL markdown with `/v1/scrape`
* [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md)
* [Rate limits](https://docs.spicrawl.com/rate-limits.md)
