# Append items to an open job

> Adds items to a job created with `open: true`. Same body shape as batchCreate: `urls` or `items`, at most 10,000 per call; a job holds at most 100,000 items across all appends (exceeding it is a 400).

Source: https://docs.spicrawl.com/api-reference/batch/batch-append

## POST /v1/batch/{batchID}/items

Operation ID: `batchAppend`. API key scope: `batch`.

Adds items to a job created with `open: true`. Same body shape as batchCreate: `urls` or `items`, at most 10,000
per call; a job holds at most 100,000 items across all appends (exceeding it is a 400). New items get `seq`
numbers continuing from the job's current total.
Only per-item values are stored: a new item runs with the job's ORIGINAL job-level parameters plus its own
overrides. Top-level scrape parameters in this body are used only to validate the new items and are otherwise
discarded; run controls (`name`, `concurrency`, `priority`, `max_attempts`, `failure_threshold`, `credit_budget`,
`webhook_endpoint_id`, `open`) are accepted and ignored. Appended items are held against the monthly allowance
exactly like submitted ones (`402` if it cannot cover them, and then nothing is added); `estimated_credits`
stays the submission's and is not updated.
409 when the job is closed (never opened, or already closed). When finished appending, call
`POST /v1/batch/{batchID}/close`.

### Example

```bash
curl -X POST "https://api.spicrawl.com/v1/batch/{batchID}/items" \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"items":[{"url":"https://example.com/discovered/17","external_id":"page-17"},{"url":"https://example.com/discovered/18"}]}'
```

### Parameters

| Name | In | Type | Required | Description |
|---|---|---|---|---|
| `batchID` | path | string | yes | The job `id` (26-character ULID as returned). The 32/36-character UUID form of the same id is also accepted. A malformed id and another tenant's id both answer 404. |

### Request body

`application/json`, required.

| Field | Type | Required | Description |
|---|---|---|---|
| `name` | string | no | Free-text label shown on the job object and in the dashboard. No uniqueness; omit for none. |
| `urls` | array of string (uri) | no | Shorthand list of targets that all use the job-level settings. Mutually exclusive with `items` (both → 400 INCOMPATIBLE_FLAGS; neither → 400 MISSING_PARAMETER). Order defines `seq` (0-based). |
| `items` | array of object (`BatchItem`) | no | Long form, one object per target with optional per-item overrides and `external_id`. Mutually exclusive with `urls`. Order defines `seq` (0-based). |
| `concurrency` | integer; default `10` | no | Max items of this job in flight at once. Values below 1 become 1; values above 50 are clamped to 50 and reported in the job's `warnings` (not an error). |
| `priority` | integer; default `100` | no | Scheduling priority relative to your other jobs. Outside 0-1000 is a 400. |
| `max_attempts` | integer; default `3` | no | Attempts per item before it is marked failed (only retryable failures are re-attempted). Outside 1-10 is a 400. |
| `failure_threshold` | integer (int64) | no | Abort the job as `failed` once this many items have failed. Omit to never abort early. 0 or negative is a 400. |
| `credit_budget` | integer | no | Ceiling on the WHOLE run in whole credits; the worker stops charging work beyond it. Omit to use the job's projected cost; an `open` job then has no ceiling, because its projection covers only the items it was submitted with. Stored as given in `params.credit_budget_micro`. 0 or negative is a 400. Different from `max_cost`, which is per item. |
| `webhook_endpoint_id` | string | no | Id of a webhook endpoint registered for this organization; it receives the job object when the job finishes (notify-only — fetch results yourself). A value that is not a valid id is a 400; blank is treated as absent. |
| `open` | boolean; default `false` | no | Keep the job accepting more items via POST /v1/batch/{batchID}/items. An open job never completes on its own — call POST /v1/batch/{batchID}/close when done. Total across appends is capped at 100,000 items. |
| `method` | string: `GET`, `POST`, `PUT`, `PATCH`, `DELETE`, `HEAD`, `OPTIONS`; default `"GET"` | no | HTTP method sent to the target (case-insensitive). There is no request-body parameter, so non-GET methods are sent without a body; only GET/HEAD/OPTIONS are retried on transient failure and only GET results are cached. |
| `js_render` | boolean; default `false` | no | Render in a browser engine (obscura; 3 credits datacenter, 25 residential) so JavaScript runs. Required by every browser-only flag unless stealth or mode=auto is set; cannot be combined with mode=auto. If omitted, the project's stored default applies. |
| `stealth` | boolean; default `false` | no | Coming soon. Render in the hardened camoufox browser, flat 25 credits at any proxy tier; wins over js_render when both are set. Cannot be combined with mode=auto or with an engine pin other than camoufox; returns 503 ERR::ENGINE::UNAVAILABLE where camoufox is not deployed. If omitted, the project's stored default applies. |
| `impersonate` | boolean; default `true` | no | Fetch tier only: present a current Chrome TLS/HTTP2 fingerprint, with matching browser headers and Accept-Encoding, to pass passive fingerprint checks. On by default, free, no JavaScript; `false` turns it off and the fetch uses a non-browser TLS handshake. Ignored on render engines. If omitted, the project's stored default applies, and `true` when the project sets none. |
| `mode` | string: `auto` | no | `auto` escalates through the available engines (fetch -> obscura) until a rung returns a usable page, billing only the rung that succeeded (0 if all fail) while reserving the dearest rung against quota. Incompatible with js_render, stealth and engine. If omitted, the project's stored default applies. |
| `engine` | string: `fetch`, `obscura`, `camoufox`, `chromium` | no | Engine `camoufox` is coming soon. Pin the execution engine (case-insensitive); disables escalation and is never substituted. chromium (8 credits datacenter, 32 residential) is the only engine supporting headless=false; screenshots and pdf need camoufox or chromium. An unentitled engine is 403 ERR::AUTH::ENGINE_NOT_ENTITLED, an undeployed one 503 ERR::ENGINE::UNAVAILABLE, both at 0 credits. If omitted, the project's stored default applies. |
| `premium_proxy` | boolean; default `false` | no | Coming soon. Use residential exits from the managed pool: fetch 10, obscura 25, chromium 32 credits (camoufox stays 25). If no residential exit is available the request fails 503 ERR::PROXY::EXHAUSTED rather than falling back; ignored (with X-Warning PROXY_FLAG_IGNORED) when `proxy` is set. A bot challenge from a named vendor is retried on a different exit (fetch tier), then climbs to camoufox; see "Bot challenges" in the operation description. If omitted, the project's stored default applies. |
| `proxy_country` | string | no | Coming soon. ISO 3166-1 alpha-2 exit country (case-insensitive), within the residential pool. Requires premium_proxy=true (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise) unless `proxy` is set, in which case it is ignored with a warning; an unavailable country is 503 ERR::PROXY::EXHAUSTED, never substituted. A browser render takes its timezone and locale from this country or, when it is omitted, from the country of the pool exit it was assigned. If omitted, the project's stored default applies. |
| `proxy` | string | no | Your own proxy URL (http, https, socks5 or socks5h; credentials in userinfo). Takes precedence over premium_proxy/proxy_country and adds 0 proxy credits; loopback hosts are refused with 400 ERR::SECURITY::SSRF_BLOCKED. An unreachable proxy fails 502 ERR::PROXY::UNREACHABLE. A bot challenge from a named vendor climbs to camoufox when the proxy is http://; see "Bot challenges" in the operation description. |
| `proxy_verify` | boolean; default `false` | no | Probe the custom `proxy` before spending the engine cost so a dead proxy fails fast. Requires `proxy` (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise). |
| `session_id` | string | no | Reuse a session's cookies and storage (created via /v1/sessions); the session's engine is used and pinning a different engine is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. Checked before any work, at 0 credits and not retryable - an unknown, malformed or other organization's id is 404 ERR::SESSION::NOT_FOUND, a released session 410 ERR::SESSION::RELEASED, an expired one 410 ERR::SESSION::EXPIRED. A successful request increments the session's `usage_count`, sets `last_used_at` and slides `expires_at`. Requests with a session are never cached. |
| `sticky_key` | string | no | Coming soon. Any string; requests sharing it reuse the same pool exit. Forces a pool exit even on deployments whose default is direct egress, so it can fail 503 ERR::PROXY::EXHAUSTED where no pool exists. |
| `wait` | integer; default `0` | no | Milliseconds to wait after load before capture. Browser-only: needs js_render, stealth, an engine pin with a browser, or mode=auto, else 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. |
| `wait_for` | string | no | CSS selector to wait for before capture. Browser-only. If it never appears the page is returned as it stood with an X-Warning RENDER_DEGRADED; if the render times out waiting it fails 504 ERR::UPSTREAM::TIMEOUT - raise wait_for_timeout or fix the selector. |
| `wait_for_timeout` | integer; default `0` | no | Milliseconds to wait for `wait_for`; 0 uses the engine default. Requires wait_for (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise). |
| `actions` | array of object (`actions`) | no | Ordered browser workflow (click, fill, scroll, ...), validated before anything is spent. Browser-only; disables the default image/font blocking and makes the request uncacheable. A failed step ends the request at 0 credits unless that step has on_error=skip. |
| `block_resources` | array of string: `none`, `document`, `documents`, `stylesheet`, `stylesheets`, `css`, `image`, `images`, `media`, `font`, `fonts`, `script`, `scripts`, `js`, `xhr`, `fetch`, `websocket`, `websockets`, `other` (`block_resources`) | no | Subresource classes the browser must not load (plural and short aliases accepted, case-insensitive). If omitted, renders block image+font except for screenshot, pdf, actions, or a network_capture of images/fonts; an explicit list replaces that default and `none` (alone) blocks nothing. Any value other than `none` is browser-only. If omitted, the project's stored default applies. |
| `headless` | boolean | no | false runs the browser on a real display (1920x1080 screen instead of 800x600); only engine=chromium supports it, anything else is 400 ERR::REQUEST::CAPABILITY_UNSUPPORTED at 0 credits. Either value requires a browser engine; omit to use the deployment default (headless). Headful renders skip the warm pool and add ~1s. If omitted, the project's stored default applies. |
| `screenshot` | boolean; default `false` | no | Capture a screenshot; forces the JSON envelope (image base64 under `screenshots`). Needs a rasterising engine (camoufox or chromium) - otherwise 400 ERR::REQUEST::CAPABILITY_UNSUPPORTED at 0 credits. mode=auto cannot satisfy it (its fetch rung has no rasteriser). If the image is not delivered the page is still returned but charged 0. |
| `screenshot_fullpage` | boolean; default `false` | no | Capture the full scrollable page. Requires screenshot=true; mutually exclusive with screenshot_selector. |
| `screenshot_selector` | string | no | CSS selector of the element to capture. Requires screenshot=true; mutually exclusive with screenshot_fullpage. |
| `screenshot_format` | string: `png`, `jpeg`, `jpg`, `webp` | no | Image format; omitted uses the engine default (png). Not validated by the API - an unrecognised value silently falls back to the engine default. |
| `screenshot_quality` | integer; default `0` | no | Lossy quality for jpeg/webp; 0 uses the engine default. A value above 0 requires screenshot=true. |
| `custom_headers` | object (`custom_headers`) | no | Headers sent to the target, as name -> string value. Hop-by-hop headers (Connection, Keep-Alive, Proxy-Authenticate, Proxy-Authorization, TE, Trailer, Transfer-Encoding, Upgrade, Host, Content-Length) are refused with 400 ERR::REQUEST::INVALID_PARAMETER. If omitted, the project's stored default applies. |
| `autoparse` | boolean; default `false` | no | Harvest the page's own structured metadata (JSON-LD, OpenGraph, Twitter, microdata, RDFa, meta, embedded SPA state) into `data`, no selectors needed. Forces the JSON envelope; no extra credits; refused on a PDF. |
| `links` | boolean; default `false` | no | Return the page's absolute, de-duplicated <a href> targets under `links`. Forces the JSON envelope; no extra credits; refused on a PDF. If omitted, the project's stored default applies. |
| `network_capture` | object (`network_capture`) | no | Record the XHR/fetch responses the page itself made (often cleaner JSON than the DOM), returned under `network` in the JSON envelope. Browser-only; no extra credits. |
| `extract` | object (`extract`) | no | Selector extraction into `data`: a selector map, or a JSON Schema whose properties carry `selector`. Forces the JSON envelope, validated before fetching, no extra credits; mutually exclusive with extract_preset. Rules that match nothing are listed in `empty_fields`. |
| `extract_preset` | string | no | Coming soon. Name of a server-side extraction preset. This deployment has no preset store, so any value fails with 400 ERR::EXTRACT::INVALID_RULES (at 0 credits, after the fetch); use `extract` inline. Mutually exclusive with extract. |
| `ai_extract` | object (`ai_extract`) | no | Coming soon. Model-driven extraction into `data`; adds 4 credits. Refused with 503 ERR::INTERNAL::UNAVAILABLE where no model is configured, and a model failure is 502 ERR::EXTRACT::FAILED at 0 credits. Forces the JSON envelope. |
| `main_content_only` | boolean | no | true strips nav/footer/aside to the main article; false keeps the whole document. Omitted keeps the pipeline default (main-content isolation on for markdown). Applies to markdown/text/cleaned output. If omitted, the project's stored default applies. |
| `include_tags` | array of string (`include_tags`) | no | CSS selectors; keep only matching subtrees before extraction and markdown conversion. Applied before exclude_tags and main_content_only. |
| `exclude_tags` | array of string (`exclude_tags`) | no | CSS selectors; remove matching elements, applied after include_tags and before main_content_only. |
| `parse_pdf` | boolean; default `true` | no | When the target returns a PDF and markdown/text is requested, parse it to text. false refuses it with 502 ERR::EXTRACT::FAILED (0 credits); response_format=html always returns the raw bytes. If omitted, the project's stored default applies. |
| `response_format` | string: `html`, `markdown`, `text`, `json`, `pdf`; default `"html"` | no | Body shape: html/markdown/text return the document itself, json the envelope, pdf a browser-printed PDF (needs camoufox or chromium). extract, autoparse, links, ai_extract, network_capture or screenshot override this to the JSON envelope with X-Warning FORMAT_COERCED. A non-HTML target (image, binary) requested as markdown/text fails 502 ERR::EXTRACT::FAILED at 0 credits. If omitted, the project's stored default applies. |
| `cache` | boolean; default `true` | no | Serve a stored result younger than cache_ttl (Cache-State: hit, 0 credits). Only billable successes of GET-method requests without session_id or actions are stored. Set false for time-sensitive data; cache=false with cache_ttl>0 is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. |
| `cache_ttl` | integer; default `172800` | no | Maximum acceptable age of a cached result, in seconds. Default and cap are deployment settings (48h by default); larger values are clamped with X-Warning CACHE_TTL_CLAMPED. 0 disables caching for this request. |
| `max_cost` | integer; default `0` | no | Credit ceiling; 0 means none. Checked against the dearest reachable outcome (the top rung under mode=auto) before anything runs, failing 400 ERR::LIMIT::MAX_COST_EXCEEDED at 0 credits. |
| `original_status` | boolean; default `false` | no | On success, use the target's status as this response's HTTP status instead of 200. Leave off unless you need it: it makes a target 404 indistinguishable from an API error by status alone. |
| `allowed_status_codes` | array of integer (`allowed_status_codes`) | no | Extra target statuses to treat as billable successes (200, 404 and 410 always are). Other statuses return the target's body at 0 credits, and under mode=auto trigger escalation. |

```json
{
  "items": [
    {
      "url": "https://example.com/discovered/17",
      "external_id": "page-17"
    },
    {
      "url": "https://example.com/discovered/18"
    }
  ]
}
```

### Responses

#### 202

Items written. If `items_dispatched` < `items_added` the remainder is recovered by the worker's sweep; do NOT re-append them (they would run twice).

Headers: `X-Request-Cost`, `X-Credits-Charged`, `X-Credits-Remaining`, `X-Request-Id`.

`application/json`.

| Field | Type | Required | Description |
|---|---|---|---|
| `job` | object (`BatchJob`) | yes | The batch job, identical in shape across create, list, get, cancel and close. Poll `status_url` until `status` is terminal (`completed`, `failed`, `cancelled`), then read `results_url` before `results_expire_at`. There is no `credits_held` field: the hold is `estimated_credits`, the spend is `progress.credits_charged`. |
| `items_added` | integer | yes | Items written to the job. Their `seq` values continue from the previous `total_items`. |
| `items_dispatched` | integer | yes | Of those, how many reached the dispatch queue now. The rest are recovered by the worker's sweep; do not re-append them. |

#### 400

The request is invalid. Not retryable; fix the request using `code` and `diagnostics.hint`.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 401

Missing, invalid, revoked or expired key. Not retryable with the same key.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 402

`ERR::LIMIT::QUOTA_EXCEEDED`: the organization is out of credits. Not retryable until credits are added.

Headers: `X-Request-Id`, `X-Credits-Remaining`.

`application/problem+json` (`Problem` schema).

#### 403

The key lacks the scope this route requires (`ERR::AUTH::INSUFFICIENT_SCOPE`), or the action is not permitted.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 404

No such resource for this key's organization.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 409

Conflicts with the resource's current state, for example a session in use (`ERR::SESSION::BUSY`, retryable).

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 413

The request body exceeds the size limit.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 500

Internal error. Retryable when `retryable` is true.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 503

No proxy exit or engine capacity was available. Retryable after `Retry-After`; 0 credits.

Headers: `Retry-After`, `X-Request-Id`.

`application/problem+json` (`Problem` schema).

Full OpenAPI spec: https://docs.spicrawl.com/openapi.yaml
