# Scrape a URL

> Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data.

Source: https://docs.spicrawl.com/api-reference/scrape/scrape-post

## POST /v1/scrape

Operation ID: `scrapePost`. API key scope: `scrape`.

Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data.

Choosing an engine (cheapest first):
- Default (fetch, 1 credit): plain HTTP, no JavaScript, presenting a current Chrome TLS/HTTP2 fingerprint (`impersonate`, on by default and free; `impersonate=false` turns it off).
- `js_render=true` (obscura, 3): content is built by JavaScript, or you need wait/wait_for/actions/block_resources/network_capture.
- `engine=chromium` (8): needed for screenshot and response_format=pdf. `stealth=true` (camoufox) is coming soon.
- `mode=auto`: unsure - escalates through the available engines, bills only the rung that worked.
- Managed proxy pools (`premium_proxy`) are coming soon; during the beta pass your own proxy with `proxy`.
Browser-only flags on the fetch tier are a 400, never silently ignored.

Billing: failures cost 0 (errors, bot challenges, timeouts, target statuses other than 200/404/410 unless listed in allowed_status_codes, undelivered screenshots); cache hits cost 0. Check X-Credits-Charged.

Bot challenges: a challenge page attributed to a named vendor (Cloudflare, DataDome, PerimeterX, Akamai) is never returned as a success. A fetch through a managed pool exit that meets one is retried once on a different exit, free (GET/HEAD/OPTIONS only, not with `sticky_key` or `session_id`, never on your own `proxy`). With `premium_proxy=true` or your own `proxy`, no `engine` pin and no `mode=auto`, a request whose page is still a challenge then climbs to camoufox, which waits out or clicks through the interstitial. Only the rung that returned the page is billed (camoufox 25, or the first rung's price), and 0 if every rung was challenged (502 ERR::UPSTREAM::CHALLENGE). A transport error, a refusal with no vendor marker (such as a bare 403) or a non-billable status such as 404 does not climb. The climb is left off for a non-GET method, headless=false, response_format=pdf, a screenshot, actions, a `proxy` that is not http://, a max_cost below the camoufox price, or a deployment without camoufox; an allowance that covers the first rung but not camoufox runs the request without it. X-Engine names the engine that served, each abandoned rung adds an `X-Warning: ESCALATED`, and diagnostics.attempts lists every attempt (in the JSON envelope on a successful climb). Batch items never climb.

Status: the HTTP status is the platform's. A blocked or erroring site is usually still HTTP 200 - read envelope `status` or X-Target-Status for the site's answer. Errors are application/problem+json; retry only when `retryable` is true.

Not idempotent: there is no idempotency key, and each call performs (and bills) a new scrape unless served from cache. GET and POST accept the same parameters and behave identically.

### Example

```bash
curl -X POST "https://api.spicrawl.com/v1/scrape" \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com/blog/launch","response_format":"markdown"}'
```

### Request body

`application/json`, required.

| Field | Type | Required | Description |
|---|---|---|---|
| `url` | string (uri) | yes | Target URL; required, must be http or https with a host. A private or reserved address fails with 400 ERR::SECURITY::SSRF_BLOCKED at 0 credits. |
| `method` | string: `GET`, `POST`, `PUT`, `PATCH`, `DELETE`, `HEAD`, `OPTIONS`; default `"GET"` | no | HTTP method sent to the target (case-insensitive). There is no request-body parameter, so non-GET methods are sent without a body; only GET/HEAD/OPTIONS are retried on transient failure and only GET results are cached. |
| `js_render` | boolean; default `false` | no | Render in a browser engine (obscura; 3 credits datacenter, 25 residential) so JavaScript runs. Required by every browser-only flag unless stealth or mode=auto is set; cannot be combined with mode=auto. If omitted, the project's stored default applies. |
| `stealth` | boolean; default `false` | no | Coming soon. Render in the hardened camoufox browser, flat 25 credits at any proxy tier; wins over js_render when both are set. Cannot be combined with mode=auto or with an engine pin other than camoufox; returns 503 ERR::ENGINE::UNAVAILABLE where camoufox is not deployed. If omitted, the project's stored default applies. |
| `impersonate` | boolean; default `true` | no | Fetch tier only: present a current Chrome TLS/HTTP2 fingerprint, with matching browser headers and Accept-Encoding, to pass passive fingerprint checks. On by default, free, no JavaScript; `false` turns it off and the fetch uses a non-browser TLS handshake. Ignored on render engines. If omitted, the project's stored default applies, and `true` when the project sets none. |
| `mode` | string: `auto` | no | `auto` escalates through the available engines (fetch -> obscura) until a rung returns a usable page, billing only the rung that succeeded (0 if all fail) while reserving the dearest rung against quota. Incompatible with js_render, stealth and engine. If omitted, the project's stored default applies. |
| `engine` | string: `fetch`, `obscura`, `camoufox`, `chromium` | no | Engine `camoufox` is coming soon. Pin the execution engine (case-insensitive); disables escalation and is never substituted. chromium (8 credits datacenter, 32 residential) is the only engine supporting headless=false; screenshots and pdf need camoufox or chromium. An unentitled engine is 403 ERR::AUTH::ENGINE_NOT_ENTITLED, an undeployed one 503 ERR::ENGINE::UNAVAILABLE, both at 0 credits. If omitted, the project's stored default applies. |
| `premium_proxy` | boolean; default `false` | no | Coming soon. Use residential exits from the managed pool: fetch 10, obscura 25, chromium 32 credits (camoufox stays 25). If no residential exit is available the request fails 503 ERR::PROXY::EXHAUSTED rather than falling back; ignored (with X-Warning PROXY_FLAG_IGNORED) when `proxy` is set. A bot challenge from a named vendor is retried on a different exit (fetch tier), then climbs to camoufox; see "Bot challenges" in the operation description. If omitted, the project's stored default applies. |
| `proxy_country` | string | no | Coming soon. ISO 3166-1 alpha-2 exit country (case-insensitive), within the residential pool. Requires premium_proxy=true (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise) unless `proxy` is set, in which case it is ignored with a warning; an unavailable country is 503 ERR::PROXY::EXHAUSTED, never substituted. A browser render takes its timezone and locale from this country or, when it is omitted, from the country of the pool exit it was assigned. If omitted, the project's stored default applies. |
| `proxy` | string | no | Your own proxy URL (http, https, socks5 or socks5h; credentials in userinfo). Takes precedence over premium_proxy/proxy_country and adds 0 proxy credits; loopback hosts are refused with 400 ERR::SECURITY::SSRF_BLOCKED. An unreachable proxy fails 502 ERR::PROXY::UNREACHABLE. A bot challenge from a named vendor climbs to camoufox when the proxy is http://; see "Bot challenges" in the operation description. |
| `proxy_verify` | boolean; default `false` | no | Probe the custom `proxy` before spending the engine cost so a dead proxy fails fast. Requires `proxy` (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise). |
| `session_id` | string | no | Reuse a session's cookies and storage (created via /v1/sessions); the session's engine is used and pinning a different engine is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. Checked before any work, at 0 credits and not retryable - an unknown, malformed or other organization's id is 404 ERR::SESSION::NOT_FOUND, a released session 410 ERR::SESSION::RELEASED, an expired one 410 ERR::SESSION::EXPIRED. A successful request increments the session's `usage_count`, sets `last_used_at` and slides `expires_at`. Requests with a session are never cached. |
| `sticky_key` | string | no | Coming soon. Any string; requests sharing it reuse the same pool exit. Forces a pool exit even on deployments whose default is direct egress, so it can fail 503 ERR::PROXY::EXHAUSTED where no pool exists. |
| `wait` | integer; default `0` | no | Milliseconds to wait after load before capture. Browser-only: needs js_render, stealth, an engine pin with a browser, or mode=auto, else 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. |
| `wait_for` | string | no | CSS selector to wait for before capture. Browser-only. If it never appears the page is returned as it stood with an X-Warning RENDER_DEGRADED; if the render times out waiting it fails 504 ERR::UPSTREAM::TIMEOUT - raise wait_for_timeout or fix the selector. |
| `wait_for_timeout` | integer; default `0` | no | Milliseconds to wait for `wait_for`; 0 uses the engine default. Requires wait_for (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise). |
| `actions` | array of object (`ScrapeAction`) | no | Ordered browser workflow (click, fill, scroll, ...), validated before anything is spent. Browser-only; disables the default image/font blocking and makes the request uncacheable. A failed step ends the request at 0 credits unless that step has on_error=skip. |
| `block_resources` | array of string: `none`, `document`, `documents`, `stylesheet`, `stylesheets`, `css`, `image`, `images`, `media`, `font`, `fonts`, `script`, `scripts`, `js`, `xhr`, `fetch`, `websocket`, `websockets`, `other` | no | Subresource classes the browser must not load (plural and short aliases accepted, case-insensitive). If omitted, renders block image+font except for screenshot, pdf, actions, or a network_capture of images/fonts; an explicit list replaces that default and `none` (alone) blocks nothing. Any value other than `none` is browser-only. If omitted, the project's stored default applies. |
| `headless` | boolean | no | false runs the browser on a real display (1920x1080 screen instead of 800x600); only engine=chromium supports it, anything else is 400 ERR::REQUEST::CAPABILITY_UNSUPPORTED at 0 credits. Either value requires a browser engine; omit to use the deployment default (headless). Headful renders skip the warm pool and add ~1s. If omitted, the project's stored default applies. |
| `screenshot` | boolean; default `false` | no | Capture a screenshot; forces the JSON envelope (image base64 under `screenshots`). Needs a rasterising engine (camoufox or chromium) - otherwise 400 ERR::REQUEST::CAPABILITY_UNSUPPORTED at 0 credits. mode=auto cannot satisfy it (its fetch rung has no rasteriser). If the image is not delivered the page is still returned but charged 0. |
| `screenshot_fullpage` | boolean; default `false` | no | Capture the full scrollable page. Requires screenshot=true; mutually exclusive with screenshot_selector. |
| `screenshot_selector` | string | no | CSS selector of the element to capture. Requires screenshot=true; mutually exclusive with screenshot_fullpage. |
| `screenshot_format` | string: `png`, `jpeg`, `jpg`, `webp` | no | Image format; omitted uses the engine default (png). Not validated by the API - an unrecognised value silently falls back to the engine default. |
| `screenshot_quality` | integer; default `0` | no | Lossy quality for jpeg/webp; 0 uses the engine default. A value above 0 requires screenshot=true. |
| `custom_headers` | object | no | Headers sent to the target, as name -> string value. Hop-by-hop headers (Connection, Keep-Alive, Proxy-Authenticate, Proxy-Authorization, TE, Trailer, Transfer-Encoding, Upgrade, Host, Content-Length) are refused with 400 ERR::REQUEST::INVALID_PARAMETER. If omitted, the project's stored default applies. |
| `network_capture` |  | no | Record the XHR/fetch responses the page itself made (often cleaner JSON than the DOM), returned under `network` in the JSON envelope. Browser-only; no extra credits. |
| `autoparse` | boolean; default `false` | no | Harvest the page's own structured metadata (JSON-LD, OpenGraph, Twitter, microdata, RDFa, meta, embedded SPA state) into `data`, no selectors needed. Forces the JSON envelope; no extra credits; refused on a PDF. |
| `links` | boolean; default `false` | no | Return the page's absolute, de-duplicated <a href> targets under `links`. Forces the JSON envelope; no extra credits; refused on a PDF. If omitted, the project's stored default applies. |
| `extract` |  | no | Selector extraction into `data`: a selector map, or a JSON Schema whose properties carry `selector`. Forces the JSON envelope, validated before fetching, no extra credits; mutually exclusive with extract_preset. Rules that match nothing are listed in `empty_fields`. |
| `extract_preset` | string | no | Coming soon. Name of a server-side extraction preset. This deployment has no preset store, so any value fails with 400 ERR::EXTRACT::INVALID_RULES (at 0 credits, after the fetch); use `extract` inline. Mutually exclusive with extract. |
| `ai_extract` |  | no | Coming soon. Model-driven extraction into `data`; adds 4 credits. Refused with 503 ERR::INTERNAL::UNAVAILABLE where no model is configured, and a model failure is 502 ERR::EXTRACT::FAILED at 0 credits. Forces the JSON envelope. |
| `main_content_only` | boolean | no | true strips nav/footer/aside to the main article; false keeps the whole document. Omitted keeps the pipeline default (main-content isolation on for markdown). Applies to markdown/text/cleaned output. If omitted, the project's stored default applies. |
| `include_tags` | array of string | no | CSS selectors; keep only matching subtrees before extraction and markdown conversion. Applied before exclude_tags and main_content_only. |
| `exclude_tags` | array of string | no | CSS selectors; remove matching elements, applied after include_tags and before main_content_only. |
| `parse_pdf` | boolean; default `true` | no | When the target returns a PDF and markdown/text is requested, parse it to text. false refuses it with 502 ERR::EXTRACT::FAILED (0 credits); response_format=html always returns the raw bytes. If omitted, the project's stored default applies. |
| `response_format` | string: `html`, `markdown`, `text`, `json`, `pdf`; default `"html"` | no | Body shape: html/markdown/text return the document itself, json the envelope, pdf a browser-printed PDF (needs camoufox or chromium). extract, autoparse, links, ai_extract, network_capture or screenshot override this to the JSON envelope with X-Warning FORMAT_COERCED. A non-HTML target (image, binary) requested as markdown/text fails 502 ERR::EXTRACT::FAILED at 0 credits. If omitted, the project's stored default applies. |
| `cache` | boolean; default `true` | no | Serve a stored result younger than cache_ttl (Cache-State: hit, 0 credits). Only billable successes of GET-method requests without session_id or actions are stored. Set false for time-sensitive data; cache=false with cache_ttl>0 is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. |
| `cache_ttl` | integer; default `172800` | no | Maximum acceptable age of a cached result, in seconds. Default and cap are deployment settings (48h by default); larger values are clamped with X-Warning CACHE_TTL_CLAMPED. 0 disables caching for this request. |
| `max_cost` | integer; default `0` | no | Credit ceiling; 0 means none. Checked against the dearest reachable outcome (the top rung under mode=auto) before anything runs, failing 400 ERR::LIMIT::MAX_COST_EXCEEDED at 0 credits. |
| `original_status` | boolean; default `false` | no | On success, use the target's status as this response's HTTP status instead of 200. Leave off unless you need it: it makes a target 404 indistinguishable from an API error by status alone. |
| `allowed_status_codes` | array of integer | no | Extra target statuses to treat as billable successes (200, 404 and 410 always are). Other statuses return the target's body at 0 credits, and under mode=auto trigger escalation. |

Example `markdown`: Page as markdown

```json
{
  "url": "https://example.com/blog/launch",
  "response_format": "markdown"
}
```

Example `selector_extract`: Selector extraction

```json
{
  "url": "https://example.com/product/1",
  "extract": {
    "title": "h1",
    "price": {
      "selector": ".price",
      "output": "text"
    },
    "images": {
      "selector": ".gallery img",
      "kind": "list",
      "output": "attr",
      "attribute": "src"
    }
  }
}
```

Example `rendered_actions`: Rendered with actions and a screenshot

```json
{
  "url": "https://example.com/pricing",
  "engine": "chromium",
  "wait_for": ".plans",
  "actions": [
    {
      "click": {
        "selector": "#accept-cookies"
      },
      "on_error": "skip"
    },
    {
      "click": ".toggle-annual",
      "label": "annual"
    },
    {
      "scroll": {
        "to_bottom": true
      }
    }
  ],
  "screenshot": true,
  "response_format": "markdown",
  "max_cost": 10
}
```

### Responses

#### 200

Scrape completed. The HTTP status is the platform's: the site's status is envelope `status` / X-Target-Status, and a non-billable target status (e.g. 403, 500) still returns 200 with the site's body at 0 credits. With original_status=true this status line carries the target's status instead.

Headers: `X-Request-Id`, `X-Credits-Charged`, `X-Request-Cost`, `X-Credits-Remaining`, `X-Engine`, `X-Proxy-Source`, `X-Target-Status`, `X-Proxy-Endpoint`, `X-Final-Url`, `X-RateLimit-Limit`, `X-RateLimit-Remaining`, `X-RateLimit-Reset`, `Concurrency-Limit`, `Concurrency-Remaining`, `Cache-State`, `X-Warning`.

`application/json`.

| Field | Type | Required | Description |
|---|---|---|---|
| `url` | string | yes | The requested URL. |
| `final_url` | string | yes | URL after redirects; may be empty. |
| `status` | integer | yes | The TARGET's HTTP status (not this response's). Anything other than 200/404/410 or allowed_status_codes was charged 0. |
| `content` | string | yes | The page in the requested rendering: html (cleaned html if extraction produced one), markdown or text. |
| `headers` | object \| null | yes | The target's response headers. |
| `truncated` | boolean | yes | The target body exceeded the size cap and was cut. |
| `credits` | integer | yes | Credits charged for this request (same as X-Credits-Charged). |
| `engine` | string: `fetch`, `obscura`, `camoufox`, `chromium` | yes | Engine that produced the page. |
| `proxy_source` | string: `pool`, `custom`, `direct` | yes |  |
| `warnings` | array of string | yes | Human-readable warnings (messages only; codes are on X-Warning). |
| `data` | object | no | Output of extract, autoparse or ai_extract. |
| `empty_fields` | array of string | no | Extraction rules that matched nothing - the usual sign a selector broke. |
| `links` | array of string (uri) | no | Present (possibly empty) when links=true. |
| `screenshots` | array of object (`ScrapeScreenshot`) | no | Delivered images only; undelivered ones are reported in warnings and charged 0. |
| `network` | array of object (`ScrapeNetworkResponse`) | no | Captured responses when network_capture was set. |
| `diagnostics` | object (`ScrapeDiagnostics`) | no | Present on a successful mode=auto escalation: which rungs were tried and why they were abandoned. |

Example `extracted`: Selector extraction

```json
{
  "url": "https://example.com/product/1",
  "final_url": "https://example.com/product/1",
  "status": 200,
  "content": "<html><head><title>Walnut Desk</title></head><body>...</body></html>",
  "headers": {
    "content-type": "text/html; charset=utf-8"
  },
  "truncated": false,
  "credits": 1,
  "engine": "fetch",
  "proxy_source": "pool",
  "warnings": [],
  "data": {
    "title": "Walnut Desk",
    "price": "$349.00",
    "images": [
      "https://example.com/img/desk-1.jpg"
    ]
  }
}
```

Example `screenshot`: Rendered with a screenshot

```json
{
  "url": "https://example.com/pricing",
  "final_url": "https://example.com/pricing",
  "status": 200,
  "content": "# Pricing\n\n| Plan | Price |\n|---|---|\n| Starter | $29 |",
  "headers": {
    "content-type": "text/html"
  },
  "truncated": false,
  "credits": 8,
  "engine": "chromium",
  "proxy_source": "pool",
  "warnings": [
    "`screenshot` returns an image, which has no representation in a markdown body, so the response is the JSON envelope with the rendering under `content` and the image under `screenshots`."
  ],
  "screenshots": [
    {
      "label": "final",
      "encoding": "base64",
      "data": "iVBORw0KGgoAAAANSUhEUg...",
      "size_bytes": 184233,
      "format": "png",
      "width": 1280,
      "height": 3400
    }
  ]
}
```

`text/html`.

Type: string.

`text/markdown`.

Type: string.

`text/plain`.

Type: string.

`application/pdf`.

Type: string.

#### 400

The request is invalid. Not retryable; fix the request using `code` and `diagnostics.hint`.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 401

Missing, invalid, revoked or expired key. Not retryable with the same key.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 402

`ERR::LIMIT::QUOTA_EXCEEDED`: the organization is out of credits. Not retryable until credits are added.

Headers: `X-Request-Id`, `X-Credits-Remaining`.

`application/problem+json` (`Problem` schema).

#### 403

The key lacks the scope this route requires (`ERR::AUTH::INSUFFICIENT_SCOPE`), or the action is not permitted.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 404

No such resource for this key's organization.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 410

The resource existed but is past its retention window or was purged. Not retryable.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 413

The request body exceeds the size limit.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 429

Rate, concurrency or live-session limit reached. Retryable after `Retry-After`.

Headers: `Retry-After`, `X-RateLimit-Limit`, `X-RateLimit-Remaining`, `X-RateLimit-Reset`, `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 500

Internal error. Retryable when `retryable` is true.

Headers: `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 502

The target, the proxy or the render engine failed. Check `code` and `target_status`; most are retryable and cost 0 credits.

Headers: `Retry-After`, `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 503

No proxy exit or engine capacity was available. Retryable after `Retry-After`; 0 credits.

Headers: `Retry-After`, `X-Request-Id`.

`application/problem+json` (`Problem` schema).

#### 504

The target did not answer in time (`ERR::UPSTREAM::TIMEOUT`). Retryable; 0 credits. See `diagnostics.hint`.

Headers: `Retry-After`, `X-Request-Id`.

`application/problem+json` (`Problem` schema).

Full OpenAPI spec: https://docs.spicrawl.com/openapi.yaml
