spicrawlspicrawlDocs
Scrape

Scrape a URL (query-string form)

Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data.

Requires scope: scrape

Credits (credits per successful request)
enginedirectdatacenterresidentialmobile
fetch111010
obscura332525
camoufox25252525
chromium883232

ai_extract Coming soon adds 4. Failures cost 0; cache hits cost 0. The billed amount is in the X-Credits-Charged response header, and the monthly allowance left after it in X-Credits-Remaining.

GET
/v1/scrape

Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data.

Choosing an engine (cheapest first):

  • Default (fetch, 1 credit): plain HTTP, no JavaScript, presenting a current Chrome TLS/HTTP2 fingerprint (impersonate, on by default and free; impersonate=false turns it off).
  • js_render=true (obscura, 3): content is built by JavaScript, or you need wait/wait_for/actions/block_resources/network_capture.
  • engine=chromium (8): needed for screenshot and response_format=pdf. stealth=true (camoufox) is coming soon.
  • mode=auto: unsure - escalates through the available engines, bills only the rung that worked.
  • Managed proxy pools (premium_proxy) are coming soon; during the beta pass your own proxy with proxy. Browser-only flags on the fetch tier are a 400, never silently ignored.

Billing: failures cost 0 (errors, bot challenges, timeouts, target statuses other than 200/404/410 unless listed in allowed_status_codes, undelivered screenshots); cache hits cost 0. Check X-Credits-Charged.

Bot challenges: a challenge page attributed to a named vendor (Cloudflare, DataDome, PerimeterX, Akamai) is never returned as a success. A fetch through a managed pool exit that meets one is retried once on a different exit, free (GET/HEAD/OPTIONS only, not with sticky_key or session_id, never on your own proxy). With premium_proxy=true or your own proxy, no engine pin and no mode=auto, a request whose page is still a challenge then climbs to camoufox, which waits out or clicks through the interstitial. Only the rung that returned the page is billed (camoufox 25, or the first rung's price), and 0 if every rung was challenged (502 ERR::UPSTREAM::CHALLENGE). A transport error, a refusal with no vendor marker (such as a bare 403) or a non-billable status such as 404 does not climb. The climb is left off for a non-GET method, headless=false, response_format=pdf, a screenshot, actions, a proxy that is not http://, a max_cost below the camoufox price, or a deployment without camoufox; an allowance that covers the first rung but not camoufox runs the request without it. X-Engine names the engine that served, each abandoned rung adds an X-Warning: ESCALATED, and diagnostics.attempts lists every attempt (in the JSON envelope on a successful climb). Batch items never climb.

Status: the HTTP status is the platform's. A blocked or erroring site is usually still HTTP 200 - read envelope status or X-Target-Status for the site's answer. Errors are application/problem+json; retry only when retryable is true.

Not idempotent: there is no idempotency key, and each call performs (and bills) a new scrape unless served from cache. GET and POST accept the same parameters and behave identically.

Authorization

bearerAuth
AuthorizationBearer <token>

Authorization: Bearer <key>. Read the key from the SPICRAWL_API_KEY environment variable; never hard-code or log it. spicrawl_test_… keys can never spend live credits. Scopes: scrape, batch, sessions (granted by default), browser and read (granted deliberately). A missing scope is 403 ERR::AUTH::INSUFFICIENT_SCOPE naming the scope.

In: header

Query Parameters

url*string

Target URL; required, must be http or https with a host. A private or reserved address fails with 400 ERR::SECURITY::SSRF_BLOCKED at 0 credits.

Formaturi
Lengthlength <= 8192
method?string

HTTP method sent to the target (case-insensitive). There is no request-body parameter, so non-GET methods are sent without a body; only GET/HEAD/OPTIONS are retried on transient failure and only GET results are cached.

Default"GET"

Value in

  • "GET"
  • "POST"
  • "PUT"
  • "PATCH"
  • "DELETE"
  • "HEAD"
  • "OPTIONS"
js_render?boolean

Render in a browser engine (obscura; 3 credits datacenter, 25 residential) so JavaScript runs. Required by every browser-only flag unless stealth or mode=auto is set; cannot be combined with mode=auto. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?js_render means true.

Defaultfalse
stealth?boolean

Coming soon Render in the hardened camoufox browser, flat 25 credits at any proxy tier; wins over js_render when both are set. Cannot be combined with mode=auto or with an engine pin other than camoufox; returns 503 ERR::ENGINE::UNAVAILABLE where camoufox is not deployed. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?stealth means true.

Defaultfalse
impersonate?boolean

Fetch tier only. Present a current Chrome TLS/HTTP2 fingerprint, with matching browser headers and Accept-Encoding, to pass passive fingerprint checks. On by default, free, no JavaScript; impersonate=false turns it off and the fetch uses a non-browser TLS handshake. Ignored on render engines. If omitted, the project's stored default applies, and true when the project sets none. Accepts true/false/1/0; a bare ?impersonate means true.

Defaulttrue
mode?"auto"

auto escalates through the available engines (fetch -> obscura) until a rung returns a usable page, billing only the rung that succeeded (0 if all fail) while reserving the dearest rung against quota. Incompatible with js_render, stealth and engine. If omitted, the project's stored default applies.

Value in

  • "auto"
engine?string

Engine camoufox is coming soon. Pin the execution engine (case-insensitive); disables escalation and is never substituted. chromium (8 credits datacenter, 32 residential) is the only engine supporting headless=false; screenshots and pdf need camoufox or chromium. An unentitled engine is 403 ERR::AUTH::ENGINE_NOT_ENTITLED, an undeployed one 503 ERR::ENGINE::UNAVAILABLE, both at 0 credits. If omitted, the project's stored default applies.

Value in

  • "fetch"
  • "obscura"
  • "camoufox"
  • "chromium"
premium_proxy?boolean

Coming soon Use residential exits from the managed pool: fetch 10, obscura 25, chromium 32 credits (camoufox stays 25). If no residential exit is available the request fails 503 ERR::PROXY::EXHAUSTED rather than falling back; ignored (with X-Warning PROXY_FLAG_IGNORED) when proxy is set. A bot challenge from a named vendor is retried on a different exit (fetch tier), then climbs to camoufox; see "Bot challenges" in the operation description. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?premium_proxy means true.

Defaultfalse
proxy_country?string

Coming soon ISO 3166-1 alpha-2 exit country (case-insensitive), within the residential pool. Requires premium_proxy=true (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise) unless proxy is set, in which case it is ignored with a warning; an unavailable country is 503 ERR::PROXY::EXHAUSTED, never substituted. A browser render takes its timezone and locale from this country or, when it is omitted, from the country of the pool exit it was assigned. If omitted, the project's stored default applies.

Match^[A-Za-z]{2}$
proxy?string

Your own proxy URL (http, https, socks5 or socks5h; credentials in userinfo). Takes precedence over premium_proxy/proxy_country and adds 0 proxy credits; loopback hosts are refused with 400 ERR::SECURITY::SSRF_BLOCKED. An unreachable proxy fails 502 ERR::PROXY::UNREACHABLE. A bot challenge from a named vendor climbs to camoufox when the proxy is http://; see "Bot challenges" in the operation description.

proxy_verify?boolean

Probe the custom proxy before spending the engine cost so a dead proxy fails fast. Requires proxy (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise). Accepts true/false/1/0; a bare ?proxy_verify means true.

Defaultfalse
session_id?string

Reuse a session's cookies and storage (created via /v1/sessions); the session's engine is used and pinning a different engine is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. Checked before any work, at 0 credits and not retryable - an unknown, malformed or other organization's id is 404 ERR::SESSION::NOT_FOUND, a released session 410 ERR::SESSION::RELEASED, an expired one 410 ERR::SESSION::EXPIRED. A successful request increments the session's usage_count, sets last_used_at and slides expires_at. Requests with a session are never cached.

sticky_key?string

Coming soon Any string; requests sharing it reuse the same pool exit. Forces a pool exit even on deployments whose default is direct egress, so it can fail 503 ERR::PROXY::EXHAUSTED where no pool exists.

wait?integer

Milliseconds to wait after load before capture. Browser-only: needs js_render, stealth, an engine pin with a browser, or mode=auto, else 400 ERR::REQUEST::INCOMPATIBLE_FLAGS.

Range0 <= value <= 30000
Default0
wait_for?string

CSS selector to wait for before capture. Browser-only. If it never appears the page is returned as it stood with an X-Warning RENDER_DEGRADED; if the render times out waiting it fails 504 ERR::UPSTREAM::TIMEOUT - raise wait_for_timeout or fix the selector.

wait_for_timeout?integer

Milliseconds to wait for wait_for; 0 uses the engine default. Requires wait_for (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise).

Range0 <= value <= 120000
Default0
actions?string

URL-encoded JSON array. Ordered browser workflow (click, fill, scroll, ...), validated before anything is spent. Browser-only; disables the default image/font blocking and makes the request uncacheable. A failed step ends the request at 0 credits unless that step has on_error=skip.

block_resources?array<>

Comma-separated list (blank items dropped). Subresource classes the browser must not load (plural and short aliases accepted, case-insensitive). If omitted, renders block image+font except for screenshot, pdf, actions, or a network_capture of images/fonts; an explicit list replaces that default and none (alone) blocks nothing. Any value other than none is browser-only. If omitted, the project's stored default applies.

headless?boolean

false runs the browser on a real display (1920x1080 screen instead of 800x600); only engine=chromium supports it, anything else is 400 ERR::REQUEST::CAPABILITY_UNSUPPORTED at 0 credits. Either value requires a browser engine; omit to use the deployment default (headless). Headful renders skip the warm pool and add ~1s. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?headless means true.

screenshot?boolean

Capture a screenshot; forces the JSON envelope (image base64 under screenshots). Needs a rasterising engine (camoufox or chromium) - otherwise 400 ERR::REQUEST::CAPABILITY_UNSUPPORTED at 0 credits. mode=auto cannot satisfy it (its fetch rung has no rasteriser). If the image is not delivered the page is still returned but charged 0. Accepts true/false/1/0; a bare ?screenshot means true.

Defaultfalse
screenshot_fullpage?boolean

Capture the full scrollable page. Requires screenshot=true; mutually exclusive with screenshot_selector. Accepts true/false/1/0; a bare ?screenshot_fullpage means true.

Defaultfalse
screenshot_selector?string

CSS selector of the element to capture. Requires screenshot=true; mutually exclusive with screenshot_fullpage.

screenshot_format?string

Image format; omitted uses the engine default (png). Not validated by the API - an unrecognised value silently falls back to the engine default.

Value in

  • "png"
  • "jpeg"
  • "jpg"
  • "webp"
screenshot_quality?integer

Lossy quality for jpeg/webp; 0 uses the engine default. A value above 0 requires screenshot=true.

Range0 <= value <= 100
Default0
custom_headers?string

URL-encoded JSON object. Headers sent to the target, as name -> string value. Hop-by-hop headers (Connection, Keep-Alive, Proxy-Authenticate, Proxy-Authorization, TE, Trailer, Transfer-Encoding, Upgrade, Host, Content-Length) are refused with 400 ERR::REQUEST::INVALID_PARAMETER. If omitted, the project's stored default applies.

network_capture?string

URL-encoded JSON object. Record the XHR/fetch responses the page itself made (often cleaner JSON than the DOM), returned under network in the JSON envelope. Browser-only; no extra credits.

autoparse?boolean

Harvest the page's own structured metadata (JSON-LD, OpenGraph, Twitter, microdata, RDFa, meta, embedded SPA state) into data, no selectors needed. Forces the JSON envelope; no extra credits; refused on a PDF. Accepts true/false/1/0; a bare ?autoparse means true.

Defaultfalse
extract?string

URL-encoded JSON object. Selector extraction into data: a selector map, or a JSON Schema whose properties carry selector. Forces the JSON envelope, validated before fetching, no extra credits; mutually exclusive with extract_preset. Rules that match nothing are listed in empty_fields.

extract_preset?string

Coming soon Name of a server-side extraction preset. This deployment has no preset store, so any value fails with 400 ERR::EXTRACT::INVALID_RULES (at 0 credits, after the fetch); use extract inline. Mutually exclusive with extract.

ai_extract?string

Coming soon URL-encoded JSON object. Model-driven extraction into data; adds 4 credits. Refused with 503 ERR::INTERNAL::UNAVAILABLE where no model is configured, and a model failure is 502 ERR::EXTRACT::FAILED at 0 credits. Forces the JSON envelope.

main_content_only?boolean

true strips nav/footer/aside to the main article; false keeps the whole document. Omitted keeps the pipeline default (main-content isolation on for markdown). Applies to markdown/text/cleaned output. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?main_content_only means true.

include_tags?array<string>

Comma-separated list (blank items dropped). CSS selectors; keep only matching subtrees before extraction and markdown conversion. Applied before exclude_tags and main_content_only.

exclude_tags?array<string>

Comma-separated list (blank items dropped). CSS selectors; remove matching elements, applied after include_tags and before main_content_only.

parse_pdf?boolean

When the target returns a PDF and markdown/text is requested, parse it to text. false refuses it with 502 ERR::EXTRACT::FAILED (0 credits); response_format=html always returns the raw bytes. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?parse_pdf means true.

Defaulttrue
response_format?string

Body shape: html/markdown/text return the document itself, json the envelope, pdf a browser-printed PDF (needs camoufox or chromium). extract, autoparse, links, ai_extract, network_capture or screenshot override this to the JSON envelope with X-Warning FORMAT_COERCED. A non-HTML target (image, binary) requested as markdown/text fails 502 ERR::EXTRACT::FAILED at 0 credits. If omitted, the project's stored default applies.

Default"html"

Value in

  • "html"
  • "markdown"
  • "text"
  • "json"
  • "pdf"
cache?boolean

Serve a stored result younger than cache_ttl (Cache-State: hit, 0 credits). Only billable successes of GET-method requests without session_id or actions are stored. Set false for time-sensitive data; cache=false with cache_ttl>0 is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. Accepts true/false/1/0; a bare ?cache means true.

Defaulttrue
cache_ttl?integer

Maximum acceptable age of a cached result, in seconds. Default and cap are deployment settings (48h by default); larger values are clamped with X-Warning CACHE_TTL_CLAMPED. 0 disables caching for this request.

Range0 <= value
Default172800
max_cost?integer

Credit ceiling; 0 means none. Checked against the dearest reachable outcome (the top rung under mode=auto) before anything runs, failing 400 ERR::LIMIT::MAX_COST_EXCEEDED at 0 credits.

Range0 <= value
Default0
original_status?boolean

On success, use the target's status as this response's HTTP status instead of 200. Leave off unless you need it: it makes a target 404 indistinguishable from an API error by status alone. Accepts true/false/1/0; a bare ?original_status means true.

Defaultfalse
allowed_status_codes?array<>

Comma-separated list (blank items dropped). Extra target statuses to treat as billable successes (200, 404 and 410 always are). Other statuses return the target's body at 0 credits, and under mode=auto trigger escalation.

apikey?string
Deprecated

Accepted by the query parser but NOT honoured for authentication on this route (only /v1/browser reads a query key); a request relying on it gets 401 ERR::AUTH::MISSING_KEY. Send Authorization: Bearer <key> instead.

api_key?string
Deprecated

Accepted by the query parser but NOT honoured for authentication on this route (only /v1/browser reads a query key); a request relying on it gets 401 ERR::AUTH::MISSING_KEY. Send Authorization: Bearer <key> instead.

Response Body

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

curl -X GET "https://example.com/v1/scrape?url=http%3A%2F%2Fexample.com"

{  "url": "https://example.com/product/1",  "final_url": "https://example.com/product/1",  "status": 200,  "content": "<html><head><title>Walnut Desk</title></head><body>...</body></html>",  "headers": {    "content-type": "text/html; charset=utf-8"  },  "truncated": false,  "credits": 1,  "engine": "fetch",  "proxy_source": "pool",  "warnings": [],  "data": {    "title": "Walnut Desk",    "price": "$349.00",    "images": [      "https://example.com/img/desk-1.jpg"    ]  }}