Scrape a URL (query-string form)
Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data.
Requires scope: scrape
Credits (credits per successful request)
| engine | direct | datacenter | residential | mobile |
|---|---|---|---|---|
| fetch | 1 | 1 | 10 | 10 |
| obscura | 3 | 3 | 25 | 25 |
| camoufox | 25 | 25 | 25 | 25 |
| chromium | 8 | 8 | 32 | 32 |
ai_extract Coming soon adds 4. Failures cost 0; cache hits cost 0. The billed amount is in the X-Credits-Charged response header, and the monthly allowance left after it in X-Credits-Remaining.
Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data.
Choosing an engine (cheapest first):
- Default (fetch, 1 credit): plain HTTP, no JavaScript, presenting a current Chrome TLS/HTTP2 fingerprint (
impersonate, on by default and free;impersonate=falseturns it off). js_render=true(obscura, 3): content is built by JavaScript, or you need wait/wait_for/actions/block_resources/network_capture.engine=chromium(8): needed for screenshot and response_format=pdf.stealth=true(camoufox) is coming soon.mode=auto: unsure - escalates through the available engines, bills only the rung that worked.- Managed proxy pools (
premium_proxy) are coming soon; during the beta pass your own proxy withproxy. Browser-only flags on the fetch tier are a 400, never silently ignored.
Billing: failures cost 0 (errors, bot challenges, timeouts, target statuses other than 200/404/410 unless listed in allowed_status_codes, undelivered screenshots); cache hits cost 0. Check X-Credits-Charged.
Bot challenges: a challenge page attributed to a named vendor (Cloudflare, DataDome, PerimeterX, Akamai) is never returned as a success. A fetch through a managed pool exit that meets one is retried once on a different exit, free (GET/HEAD/OPTIONS only, not with sticky_key or session_id, never on your own proxy). With premium_proxy=true or your own proxy, no engine pin and no mode=auto, a request whose page is still a challenge then climbs to camoufox, which waits out or clicks through the interstitial. Only the rung that returned the page is billed (camoufox 25, or the first rung's price), and 0 if every rung was challenged (502 ERR::UPSTREAM::CHALLENGE). A transport error, a refusal with no vendor marker (such as a bare 403) or a non-billable status such as 404 does not climb. The climb is left off for a non-GET method, headless=false, response_format=pdf, a screenshot, actions, a proxy that is not http://, a max_cost below the camoufox price, or a deployment without camoufox; an allowance that covers the first rung but not camoufox runs the request without it. X-Engine names the engine that served, each abandoned rung adds an X-Warning: ESCALATED, and diagnostics.attempts lists every attempt (in the JSON envelope on a successful climb). Batch items never climb.
Status: the HTTP status is the platform's. A blocked or erroring site is usually still HTTP 200 - read envelope status or X-Target-Status for the site's answer. Errors are application/problem+json; retry only when retryable is true.
Not idempotent: there is no idempotency key, and each call performs (and bills) a new scrape unless served from cache. GET and POST accept the same parameters and behave identically.
Authorization
bearerAuth Authorization: Bearer <key>. Read the key from the SPICRAWL_API_KEY environment
variable; never hard-code or log it. spicrawl_test_… keys can never spend live
credits. Scopes: scrape, batch, sessions (granted by default), browser
and read (granted deliberately). A missing scope is 403 ERR::AUTH::INSUFFICIENT_SCOPE naming the scope.
In: header
Query Parameters
Target URL; required, must be http or https with a host. A private or reserved address fails with 400 ERR::SECURITY::SSRF_BLOCKED at 0 credits.
urilength <= 8192HTTP method sent to the target (case-insensitive). There is no request-body parameter, so non-GET methods are sent without a body; only GET/HEAD/OPTIONS are retried on transient failure and only GET results are cached.
"GET"Value in
- "GET"
- "POST"
- "PUT"
- "PATCH"
- "DELETE"
- "HEAD"
- "OPTIONS"
Render in a browser engine (obscura; 3 credits datacenter, 25 residential) so JavaScript runs. Required by every browser-only flag unless stealth or mode=auto is set; cannot be combined with mode=auto. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?js_render means true.
falseComing soon Render in the hardened camoufox browser, flat 25 credits at any proxy tier; wins over js_render when both are set. Cannot be combined with mode=auto or with an engine pin other than camoufox; returns 503 ERR::ENGINE::UNAVAILABLE where camoufox is not deployed. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?stealth means true.
falseFetch tier only. Present a current Chrome TLS/HTTP2 fingerprint, with matching browser headers and Accept-Encoding, to pass passive fingerprint checks. On by default, free, no JavaScript; impersonate=false turns it off and the fetch uses a non-browser TLS handshake. Ignored on render engines. If omitted, the project's stored default applies, and true when the project sets none. Accepts true/false/1/0; a bare ?impersonate means true.
trueauto escalates through the available engines (fetch -> obscura) until a rung returns a usable page, billing only the rung that succeeded (0 if all fail) while reserving the dearest rung against quota. Incompatible with js_render, stealth and engine. If omitted, the project's stored default applies.
Value in
- "auto"
Engine camoufox is coming soon. Pin the execution engine (case-insensitive); disables escalation and is never substituted. chromium (8 credits datacenter, 32 residential) is the only engine supporting headless=false; screenshots and pdf need camoufox or chromium. An unentitled engine is 403 ERR::AUTH::ENGINE_NOT_ENTITLED, an undeployed one 503 ERR::ENGINE::UNAVAILABLE, both at 0 credits. If omitted, the project's stored default applies.
Value in
- "fetch"
- "obscura"
- "camoufox"
- "chromium"
Coming soon ISO 3166-1 alpha-2 exit country (case-insensitive), within the residential pool. Requires premium_proxy=true (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise) unless proxy is set, in which case it is ignored with a warning; an unavailable country is 503 ERR::PROXY::EXHAUSTED, never substituted. A browser render takes its timezone and locale from this country or, when it is omitted, from the country of the pool exit it was assigned. If omitted, the project's stored default applies.
^[A-Za-z]{2}$Your own proxy URL (http, https, socks5 or socks5h; credentials in userinfo). Takes precedence over premium_proxy/proxy_country and adds 0 proxy credits; loopback hosts are refused with 400 ERR::SECURITY::SSRF_BLOCKED. An unreachable proxy fails 502 ERR::PROXY::UNREACHABLE. A bot challenge from a named vendor climbs to camoufox when the proxy is http://; see "Bot challenges" in the operation description.
Probe the custom proxy before spending the engine cost so a dead proxy fails fast. Requires proxy (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise). Accepts true/false/1/0; a bare ?proxy_verify means true.
falseReuse a session's cookies and storage (created via /v1/sessions); the session's engine is used and pinning a different engine is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. Checked before any work, at 0 credits and not retryable - an unknown, malformed or other organization's id is 404 ERR::SESSION::NOT_FOUND, a released session 410 ERR::SESSION::RELEASED, an expired one 410 ERR::SESSION::EXPIRED. A successful request increments the session's usage_count, sets last_used_at and slides expires_at. Requests with a session are never cached.
Coming soon Any string; requests sharing it reuse the same pool exit. Forces a pool exit even on deployments whose default is direct egress, so it can fail 503 ERR::PROXY::EXHAUSTED where no pool exists.
Milliseconds to wait after load before capture. Browser-only: needs js_render, stealth, an engine pin with a browser, or mode=auto, else 400 ERR::REQUEST::INCOMPATIBLE_FLAGS.
0 <= value <= 300000CSS selector to wait for before capture. Browser-only. If it never appears the page is returned as it stood with an X-Warning RENDER_DEGRADED; if the render times out waiting it fails 504 ERR::UPSTREAM::TIMEOUT - raise wait_for_timeout or fix the selector.
Milliseconds to wait for wait_for; 0 uses the engine default. Requires wait_for (400 ERR::REQUEST::INCOMPATIBLE_FLAGS otherwise).
0 <= value <= 1200000URL-encoded JSON array. Ordered browser workflow (click, fill, scroll, ...), validated before anything is spent. Browser-only; disables the default image/font blocking and makes the request uncacheable. A failed step ends the request at 0 credits unless that step has on_error=skip.
Comma-separated list (blank items dropped). Subresource classes the browser must not load (plural and short aliases accepted, case-insensitive). If omitted, renders block image+font except for screenshot, pdf, actions, or a network_capture of images/fonts; an explicit list replaces that default and none (alone) blocks nothing. Any value other than none is browser-only. If omitted, the project's stored default applies.
false runs the browser on a real display (1920x1080 screen instead of 800x600); only engine=chromium supports it, anything else is 400 ERR::REQUEST::CAPABILITY_UNSUPPORTED at 0 credits. Either value requires a browser engine; omit to use the deployment default (headless). Headful renders skip the warm pool and add ~1s. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?headless means true.
Capture a screenshot; forces the JSON envelope (image base64 under screenshots). Needs a rasterising engine (camoufox or chromium) - otherwise 400 ERR::REQUEST::CAPABILITY_UNSUPPORTED at 0 credits. mode=auto cannot satisfy it (its fetch rung has no rasteriser). If the image is not delivered the page is still returned but charged 0. Accepts true/false/1/0; a bare ?screenshot means true.
falseCapture the full scrollable page. Requires screenshot=true; mutually exclusive with screenshot_selector. Accepts true/false/1/0; a bare ?screenshot_fullpage means true.
falseCSS selector of the element to capture. Requires screenshot=true; mutually exclusive with screenshot_fullpage.
Image format; omitted uses the engine default (png). Not validated by the API - an unrecognised value silently falls back to the engine default.
Value in
- "png"
- "jpeg"
- "jpg"
- "webp"
Lossy quality for jpeg/webp; 0 uses the engine default. A value above 0 requires screenshot=true.
0 <= value <= 1000URL-encoded JSON object. Headers sent to the target, as name -> string value. Hop-by-hop headers (Connection, Keep-Alive, Proxy-Authenticate, Proxy-Authorization, TE, Trailer, Transfer-Encoding, Upgrade, Host, Content-Length) are refused with 400 ERR::REQUEST::INVALID_PARAMETER. If omitted, the project's stored default applies.
URL-encoded JSON object. Record the XHR/fetch responses the page itself made (often cleaner JSON than the DOM), returned under network in the JSON envelope. Browser-only; no extra credits.
Harvest the page's own structured metadata (JSON-LD, OpenGraph, Twitter, microdata, RDFa, meta, embedded SPA state) into data, no selectors needed. Forces the JSON envelope; no extra credits; refused on a PDF. Accepts true/false/1/0; a bare ?autoparse means true.
falseReturn the page's absolute, de-duplicated targets under links. Forces the JSON envelope; no extra credits; refused on a PDF. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?links means true.
falseURL-encoded JSON object. Selector extraction into data: a selector map, or a JSON Schema whose properties carry selector. Forces the JSON envelope, validated before fetching, no extra credits; mutually exclusive with extract_preset. Rules that match nothing are listed in empty_fields.
Coming soon Name of a server-side extraction preset. This deployment has no preset store, so any value fails with 400 ERR::EXTRACT::INVALID_RULES (at 0 credits, after the fetch); use extract inline. Mutually exclusive with extract.
Coming soon URL-encoded JSON object. Model-driven extraction into data; adds 4 credits. Refused with 503 ERR::INTERNAL::UNAVAILABLE where no model is configured, and a model failure is 502 ERR::EXTRACT::FAILED at 0 credits. Forces the JSON envelope.
true strips nav/footer/aside to the main article; false keeps the whole document. Omitted keeps the pipeline default (main-content isolation on for markdown). Applies to markdown/text/cleaned output. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?main_content_only means true.
When the target returns a PDF and markdown/text is requested, parse it to text. false refuses it with 502 ERR::EXTRACT::FAILED (0 credits); response_format=html always returns the raw bytes. If omitted, the project's stored default applies. Accepts true/false/1/0; a bare ?parse_pdf means true.
trueBody shape: html/markdown/text return the document itself, json the envelope, pdf a browser-printed PDF (needs camoufox or chromium). extract, autoparse, links, ai_extract, network_capture or screenshot override this to the JSON envelope with X-Warning FORMAT_COERCED. A non-HTML target (image, binary) requested as markdown/text fails 502 ERR::EXTRACT::FAILED at 0 credits. If omitted, the project's stored default applies.
"html"Value in
- "html"
- "markdown"
- "text"
- "json"
- "pdf"
Serve a stored result younger than cache_ttl (Cache-State: hit, 0 credits). Only billable successes of GET-method requests without session_id or actions are stored. Set false for time-sensitive data; cache=false with cache_ttl>0 is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. Accepts true/false/1/0; a bare ?cache means true.
trueMaximum acceptable age of a cached result, in seconds. Default and cap are deployment settings (48h by default); larger values are clamped with X-Warning CACHE_TTL_CLAMPED. 0 disables caching for this request.
0 <= value172800Credit ceiling; 0 means none. Checked against the dearest reachable outcome (the top rung under mode=auto) before anything runs, failing 400 ERR::LIMIT::MAX_COST_EXCEEDED at 0 credits.
0 <= value0On success, use the target's status as this response's HTTP status instead of 200. Leave off unless you need it: it makes a target 404 indistinguishable from an API error by status alone. Accepts true/false/1/0; a bare ?original_status means true.
falseComma-separated list (blank items dropped). Extra target statuses to treat as billable successes (200, 404 and 410 always are). Other statuses return the target's body at 0 credits, and under mode=auto trigger escalation.
Accepted by the query parser but NOT honoured for authentication on this route (only /v1/browser reads a query key); a request relying on it gets 401 ERR::AUTH::MISSING_KEY. Send Authorization: Bearer <key> instead.
Accepted by the query parser but NOT honoured for authentication on this route (only /v1/browser reads a query key); a request relying on it gets 401 ERR::AUTH::MISSING_KEY. Send Authorization: Bearer <key> instead.
Response Body
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
curl -X GET "https://example.com/v1/scrape?url=http%3A%2F%2Fexample.com"{ "url": "https://example.com/product/1", "final_url": "https://example.com/product/1", "status": 200, "content": "<html><head><title>Walnut Desk</title></head><body>...</body></html>", "headers": { "content-type": "text/html; charset=utf-8" }, "truncated": false, "credits": 1, "engine": "fetch", "proxy_source": "pool", "warnings": [], "data": { "title": "Walnut Desk", "price": "$349.00", "images": [ "https://example.com/img/desk-1.jpg" ] }}