Scrape a URL
Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data.
Requires scope: scrape
Credits (credits per successful request)
| engine | direct | datacenter | residential | mobile |
|---|---|---|---|---|
| fetch | 1 | 1 | 10 | 10 |
| obscura | 3 | 3 | 25 | 25 |
| camoufox | 25 | 25 | 25 | 25 |
| chromium | 8 | 8 | 32 | 32 |
ai_extract Coming soon adds 4. Failures cost 0; cache hits cost 0. The billed amount is in the X-Credits-Charged response header, and the monthly allowance left after it in X-Credits-Remaining.
Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data.
Choosing an engine (cheapest first):
- Default (fetch, 1 credit): plain HTTP, no JavaScript, presenting a current Chrome TLS/HTTP2 fingerprint (
impersonate, on by default and free;impersonate=falseturns it off). js_render=true(obscura, 3): content is built by JavaScript, or you need wait/wait_for/actions/block_resources/network_capture.engine=chromium(8): needed for screenshot and response_format=pdf.stealth=true(camoufox) is coming soon.mode=auto: unsure - escalates through the available engines, bills only the rung that worked.- Managed proxy pools (
premium_proxy) are coming soon; during the beta pass your own proxy withproxy. Browser-only flags on the fetch tier are a 400, never silently ignored.
Billing: failures cost 0 (errors, bot challenges, timeouts, target statuses other than 200/404/410 unless listed in allowed_status_codes, undelivered screenshots); cache hits cost 0. Check X-Credits-Charged.
Bot challenges: a challenge page attributed to a named vendor (Cloudflare, DataDome, PerimeterX, Akamai) is never returned as a success. A fetch through a managed pool exit that meets one is retried once on a different exit, free (GET/HEAD/OPTIONS only, not with sticky_key or session_id, never on your own proxy). With premium_proxy=true or your own proxy, no engine pin and no mode=auto, a request whose page is still a challenge then climbs to camoufox, which waits out or clicks through the interstitial. Only the rung that returned the page is billed (camoufox 25, or the first rung's price), and 0 if every rung was challenged (502 ERR::UPSTREAM::CHALLENGE). A transport error, a refusal with no vendor marker (such as a bare 403) or a non-billable status such as 404 does not climb. The climb is left off for a non-GET method, headless=false, response_format=pdf, a screenshot, actions, a proxy that is not http://, a max_cost below the camoufox price, or a deployment without camoufox; an allowance that covers the first rung but not camoufox runs the request without it. X-Engine names the engine that served, each abandoned rung adds an X-Warning: ESCALATED, and diagnostics.attempts lists every attempt (in the JSON envelope on a successful climb). Batch items never climb.
Status: the HTTP status is the platform's. A blocked or erroring site is usually still HTTP 200 - read envelope status or X-Target-Status for the site's answer. Errors are application/problem+json; retry only when retryable is true.
Not idempotent: there is no idempotency key, and each call performs (and bills) a new scrape unless served from cache. GET and POST accept the same parameters and behave identically.
Authorization
bearerAuth Authorization: Bearer <key>. Read the key from the SPICRAWL_API_KEY environment
variable; never hard-code or log it. spicrawl_test_… keys can never spend live
credits. Scopes: scrape, batch, sessions (granted by default), browser
and read (granted deliberately). A missing scope is 403 ERR::AUTH::INSUFFICIENT_SCOPE naming the scope.
In: header
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
Body of POST /v1/scrape. Unknown fields (including nested ones) are rejected with 400 ERR::REQUEST::INVALID_PARAMETER; a wrong JSON type is 400 naming the field. Max body 1 MiB (413).
Response Body
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
application/problem+json
curl -X POST "https://example.com/v1/scrape" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com/blog/launch", "response_format": "markdown" }'{ "url": "https://example.com/product/1", "final_url": "https://example.com/product/1", "status": 200, "content": "<html><head><title>Walnut Desk</title></head><body>...</body></html>", "headers": { "content-type": "text/html; charset=utf-8" }, "truncated": false, "credits": 1, "engine": "fetch", "proxy_source": "pool", "warnings": [], "data": { "title": "Walnut Desk", "price": "$349.00", "images": [ "https://example.com/img/desk-1.jpg" ] }}