spicrawlspicrawlDocs
Scrape

Scrape a URL

Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data.

Requires scope: scrape

Credits (credits per successful request)
enginedirectdatacenterresidentialmobile
fetch111010
obscura332525
camoufox25252525
chromium883232

ai_extract Coming soon adds 4. Failures cost 0; cache hits cost 0. The billed amount is in the X-Credits-Charged response header, and the monthly allowance left after it in X-Credits-Remaining.

POST
/v1/scrape

Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data.

Choosing an engine (cheapest first):

  • Default (fetch, 1 credit): plain HTTP, no JavaScript, presenting a current Chrome TLS/HTTP2 fingerprint (impersonate, on by default and free; impersonate=false turns it off).
  • js_render=true (obscura, 3): content is built by JavaScript, or you need wait/wait_for/actions/block_resources/network_capture.
  • engine=chromium (8): needed for screenshot and response_format=pdf. stealth=true (camoufox) is coming soon.
  • mode=auto: unsure - escalates through the available engines, bills only the rung that worked.
  • Managed proxy pools (premium_proxy) are coming soon; during the beta pass your own proxy with proxy. Browser-only flags on the fetch tier are a 400, never silently ignored.

Billing: failures cost 0 (errors, bot challenges, timeouts, target statuses other than 200/404/410 unless listed in allowed_status_codes, undelivered screenshots); cache hits cost 0. Check X-Credits-Charged.

Bot challenges: a challenge page attributed to a named vendor (Cloudflare, DataDome, PerimeterX, Akamai) is never returned as a success. A fetch through a managed pool exit that meets one is retried once on a different exit, free (GET/HEAD/OPTIONS only, not with sticky_key or session_id, never on your own proxy). With premium_proxy=true or your own proxy, no engine pin and no mode=auto, a request whose page is still a challenge then climbs to camoufox, which waits out or clicks through the interstitial. Only the rung that returned the page is billed (camoufox 25, or the first rung's price), and 0 if every rung was challenged (502 ERR::UPSTREAM::CHALLENGE). A transport error, a refusal with no vendor marker (such as a bare 403) or a non-billable status such as 404 does not climb. The climb is left off for a non-GET method, headless=false, response_format=pdf, a screenshot, actions, a proxy that is not http://, a max_cost below the camoufox price, or a deployment without camoufox; an allowance that covers the first rung but not camoufox runs the request without it. X-Engine names the engine that served, each abandoned rung adds an X-Warning: ESCALATED, and diagnostics.attempts lists every attempt (in the JSON envelope on a successful climb). Batch items never climb.

Status: the HTTP status is the platform's. A blocked or erroring site is usually still HTTP 200 - read envelope status or X-Target-Status for the site's answer. Errors are application/problem+json; retry only when retryable is true.

Not idempotent: there is no idempotency key, and each call performs (and bills) a new scrape unless served from cache. GET and POST accept the same parameters and behave identically.

Authorization

bearerAuth
AuthorizationBearer <token>

Authorization: Bearer <key>. Read the key from the SPICRAWL_API_KEY environment variable; never hard-code or log it. spicrawl_test_… keys can never spend live credits. Scopes: scrape, batch, sessions (granted by default), browser and read (granted deliberately). A missing scope is 403 ERR::AUTH::INSUFFICIENT_SCOPE naming the scope.

In: header

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Body of POST /v1/scrape. Unknown fields (including nested ones) are rejected with 400 ERR::REQUEST::INVALID_PARAMETER; a wrong JSON type is 400 naming the field. Max body 1 MiB (413).

Response Body

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

application/problem+json

curl -X POST "https://example.com/v1/scrape" \  -H "Content-Type: application/json" \  -d '{    "url": "https://example.com/blog/launch",    "response_format": "markdown"  }'

{  "url": "https://example.com/product/1",  "final_url": "https://example.com/product/1",  "status": 200,  "content": "<html><head><title>Walnut Desk</title></head><body>...</body></html>",  "headers": {    "content-type": "text/html; charset=utf-8"  },  "truncated": false,  "credits": 1,  "engine": "fetch",  "proxy_source": "pool",  "warnings": [],  "data": {    "title": "Walnut Desk",    "price": "$349.00",    "images": [      "https://example.com/img/desk-1.jpg"    ]  }}