API reference
Base URL, authentication, content types, strict JSON, the GET and POST forms of /v1/scrape, pagination and retries for the Spicrawl REST API.
The Spicrawl API is a JSON-over-HTTPS REST API at https://api.spicrawl.com. Every route is under /v1, every request carries Authorization: Bearer <key>, and every error is application/problem+json with a stable code.
curl -sS "https://api.spicrawl.com/v1/scrape" \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/products/42", "response_format": "markdown"}'Base URL
https://api.spicrawl.com| Group | Routes | Scope |
|---|---|---|
| Scrape | GET, POST /v1/scrape | scrape |
| Batch | /v1/batch, /v1/batch/{batchID}/… | batch |
| Sessions | /v1/sessions, /v1/sessions/{sessionID}/… | sessions |
| Browser (CDP WebSocket) Coming soon | GET /v1/browser | browser |
| Request log | GET /v1/requests, GET /v1/requests/{id} | none beyond a valid key for the key's project; read for all_projects=true and another project's request |
| Usage | GET /v1/usage, /v1/usage/summary, /v1/usage/reconciliation | read |
| Fleet | GET /v1/workers | none beyond a valid key |
| Health | GET /healthz, GET /readyz | no auth |
Authentication
Send the key as a Bearer token. Read it from SPICRAWL_API_KEY.
Authorization: Bearer spicrawl_live_…- Keys are
spicrawl_live_…(spends credits) orspicrawl_test_…(can never spend live credits). - Every route ignores a query key (
?apikey=) and returns401 ERR::AUTH::MISSING_KEYwithout the header. - A key without the route's scope gets
403 ERR::AUTH::INSUFFICIENT_SCOPE.
See Authentication.
Content types
| Direction | Type |
|---|---|
| Request bodies | application/json. Any other Content-Type is 400 ERR::REQUEST::INVALID. |
| JSON responses | application/json |
/v1/scrape document bodies | text/html, text/markdown, text/plain or application/pdf, following response_format |
| Batch results | application/x-ndjson (one JSON object per line) |
| Errors | application/problem+json |
Strict JSON
Request bodies are decoded strictly. An unknown or misspelled field is refused with 400 ERR::REQUEST::INVALID_PARAMETER, and detail names it. The same rule applies to query parameters on GET /v1/scrape. Nothing is silently ignored.
{
"type": "https://docs.spicrawl.com/errors#REQUEST_INVALID_PARAMETER",
"title": "A parameter has an invalid value",
"status": 400,
"code": "ERR::REQUEST::INVALID_PARAMETER",
"detail": "Json: unknown field \"headers\".",
"retryable": false,
"doc_url": "https://docs.spicrawl.com/errors#REQUEST_INVALID_PARAMETER",
"target_status": null
}Headers sent to the target go in custom_headers, not headers. Browser-only flags such as wait_for on the plain fetch engine are also a 400, never dropped.
GET and POST forms of /v1/scrape
GET /v1/scrape and POST /v1/scrape accept the same parameters and behave the same. A parameter the JSON body would reject is a 400 in the query string too.
In the GET form, encode values like this:
| Parameter type | GET encoding | Example |
|---|---|---|
| String, integer, boolean | Plain value | js_render=true&wait=2000 |
List (block_resources, include_tags, exclude_tags, allowed_status_codes) | Comma-separated | block_resources=image,font |
Object or array (actions, custom_headers, extract, ai_extract, network_capture) | JSON, URL-encoded | extract=%7B%22title%22%3A%22h1%22%7D |
url | URL-encoded | url=https%3A%2F%2Fexample.com%2Fproducts%2F42 |
curl -sS -G "https://api.spicrawl.com/v1/scrape" \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
--data-urlencode "url=https://example.com/products/42" \
--data-urlencode 'extract={"title":"h1","price":".price"}' \
--data-urlencode "block_resources=image,font" \
--data-urlencode "js_render=true"Prefer POST for anything with objects or arrays. url in the GET form is limited to 8,192 characters.
Pagination
List endpoints use three cursor styles. In every case, pass the cursor back verbatim; never construct one.
| Endpoint | Page size | Cursor in | Send back as | Last page when |
|---|---|---|---|---|
GET /v1/batch | limit 1-200, default 50 | body next_cursor | cursor | next_cursor is absent |
GET /v1/sessions | limit 1-200, default 50 | body next_cursor | cursor | next_cursor is absent |
GET /v1/batch/{batchID}/results | limit 1-5000, default 500 | header X-Next-Cursor (also in Link) | cursor | X-Next-Cursor is absent |
GET /v1/requests | limit 1-200, default 50 | body page.next_before and page.next_before_id | before and before_id, both | page.has_more is false |
/v1/batchand/v1/sessionsrefuse an out-of-rangelimitwith400./v1/requestsclamps it silently.- A page can hold fewer items than
limitand still have more after it. Keep paging until the "last page" condition in the table holds. - A
/v1/requestscursor older than the retention window returns410 ERR::REQUEST::BEYOND_RETENTION. Restart from the first page. - With
GET /v1/requests?all_projects=true, sendall_projects=trueon every page: a cursor pages the list it came from. - Batch results are JSON Lines and readable while the job runs; on a running job, a later call can return more items after the last cursor.
Retries and idempotency
There is no idempotency key on any route. Every POST /v1/scrape performs and bills a new scrape unless it is served from the cache, and every POST /v1/batch creates a new job.
- Retry a request that returned an error with
retryable: true, afterRetry-After. Failures cost 0 credits, so that retry is free. - Do not retry a
retryable: falseerror unchanged. Applydiagnostics.hint. - If a connection drops before you see the response, a retry can bill twice. For a batch, check
GET /v1/batchbefore resubmitting. - If a
202fromPOST /v1/batchcarries a warning that some items were not dispatched, do not resubmit. The platform dispatches them.
Errors
Every error is application/problem+json:
{
"type": "https://docs.spicrawl.com/errors#UPSTREAM_TIMEOUT",
"title": "Target did not respond in time",
"status": 504,
"code": "ERR::UPSTREAM::TIMEOUT",
"detail": "The target did not answer within the time budget.",
"retryable": true,
"doc_url": "https://docs.spicrawl.com/errors#UPSTREAM_TIMEOUT",
"request_id": "01M0HF5WFWE7PRE8KZHDTNETWN",
"instance": "/v1/scrape",
"target_status": null,
"diagnostics": {
"hint": "The target did not answer in time. If the page is JavaScript-rendered, try `js_render=true`."
}
}Switch on code, retry on retryable, honour Retry-After, and read diagnostics.hint. status is the platform's HTTP status; target_status is the site's. The full catalogue with fixes is in Errors.
Every response, success or error, carries X-Request-Id. Pass it to GET /v1/requests/{id} for the full trace, and quote it to support.