spicrawlspicrawlDocs

API reference

Base URL, authentication, content types, strict JSON, the GET and POST forms of /v1/scrape, pagination and retries for the Spicrawl REST API.

The Spicrawl API is a JSON-over-HTTPS REST API at https://api.spicrawl.com. Every route is under /v1, every request carries Authorization: Bearer <key>, and every error is application/problem+json with a stable code.

curl -sS "https://api.spicrawl.com/v1/scrape" \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/products/42", "response_format": "markdown"}'

Base URL

https://api.spicrawl.com
GroupRoutesScope
ScrapeGET, POST /v1/scrapescrape
Batch/v1/batch, /v1/batch/{batchID}/…batch
Sessions/v1/sessions, /v1/sessions/{sessionID}/…sessions
Browser (CDP WebSocket) Coming soonGET /v1/browserbrowser
Request logGET /v1/requests, GET /v1/requests/{id}none beyond a valid key for the key's project; read for all_projects=true and another project's request
UsageGET /v1/usage, /v1/usage/summary, /v1/usage/reconciliationread
FleetGET /v1/workersnone beyond a valid key
HealthGET /healthz, GET /readyzno auth

Authentication

Send the key as a Bearer token. Read it from SPICRAWL_API_KEY.

Authorization: Bearer spicrawl_live_…
  • Keys are spicrawl_live_… (spends credits) or spicrawl_test_… (can never spend live credits).
  • Every route ignores a query key (?apikey=) and returns 401 ERR::AUTH::MISSING_KEY without the header.
  • A key without the route's scope gets 403 ERR::AUTH::INSUFFICIENT_SCOPE.

See Authentication.

Content types

DirectionType
Request bodiesapplication/json. Any other Content-Type is 400 ERR::REQUEST::INVALID.
JSON responsesapplication/json
/v1/scrape document bodiestext/html, text/markdown, text/plain or application/pdf, following response_format
Batch resultsapplication/x-ndjson (one JSON object per line)
Errorsapplication/problem+json

Strict JSON

Request bodies are decoded strictly. An unknown or misspelled field is refused with 400 ERR::REQUEST::INVALID_PARAMETER, and detail names it. The same rule applies to query parameters on GET /v1/scrape. Nothing is silently ignored.

400 ERR::REQUEST::INVALID_PARAMETER
{
  "type": "https://docs.spicrawl.com/errors#REQUEST_INVALID_PARAMETER",
  "title": "A parameter has an invalid value",
  "status": 400,
  "code": "ERR::REQUEST::INVALID_PARAMETER",
  "detail": "Json: unknown field \"headers\".",
  "retryable": false,
  "doc_url": "https://docs.spicrawl.com/errors#REQUEST_INVALID_PARAMETER",
  "target_status": null
}

Headers sent to the target go in custom_headers, not headers. Browser-only flags such as wait_for on the plain fetch engine are also a 400, never dropped.

GET and POST forms of /v1/scrape

GET /v1/scrape and POST /v1/scrape accept the same parameters and behave the same. A parameter the JSON body would reject is a 400 in the query string too.

In the GET form, encode values like this:

Parameter typeGET encodingExample
String, integer, booleanPlain valuejs_render=true&wait=2000
List (block_resources, include_tags, exclude_tags, allowed_status_codes)Comma-separatedblock_resources=image,font
Object or array (actions, custom_headers, extract, ai_extract, network_capture)JSON, URL-encodedextract=%7B%22title%22%3A%22h1%22%7D
urlURL-encodedurl=https%3A%2F%2Fexample.com%2Fproducts%2F42
curl -sS -G "https://api.spicrawl.com/v1/scrape" \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  --data-urlencode "url=https://example.com/products/42" \
  --data-urlencode 'extract={"title":"h1","price":".price"}' \
  --data-urlencode "block_resources=image,font" \
  --data-urlencode "js_render=true"

Prefer POST for anything with objects or arrays. url in the GET form is limited to 8,192 characters.

Pagination

List endpoints use three cursor styles. In every case, pass the cursor back verbatim; never construct one.

EndpointPage sizeCursor inSend back asLast page when
GET /v1/batchlimit 1-200, default 50body next_cursorcursornext_cursor is absent
GET /v1/sessionslimit 1-200, default 50body next_cursorcursornext_cursor is absent
GET /v1/batch/{batchID}/resultslimit 1-5000, default 500header X-Next-Cursor (also in Link)cursorX-Next-Cursor is absent
GET /v1/requestslimit 1-200, default 50body page.next_before and page.next_before_idbefore and before_id, bothpage.has_more is false
  • /v1/batch and /v1/sessions refuse an out-of-range limit with 400. /v1/requests clamps it silently.
  • A page can hold fewer items than limit and still have more after it. Keep paging until the "last page" condition in the table holds.
  • A /v1/requests cursor older than the retention window returns 410 ERR::REQUEST::BEYOND_RETENTION. Restart from the first page.
  • With GET /v1/requests?all_projects=true, send all_projects=true on every page: a cursor pages the list it came from.
  • Batch results are JSON Lines and readable while the job runs; on a running job, a later call can return more items after the last cursor.

Retries and idempotency

There is no idempotency key on any route. Every POST /v1/scrape performs and bills a new scrape unless it is served from the cache, and every POST /v1/batch creates a new job.

  • Retry a request that returned an error with retryable: true, after Retry-After. Failures cost 0 credits, so that retry is free.
  • Do not retry a retryable: false error unchanged. Apply diagnostics.hint.
  • If a connection drops before you see the response, a retry can bill twice. For a batch, check GET /v1/batch before resubmitting.
  • If a 202 from POST /v1/batch carries a warning that some items were not dispatched, do not resubmit. The platform dispatches them.

Errors

Every error is application/problem+json:

{
  "type": "https://docs.spicrawl.com/errors#UPSTREAM_TIMEOUT",
  "title": "Target did not respond in time",
  "status": 504,
  "code": "ERR::UPSTREAM::TIMEOUT",
  "detail": "The target did not answer within the time budget.",
  "retryable": true,
  "doc_url": "https://docs.spicrawl.com/errors#UPSTREAM_TIMEOUT",
  "request_id": "01M0HF5WFWE7PRE8KZHDTNETWN",
  "instance": "/v1/scrape",
  "target_status": null,
  "diagnostics": {
    "hint": "The target did not answer in time. If the page is JavaScript-rendered, try `js_render=true`."
  }
}

Switch on code, retry on retryable, honour Retry-After, and read diagnostics.hint. status is the platform's HTTP status; target_status is the site's. The full catalogue with fixes is in Errors.

Every response, success or error, carries X-Request-Id. Pass it to GET /v1/requests/{id} for the full trace, and quote it to support.

On this page