# API reference

> Base URL, authentication, content types, strict JSON, the GET and POST forms of /v1/scrape, pagination and retries for the Spicrawl REST API.

Source: https://docs.spicrawl.com/api-reference/introduction

The Spicrawl API is a JSON-over-HTTPS REST API at `https://api.spicrawl.com`. Every route is under `/v1`, every request carries `Authorization: Bearer <key>`, and every error is `application/problem+json` with a stable `code`.

```bash
curl -sS "https://api.spicrawl.com/v1/scrape" \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/products/42", "response_format": "markdown"}'
```

## Base URL

```text
https://api.spicrawl.com
```

| Group                                 | Routes                                                           | Scope                                                                                                       |
| ------------------------------------- | ---------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| Scrape                                | `GET`, `POST /v1/scrape`                                         | `scrape`                                                                                                    |
| Batch                                 | `/v1/batch`, `/v1/batch/{batchID}/…`                             | `batch`                                                                                                     |
| Sessions                              | `/v1/sessions`, `/v1/sessions/{sessionID}/…`                     | `sessions`                                                                                                  |
| Browser (CDP WebSocket) (coming soon) | `GET /v1/browser`                                                | `browser`                                                                                                   |
| Request log                           | `GET /v1/requests`, `GET /v1/requests/{id}`                      | none beyond a valid key for the key's project; `read` for `all_projects=true` and another project's request |
| Usage                                 | `GET /v1/usage`, `/v1/usage/summary`, `/v1/usage/reconciliation` | `read`                                                                                                      |
| Fleet                                 | `GET /v1/workers`                                                | none beyond a valid key                                                                                     |
| Health                                | `GET /healthz`, `GET /readyz`                                    | no auth                                                                                                     |

## Authentication

Send the key as a Bearer token. Read it from `SPICRAWL_API_KEY`.

```http
Authorization: Bearer spicrawl_live_…
```

* Keys are `spicrawl_live_…` (spends credits) or `spicrawl_test_…` (can never spend live credits).
* Every route ignores a query key (`?apikey=`) and returns `401 ERR::AUTH::MISSING_KEY` without the header.
* A key without the route's scope gets `403 ERR::AUTH::INSUFFICIENT_SCOPE`.

See [Authentication](https://docs.spicrawl.com/authentication.md).

## Content types

| Direction                    | Type                                                                                         |
| ---------------------------- | -------------------------------------------------------------------------------------------- |
| Request bodies               | `application/json`. Any other `Content-Type` is `400 ERR::REQUEST::INVALID`.                 |
| JSON responses               | `application/json`                                                                           |
| `/v1/scrape` document bodies | `text/html`, `text/markdown`, `text/plain` or `application/pdf`, following `response_format` |
| Batch results                | `application/x-ndjson` (one JSON object per line)                                            |
| Errors                       | `application/problem+json`                                                                   |

## Strict JSON

Request bodies are decoded strictly. An unknown or misspelled field is refused with `400 ERR::REQUEST::INVALID_PARAMETER`, and `detail` names it. The same rule applies to query parameters on `GET /v1/scrape`. Nothing is silently ignored.

```json title="400 ERR::REQUEST::INVALID_PARAMETER"
{
  "type": "https://docs.spicrawl.com/errors#REQUEST_INVALID_PARAMETER",
  "title": "A parameter has an invalid value",
  "status": 400,
  "code": "ERR::REQUEST::INVALID_PARAMETER",
  "detail": "Json: unknown field \"headers\".",
  "retryable": false,
  "doc_url": "https://docs.spicrawl.com/errors#REQUEST_INVALID_PARAMETER",
  "target_status": null
}
```

Headers sent to the target go in `custom_headers`, not `headers`. Browser-only flags such as `wait_for` on the plain `fetch` engine are also a `400`, never dropped.

## GET and POST forms of /v1/scrape

`GET /v1/scrape` and `POST /v1/scrape` accept the same parameters and behave the same. A parameter the JSON body would reject is a `400` in the query string too.

In the GET form, encode values like this:

| Parameter type                                                                            | GET encoding      | Example                                         |
| ----------------------------------------------------------------------------------------- | ----------------- | ----------------------------------------------- |
| String, integer, boolean                                                                  | Plain value       | `js_render=true&wait=2000`                      |
| List (`block_resources`, `include_tags`, `exclude_tags`, `allowed_status_codes`)          | Comma-separated   | `block_resources=image,font`                    |
| Object or array (`actions`, `custom_headers`, `extract`, `ai_extract`, `network_capture`) | JSON, URL-encoded | `extract=%7B%22title%22%3A%22h1%22%7D`          |
| `url`                                                                                     | URL-encoded       | `url=https%3A%2F%2Fexample.com%2Fproducts%2F42` |

```bash title="GET"
curl -sS -G "https://api.spicrawl.com/v1/scrape" \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  --data-urlencode "url=https://example.com/products/42" \
  --data-urlencode 'extract={"title":"h1","price":".price"}' \
  --data-urlencode "block_resources=image,font" \
  --data-urlencode "js_render=true"
```

```bash title="POST"
curl -sS "https://api.spicrawl.com/v1/scrape" \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/products/42",
    "extract": {"title": "h1", "price": ".price"},
    "block_resources": ["image", "font"],
    "js_render": true
  }'
```

Prefer POST for anything with objects or arrays. `url` in the GET form is limited to 8,192 characters.

## Pagination

List endpoints use three cursor styles. In every case, pass the cursor back verbatim; never construct one.

| Endpoint                          | Page size                   | Cursor in                                         | Send back as                   | Last page when             |
| --------------------------------- | --------------------------- | ------------------------------------------------- | ------------------------------ | -------------------------- |
| `GET /v1/batch`                   | `limit` 1-200, default 50   | body `next_cursor`                                | `cursor`                       | `next_cursor` is absent    |
| `GET /v1/sessions`                | `limit` 1-200, default 50   | body `next_cursor`                                | `cursor`                       | `next_cursor` is absent    |
| `GET /v1/batch/{batchID}/results` | `limit` 1-5000, default 500 | header `X-Next-Cursor` (also in `Link`)           | `cursor`                       | `X-Next-Cursor` is absent  |
| `GET /v1/requests`                | `limit` 1-200, default 50   | body `page.next_before` and `page.next_before_id` | `before` and `before_id`, both | `page.has_more` is `false` |

* `/v1/batch` and `/v1/sessions` refuse an out-of-range `limit` with `400`. `/v1/requests` clamps it silently.
* A page can hold fewer items than `limit` and still have more after it. Keep paging until the "last page" condition in the table holds.
* A `/v1/requests` cursor older than the retention window returns `410 ERR::REQUEST::BEYOND_RETENTION`. Restart from the first page.
* With `GET /v1/requests?all_projects=true`, send `all_projects=true` on every page: a cursor pages the list it came from.
* Batch results are JSON Lines and readable while the job runs; on a running job, a later call can return more items after the last cursor.

## Retries and idempotency

There is no idempotency key on any route. Every `POST /v1/scrape` performs and bills a new scrape unless it is served from the cache, and every `POST /v1/batch` creates a new job.

* Retry a request that returned an error with `retryable: true`, after `Retry-After`. Failures cost 0 credits, so that retry is free.
* Do not retry a `retryable: false` error unchanged. Apply `diagnostics.hint`.
* If a connection drops before you see the response, a retry can bill twice. For a batch, check `GET /v1/batch` before resubmitting.
* If a `202` from `POST /v1/batch` carries a warning that some items were not dispatched, do not resubmit. The platform dispatches them.

## Errors

Every error is `application/problem+json`:

```json
{
  "type": "https://docs.spicrawl.com/errors#UPSTREAM_TIMEOUT",
  "title": "Target did not respond in time",
  "status": 504,
  "code": "ERR::UPSTREAM::TIMEOUT",
  "detail": "The target did not answer within the time budget.",
  "retryable": true,
  "doc_url": "https://docs.spicrawl.com/errors#UPSTREAM_TIMEOUT",
  "request_id": "01M0HF5WFWE7PRE8KZHDTNETWN",
  "instance": "/v1/scrape",
  "target_status": null,
  "diagnostics": {
    "hint": "The target did not answer in time. If the page is JavaScript-rendered, try `js_render=true`."
  }
}
```

Switch on `code`, retry on `retryable`, honour `Retry-After`, and read `diagnostics.hint`. `status` is the platform's HTTP status; `target_status` is the site's. The full catalogue with fixes is in [Errors](https://docs.spicrawl.com/errors.md).

Every response, success or error, carries `X-Request-Id`. Pass it to `GET /v1/requests/{id}` for the full trace, and quote it to support.
