# MCP server

> Connect any MCP client to the hosted Spicrawl MCP server at https://mcp.spicrawl.com/mcp and use its 25 spicrawl_* tools to scrape, batch, manage sessions, inspect usage and search the docs.

Source: https://docs.spicrawl.com/agents/mcp

The Spicrawl MCP server gives an agent Spicrawl as native tools. It is hosted at:

```text
https://mcp.spicrawl.com/mcp
```

It speaks MCP Streamable HTTP and authenticates with your own API key as a bearer token. Every tool call runs under that key, exactly as if you had called the API yourself: same scopes, same limits, same credits.

```bash title="Claude Code"
claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
  --header "Authorization: Bearer $SPICRAWL_API_KEY"
```

## Authentication

Send `Authorization: Bearer spicrawl_live_…` (or a `spicrawl_test_…` key) on every request to the endpoint. The key is checked against the API when the MCP session opens:

* An unknown, revoked or expired key is refused with HTTP `401` and a `WWW-Authenticate: Bearer realm="spicrawl-mcp"` header before any session exists.
* If the API cannot be reached to check the key, the endpoint answers `503` with `Retry-After: 5`. Retry; do not rotate the key.
* A session is bound to the key that opened it.

Keys get the `scrape`, `batch` and `sessions` scopes by default. Two tool groups need a scope you grant deliberately in the dashboard:

| Tools                                                                       | Scope needed | Without it                                 |
| --------------------------------------------------------------------------- | ------------ | ------------------------------------------ |
| `spicrawl_usage`, `spicrawl_usage_summary`, `spicrawl_usage_reconciliation` | `read`       | `ERR::AUTH::INSUFFICIENT_SCOPE` (HTTP 403) |
| `spicrawl_browser_connect_url` (coming soon)                                | `browser`    | The WebSocket upgrade fails with 403       |

## Set up your client

Put the key in the `SPICRAWL_API_KEY` environment variable and reference it from the config, so the key never lands in a committed file.

**Claude Code**

One command, stored for you only (local scope):

```bash
claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
  --header "Authorization: Bearer $SPICRAWL_API_KEY"
```

Add `--scope user` to use it in every project. To share it with your team, commit a `.mcp.json` at the project root instead. Claude Code expands `${VAR}` in `url` and `headers`, so each teammate's own key is used:

```json title=".mcp.json"
{
  "mcpServers": {
    "spicrawl": {
      "type": "http",
      "url": "https://mcp.spicrawl.com/mcp",
      "headers": { "Authorization": "Bearer ${SPICRAWL_API_KEY}" }
    }
  }
}
```

Check it with `claude mcp list`, or `/mcp` inside a session. More in [Claude Code](https://docs.spicrawl.com/agents/claude-code.md).

**Claude Desktop**

Claude Desktop's `claude_desktop_config.json` starts local (stdio) servers only, and custom connectors added under Settings → Connectors cannot send an `Authorization` header. Bridge to the hosted server with [`mcp-remote`](https://www.npmjs.com/package/mcp-remote) (needs Node.js; `npx` fetches it):

```json title="claude_desktop_config.json"
{
  "mcpServers": {
    "spicrawl": {
      "command": "npx",
      "args": [
        "-y", "mcp-remote", "https://mcp.spicrawl.com/mcp",
        "--header", "Authorization:${SPICRAWL_AUTH}"
      ],
      "env": { "SPICRAWL_AUTH": "Bearer spicrawl_live_..." }
    }
  }
}
```

The header value goes through the `SPICRAWL_AUTH` variable because some clients split arguments on spaces. This file is on your machine only, so the key is written into it; keep it out of version control. Restart Claude Desktop after saving.

**Cursor**

Add to `.cursor/mcp.json` in the project, or `~/.cursor/mcp.json` for every project. Cursor resolves `${env:NAME}` in `url` and `headers`:

```json title=".cursor/mcp.json"
{
  "mcpServers": {
    "spicrawl": {
      "url": "https://mcp.spicrawl.com/mcp",
      "headers": { "Authorization": "Bearer ${env:SPICRAWL_API_KEY}" }
    }
  }
}
```

Cursor must be started from an environment where `SPICRAWL_API_KEY` is set. More in [Cursor](https://docs.spicrawl.com/agents/cursor.md).

**VS Code**

Add to `.vscode/mcp.json`. VS Code uses the `servers` key (not `mcpServers`). Use an input variable so VS Code prompts for the key once and stores it securely:

```json title=".vscode/mcp.json"
{
  "inputs": [
    {
      "type": "promptString",
      "id": "spicrawl-api-key",
      "description": "Spicrawl API key",
      "password": true
    }
  ],
  "servers": {
    "spicrawl": {
      "type": "http",
      "url": "https://mcp.spicrawl.com/mcp",
      "headers": { "Authorization": "Bearer ${input:spicrawl-api-key}" }
    }
  }
}
```

Start the server from the Chat view's tools list or with the **MCP: List Servers** command. VS Code does not expand `${env:...}` inside `headers`, so use the input variable above rather than an environment variable. More in [GitHub Copilot and VS Code](https://docs.spicrawl.com/agents/github-copilot.md).

**Codex**

Codex reads MCP servers from `~/.codex/config.toml`, or `.codex/config.toml` in a trusted project. `bearer_token_env_var` names the variable that holds the key:

```toml title="~/.codex/config.toml"
[mcp_servers.spicrawl]
url = "https://mcp.spicrawl.com/mcp"
bearer_token_env_var = "SPICRAWL_API_KEY"
```

Or from the command line:

```bash
codex mcp add spicrawl --url https://mcp.spicrawl.com/mcp --bearer-token-env-var SPICRAWL_API_KEY
```

More in [Codex](https://docs.spicrawl.com/agents/codex.md).

**Other clients**

Any MCP client that supports Streamable HTTP and custom headers works:

| Setting   | Value                              |
| --------- | ---------------------------------- |
| Transport | Streamable HTTP                    |
| URL       | `https://mcp.spicrawl.com/mcp`     |
| Header    | `Authorization: Bearer <your key>` |

A client that supports only stdio can bridge with `npx -y mcp-remote https://mcp.spicrawl.com/mcp --header "Authorization:${SPICRAWL_AUTH}"`, as in the Claude Desktop tab. A client that supports only OAuth cannot connect: the server accepts API keys only. A client that tries OAuth when the server answers `401` needs OAuth turned off for this server.

Step-by-step setup for more clients: [GitHub Copilot and VS Code](https://docs.spicrawl.com/agents/github-copilot.md), [OpenCode](https://docs.spicrawl.com/agents/opencode.md), [Gemini CLI](https://docs.spicrawl.com/agents/gemini-cli.md), [Windsurf](https://docs.spicrawl.com/agents/windsurf.md), [Cline](https://docs.spicrawl.com/agents/cline.md), [Kilo Code](https://docs.spicrawl.com/agents/kilo-code.md), [Zed](https://docs.spicrawl.com/agents/zed.md), [JetBrains](https://docs.spicrawl.com/agents/jetbrains.md), [Goose](https://docs.spicrawl.com/agents/goose.md), [Kiro](https://docs.spicrawl.com/agents/kiro.md), [Factory Droid](https://docs.spicrawl.com/agents/factory.md), [Devin](https://docs.spicrawl.com/agents/devin.md), [OpenHands](https://docs.spicrawl.com/agents/openhands.md), [AI app builders](https://docs.spicrawl.com/agents/app-builders.md), [agent frameworks](https://docs.spicrawl.com/agents/frameworks.md) and [more MCP clients](https://docs.spicrawl.com/agents/other-clients.md).

## Tools

The server exposes 25 tools. Argument names are the API's own except on `spicrawl_scrape` and `spicrawl_batch_submit`, which use `render` for `js_render` and `format` for `response_format`. Tools reject unknown arguments: a misspelt field (for example `js_render` on `spicrawl_scrape`) fails the call with an error naming it, instead of being dropped.

### Scrape

| Tool              | What it does                                                                                  | Key inputs                                                                                                                                                                                                                         | Use it when                              |
| ----------------- | --------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- |
| `spicrawl_scrape` | `POST /v1/scrape`: retrieve one URL as `markdown` (default), `text`, `html`, `json` or `pdf`. | `url` (required), `format`, `render`, `main_content_only`, `include_tags`, `exclude_tags`, `extract`, `ai_extract`, `autoparse`, `links`, `screenshot`, `mode`, `engine`, `session_id`, `wait_for`, `actions`, `cache`, `max_cost` | You need one page's content or data now. |

Every other `POST /v1/scrape` field is accepted under its API name: `impersonate`, `proxy`, `proxy_verify`, `method`, `custom_headers`, `headless`, `wait`, `wait_for_timeout`, `block_resources`, `network_capture`, `screenshot_fullpage`, `screenshot_selector`, `screenshot_format`, `screenshot_quality`, `parse_pdf`, `cache_ttl`, `original_status`, `allowed_status_codes`, `extract_preset`.

What comes back:

* With `format: "markdown"`, `"text"` or `"html"`, the tool returns the document itself. Response headers such as `X-Target-Status` and `X-Credits-Charged` are not passed through.
* With `format: "json"`, or whenever `extract`, `ai_extract`, `autoparse`, `links`, `network_capture` or `screenshot` is set, the tool returns the JSON envelope: `content`, the site's `status`, `credits`, `engine`, `warnings`, `empty_fields`, and `data` for extraction. Use it when the agent must check the site's status or the cost.
* Screenshots come back as MCP image blocks the model can see. Each `screenshots` entry in the JSON keeps its metadata and points to its block. An image over 5 MB of base64 is not attached; its entry says how to get a smaller one.
* `format: "pdf"` prints the page in a browser (set `render` or `engine: "chromium"`) and returns the file as an MCP resource block (`application/pdf`). The JSON beside it has `engine`, `status`, `credits`, `request_id` and `pdf.size_bytes`. A PDF over 10 MB of base64 is not attached; its `pdf.not_attached` says to call `POST /v1/scrape` directly.
* `actions` is typed for the tool: each step is `{"type": "click" | "fill" | "wait_for" | "wait_for_navigation" | "scroll" | "select" | "evaluate" | "screenshot", ...fields}`, with the fields of the [browser actions](https://docs.spicrawl.com/guides/browser-actions.md) verb plus `timeout_ms`, `on_error`, `label`. The tool sends them in the API's `{"click": {...}}` shape. Infinite scroll: `[{"type":"scroll","to_bottom":true},{"type":"wait_for","ms":1000}]`. A `screenshot` step returns the JSON envelope whatever `format` says.
* Error messages and warnings name the tool's arguments: the API's `js_render` reads as `render`, `response_format` as `format`.
* `ai_extract` (coming soon) takes `prompt` or `schema`, not both, and adds 4 credits. Until AI extraction launches it fails with `ERR::INTERNAL::UNAVAILABLE` at 0 credits; do not retry.

### Batch

| Tool                          | What it does                                                                                                           | Key inputs                                                                                                               | Use it when                                                                            |
| ----------------------------- | ---------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------- |
| `spicrawl_batch_submit`       | `POST /v1/batch`: queue many URLs as one async job. Returns the job (`id`, `status`, `estimated_credits`, `progress`). | `urls` or `items` (not both; up to 10,000), `render`, `block_resources`, `name`, `open`, `max_attempts`, `credit_budget` | You have tens to thousands of URLs.                                                    |
| `spicrawl_batch_list`         | List the project's jobs, newest first.                                                                                 | `status`, `limit` (max 200), `cursor`                                                                                    | You lost a job id.                                                                     |
| `spicrawl_batch_status`       | One job's status and progress.                                                                                         | `job_id`                                                                                                                 | Polling after submit, until `completed`, `failed` or `cancelled`.                      |
| `spicrawl_batch_results`      | A page of finished items (content, or a `result_url` API content path per item).                                       | `job_id`, `status`, `limit` (default 500, max 5000), `cursor`                                                            | Reading results; works while the job runs.                                             |
| `spicrawl_batch_task_content` | The full payload of one item by `seq` (0-based, submission order).                                                     | `job_id`, `seq`                                                                                                          | You need one URL's output. `ERR::REQUEST::CONFLICT` means not finished yet, or failed. |
| `spicrawl_batch_cancel`       | Cancel: unstarted items never run, finished ones keep results. Cannot be undone.                                       | `job_id`                                                                                                                 | The job is wrong or too expensive.                                                     |
| `spicrawl_batch_retry`        | Re-queue every failed item.                                                                                            | `job_id`                                                                                                                 | Items failed with retryable errors.                                                    |
| `spicrawl_batch_add_items`    | Append URLs to a job submitted with `open: true`. Not idempotent.                                                      | `job_id`, `urls` or `items`                                                                                              | You discover URLs while the job runs (crawling).                                       |
| `spicrawl_batch_close`        | Stop an open job accepting items so it can complete.                                                                   | `job_id`                                                                                                                 | You are done adding to an open job. An open job never finishes until closed.           |

> **Warning:** Batch items apply only `js_render` (`render`), `proxy` and `block_resources` today, and return the raw page. Other settings, including `format`, are not applied per item. For markdown or extraction, run `spicrawl_scrape` per URL.

`max_cost` on a batch caps one item; `credit_budget` caps the whole job. Results expire after the job's retention window (HTTP 410).

### Sessions

| Tool                       | What it does                                                             | Key inputs                                                                   | Use it when                                                                                                              |
| -------------------------- | ------------------------------------------------------------------------ | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `spicrawl_session_create`  | Create a persistent browser identity: cookies, storage, pinned engine.   | `engine`, `ttl_seconds` (min 30, default 1800), `session_context` (to clone) | Multi-step flows: log in, then read pages behind the login. Pass the returned `id` as `session_id` to `spicrawl_scrape`. |
| `spicrawl_session_list`    | List sessions, newest first. Never includes credentials.                 | `status`, `engine`, `limit`, `cursor`                                        | Finding a session.                                                                                                       |
| `spicrawl_session_get`     | One session's metadata: status, engine, exit, usage, expiry.             | `session_id`                                                                 | Checking a session is still `active`.                                                                                    |
| `spicrawl_session_context` | Dump a live session's cookies and storage. Contains secrets.             | `session_id`                                                                 | Cloning a session. Do not echo the result to the user or logs.                                                           |
| `spicrawl_session_release` | End a session; the record stays as `released` and its context is purged. | `session_id`, `force`                                                        | You are done with it.                                                                                                    |
| `spicrawl_session_delete`  | Permanently delete a session and its record.                             | `session_id`, `force`                                                        | Removing it entirely.                                                                                                    |

A session runs one request at a time. A second concurrent scrape on it fails with `ERR::SESSION::BUSY` (retryable).

### Usage and request history

| Tool                            | What it does                                                                                                                                                                                             | Key inputs                                                                                                                                                                                                          | Use it when                                    |
| ------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------- |
| `spicrawl_requests_list`        | Recent API calls in the key's project with status, error code, engine, proxy, latency, credits and `project_id`. Any key; `all_projects: true` lists every project in the organization and needs `read`. | `all_projects`, `only_errors`, `status`, `limit` (max 200), `before` + `before_id`                                                                                                                                  | Finding failing calls or patterns on one site. |
| `spicrawl_request_get`          | The full log record of one call. Any key for its own project's calls; another project's need `read`.                                                                                                     | `id` (the `X-Request-Id` or `request_id` of an error)                                                                                                                                                               | Debugging one failed or slow scrape.           |
| `spicrawl_usage`                | Billing-grade usage over a date window. Needs `read`.                                                                                                                                                    | `from`, `to` (exclusive, `YYYY-MM-DD`), `group_by` (`day`, `project`, `engine`, `feature`, `key`: per API key id/name/prefix, per day), `metrics` (billing metrics plus `requests_failed`, `feature_*`, `client_*`) | "How many credits did I use, and where?"       |
| `spicrawl_usage_summary`        | Current-period totals. Needs `read`.                                                                                                                                                                     | none                                                                                                                                                                                                                | A quick "how much have I used".                |
| `spicrawl_usage_reconciliation` | One day's drift between the request log and the billed rollup. Needs `read`.                                                                                                                             | `day` (default yesterday)                                                                                                                                                                                           | Auditing billing.                              |

### Browser

> **Coming soon:** `spicrawl_browser_connect_url` will return a URL for connecting Puppeteer or Playwright to a browser hosted by Spicrawl. Remote browsers are not available yet, so agents should not rely on this tool yet. See [CDP browser](https://docs.spicrawl.com/guides/cdp-browser.md).

### Docs

| Tool                   | What it does                                                                                                                                                                                                                                        | Key inputs                                                                                                  | Use it when                                                              |
| ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ |
| `spicrawl_docs_search` | Full-text search over these docs. Returns up to `limit` pages, each with its title, matching sections (heading, snippet, URL) and `md_url`. An error code as the query (`ERR::FAMILY::NAME`) also returns a link to its entry on [Errors](https://docs.spicrawl.com/errors.md). | `query` (required, up to 200 characters), `limit` (1-20, default 8)                                         | Before guessing a parameter name or value, and to explain an error code. |
| `spicrawl_docs_read`   | One docs page as Markdown. Pages over 60,000 characters are truncated, and the text says so.                                                                                                                                                        | `path`: a page path (`guides/anti-bot`), a docs URL from `spicrawl_docs_search`, or either with a `#anchor` | You know which page answers the question.                                |
| `spicrawl_docs_index`  | The docs' `llms.txt`: every page with its title, one-line description and `.md` URL.                                                                                                                                                                | none                                                                                                        | A search finds nothing, or the agent needs an overview.                  |

The docs tools read the public docs at `https://docs.spicrawl.com` and never send your API key there. A self-hosted server reads `$SPICRAWL_DOCS_URL`, a docs base URL such as `http://<host>:8080/docs`. When it is unset: the legacy `$SPICRAWL_DOCS_HOST` plus `/docs`, then `SPICRAWL_PUBLIC_BASE_URL`, then `SPICRAWL_BASE_URL`, plus `/docs` (a self-hosted API serves its own docs there); with none of them set, or with the public API as the base URL, the public docs site.

## How errors reach the agent

A failed call comes back as an MCP tool error (`isError: true`) whose text is the API's error `code`, its `detail`, a retryable note when `retryable` is `true`, and the `doc_url`:

```text
ERR::UPSTREAM::CHALLENGE: <the API's detail text> (retryable — the same request may succeed on a retry)
See https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE
```

The agent should act on the code: retry when the message says `retryable`, and change parameters otherwise. For `ERR::UPSTREAM::CHALLENGE`, retry once, then add `render: true` or the user's own `proxy`. Failed calls cost 0 credits. The full error table is on [Errors](https://docs.spicrawl.com/errors.md).


## Self-host over stdio

Run the server as a local subprocess of your MCP client instead of using the hosted endpoint. It needs Node.js 20 or later.

**Step 1: Build it**

From the `mcp/` directory of the Spicrawl source tree (provided to self-hosting customers; there is no public repository):

```bash
cd spicrawl/mcp
npm install
npm run build
```

This produces `dist/index.js` (stdio) and `dist/http.js` (Streamable HTTP).

**Step 2: Add it to your client**

```json
{
  "mcpServers": {
    "spicrawl": {
      "command": "node",
      "args": ["/absolute/path/to/spicrawl/mcp/dist/index.js"],
      "env": {
        "SPICRAWL_API_KEY": "spicrawl_live_...",
        "SPICRAWL_BASE_URL": "https://api.spicrawl.com"
      }
    }
  }
}
```

For Claude Code: `claude mcp add spicrawl --env SPICRAWL_API_KEY=$SPICRAWL_API_KEY --env SPICRAWL_BASE_URL=https://api.spicrawl.com -- node /absolute/path/to/spicrawl/mcp/dist/index.js`.

| Variable                   | Required | Default                    | Notes                                                                                     |
| -------------------------- | -------- | -------------------------- | ----------------------------------------------------------------------------------------- |
| `SPICRAWL_API_KEY`         | yes      | none                       | Your API key.                                                                             |
| `SPICRAWL_BASE_URL`        | no       | `https://api.spicrawl.com` | The API the tools call. Set it for a self-hosted API (`http://<host>:8080`).              |
| `SPICRAWL_PUBLIC_BASE_URL` | no       | `SPICRAWL_BASE_URL`        | Base used only in URLs handed back to you (`spicrawl_browser_connect_url` (coming soon)). |


## Troubleshooting

| Symptom                                                                            | Cause                                                                                                                             | Fix                                                                                                                                                                                                  |
| ---------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| HTTP `401` when the client connects, or "Unauthorized: send your Spicrawl API key" | No `Authorization` header, or the key is unknown, revoked or expired.                                                             | Check the header is `Authorization: Bearer spicrawl_…` and that the environment variable is set in the environment the client was started from. Create a new key in the dashboard if it was revoked. |
| The header contains a literal `${SPICRAWL_API_KEY}` or `${env:SPICRAWL_API_KEY}`   | The client did not expand the variable (wrong syntax for that client, or the variable is not set when the client starts).         | Use `${VAR}` in Claude Code, `${env:VAR}` in Cursor, `${input:id}` in VS Code, `bearer_token_env_var` in Codex. Start the client from a shell where the variable is set.                             |
| HTTP `503` with `Retry-After` on connect                                           | The MCP server could not reach the API to check your key.                                                                         | Retry after the given seconds. Your key is fine.                                                                                                                                                     |
| `ERR::AUTH::INSUFFICIENT_SCOPE` from a `spicrawl_usage*` tool                      | The key lacks the `read` scope.                                                                                                   | Grant `read` to the key in the dashboard, or use a key that has it.                                                                                                                                  |
| `ERR::LIMIT::QUOTA_EXCEEDED` (HTTP 402)                                            | Out of credits or over the monthly ceiling.                                                                                       | Stop; top up or wait for the reset. Not retryable.                                                                                                                                                   |
| "Unknown argument" or "Unrecognized key" on a tool call                            | The agent used a field the tool does not accept, often an API name where the tool uses its own (`js_render` instead of `render`). | Use the tool's argument name, listed above.                                                                                                                                                          |
| Tools do not appear                                                                | The client was not restarted, or the config is in the wrong file or key (`servers` in VS Code, `mcpServers` elsewhere).           | Restart the client and check its MCP panel or logs.                                                                                                                                                  |

Check that the endpoint itself is up with `curl -i https://mcp.spicrawl.com/mcp`. Sent without a key, it answers HTTP `401` with `Unauthorized: send your Spicrawl API key`, which means the server is reachable; a connection error or a `5xx` points at the network or the server, not your client's configuration.
