spicrawlspicrawlDocs

MCP server

Connect any MCP client to the hosted Spicrawl MCP server at https://mcp.spicrawl.com/mcp and use its 25 spicrawl_* tools to scrape, batch, manage sessions, inspect usage and search the docs.

The Spicrawl MCP server gives an agent Spicrawl as native tools. It is hosted at:

https://mcp.spicrawl.com/mcp

It speaks MCP Streamable HTTP and authenticates with your own API key as a bearer token. Every tool call runs under that key, exactly as if you had called the API yourself: same scopes, same limits, same credits.

Claude Code
claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
  --header "Authorization: Bearer $SPICRAWL_API_KEY"

Authentication

Send Authorization: Bearer spicrawl_live_… (or a spicrawl_test_… key) on every request to the endpoint. The key is checked against the API when the MCP session opens:

  • An unknown, revoked or expired key is refused with HTTP 401 and a WWW-Authenticate: Bearer realm="spicrawl-mcp" header before any session exists.
  • If the API cannot be reached to check the key, the endpoint answers 503 with Retry-After: 5. Retry; do not rotate the key.
  • A session is bound to the key that opened it.

Keys get the scrape, batch and sessions scopes by default. Two tool groups need a scope you grant deliberately in the dashboard:

ToolsScope neededWithout it
spicrawl_usage, spicrawl_usage_summary, spicrawl_usage_reconciliationreadERR::AUTH::INSUFFICIENT_SCOPE (HTTP 403)
spicrawl_browser_connect_url Coming soonbrowserThe WebSocket upgrade fails with 403

Set up your client

Put the key in the SPICRAWL_API_KEY environment variable and reference it from the config, so the key never lands in a committed file.

One command, stored for you only (local scope):

claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
  --header "Authorization: Bearer $SPICRAWL_API_KEY"

Add --scope user to use it in every project. To share it with your team, commit a .mcp.json at the project root instead. Claude Code expands ${VAR} in url and headers, so each teammate's own key is used:

.mcp.json
{
  "mcpServers": {
    "spicrawl": {
      "type": "http",
      "url": "https://mcp.spicrawl.com/mcp",
      "headers": { "Authorization": "Bearer ${SPICRAWL_API_KEY}" }
    }
  }
}

Check it with claude mcp list, or /mcp inside a session. More in Claude Code.

Tools

The server exposes 25 tools. Argument names are the API's own except on spicrawl_scrape and spicrawl_batch_submit, which use render for js_render and format for response_format. Tools reject unknown arguments: a misspelt field (for example js_render on spicrawl_scrape) fails the call with an error naming it, instead of being dropped.

Scrape

ToolWhat it doesKey inputsUse it when
spicrawl_scrapePOST /v1/scrape: retrieve one URL as markdown (default), text, html, json or pdf.url (required), format, render, main_content_only, include_tags, exclude_tags, extract, ai_extract, autoparse, links, screenshot, mode, engine, session_id, wait_for, actions, cache, max_costYou need one page's content or data now.

Every other POST /v1/scrape field is accepted under its API name: impersonate, proxy, proxy_verify, method, custom_headers, headless, wait, wait_for_timeout, block_resources, network_capture, screenshot_fullpage, screenshot_selector, screenshot_format, screenshot_quality, parse_pdf, cache_ttl, original_status, allowed_status_codes, extract_preset.

What comes back:

  • With format: "markdown", "text" or "html", the tool returns the document itself. Response headers such as X-Target-Status and X-Credits-Charged are not passed through.
  • With format: "json", or whenever extract, ai_extract, autoparse, links, network_capture or screenshot is set, the tool returns the JSON envelope: content, the site's status, credits, engine, warnings, empty_fields, and data for extraction. Use it when the agent must check the site's status or the cost.
  • Screenshots come back as MCP image blocks the model can see. Each screenshots entry in the JSON keeps its metadata and points to its block. An image over 5 MB of base64 is not attached; its entry says how to get a smaller one.
  • format: "pdf" prints the page in a browser (set render or engine: "chromium") and returns the file as an MCP resource block (application/pdf). The JSON beside it has engine, status, credits, request_id and pdf.size_bytes. A PDF over 10 MB of base64 is not attached; its pdf.not_attached says to call POST /v1/scrape directly.
  • actions is typed for the tool: each step is {"type": "click" | "fill" | "wait_for" | "wait_for_navigation" | "scroll" | "select" | "evaluate" | "screenshot", ...fields}, with the fields of the browser actions verb plus timeout_ms, on_error, label. The tool sends them in the API's {"click": {...}} shape. Infinite scroll: [{"type":"scroll","to_bottom":true},{"type":"wait_for","ms":1000}]. A screenshot step returns the JSON envelope whatever format says.
  • Error messages and warnings name the tool's arguments: the API's js_render reads as render, response_format as format.
  • ai_extract Coming soon takes prompt or schema, not both, and adds 4 credits. Until AI extraction launches it fails with ERR::INTERNAL::UNAVAILABLE at 0 credits; do not retry.

Batch

ToolWhat it doesKey inputsUse it when
spicrawl_batch_submitPOST /v1/batch: queue many URLs as one async job. Returns the job (id, status, estimated_credits, progress).urls or items (not both; up to 10,000), render, block_resources, name, open, max_attempts, credit_budgetYou have tens to thousands of URLs.
spicrawl_batch_listList the project's jobs, newest first.status, limit (max 200), cursorYou lost a job id.
spicrawl_batch_statusOne job's status and progress.job_idPolling after submit, until completed, failed or cancelled.
spicrawl_batch_resultsA page of finished items (content, or a result_url API content path per item).job_id, status, limit (default 500, max 5000), cursorReading results; works while the job runs.
spicrawl_batch_task_contentThe full payload of one item by seq (0-based, submission order).job_id, seqYou need one URL's output. ERR::REQUEST::CONFLICT means not finished yet, or failed.
spicrawl_batch_cancelCancel: unstarted items never run, finished ones keep results. Cannot be undone.job_idThe job is wrong or too expensive.
spicrawl_batch_retryRe-queue every failed item.job_idItems failed with retryable errors.
spicrawl_batch_add_itemsAppend URLs to a job submitted with open: true. Not idempotent.job_id, urls or itemsYou discover URLs while the job runs (crawling).
spicrawl_batch_closeStop an open job accepting items so it can complete.job_idYou are done adding to an open job. An open job never finishes until closed.

Batch items apply only js_render (render), proxy and block_resources today, and return the raw page. Other settings, including format, are not applied per item. For markdown or extraction, run spicrawl_scrape per URL.

max_cost on a batch caps one item; credit_budget caps the whole job. Results expire after the job's retention window (HTTP 410).

Sessions

ToolWhat it doesKey inputsUse it when
spicrawl_session_createCreate a persistent browser identity: cookies, storage, pinned engine.engine, ttl_seconds (min 30, default 1800), session_context (to clone)Multi-step flows: log in, then read pages behind the login. Pass the returned id as session_id to spicrawl_scrape.
spicrawl_session_listList sessions, newest first. Never includes credentials.status, engine, limit, cursorFinding a session.
spicrawl_session_getOne session's metadata: status, engine, exit, usage, expiry.session_idChecking a session is still active.
spicrawl_session_contextDump a live session's cookies and storage. Contains secrets.session_idCloning a session. Do not echo the result to the user or logs.
spicrawl_session_releaseEnd a session; the record stays as released and its context is purged.session_id, forceYou are done with it.
spicrawl_session_deletePermanently delete a session and its record.session_id, forceRemoving it entirely.

A session runs one request at a time. A second concurrent scrape on it fails with ERR::SESSION::BUSY (retryable).

Usage and request history

ToolWhat it doesKey inputsUse it when
spicrawl_requests_listRecent API calls in the key's project with status, error code, engine, proxy, latency, credits and project_id. Any key; all_projects: true lists every project in the organization and needs read.all_projects, only_errors, status, limit (max 200), before + before_idFinding failing calls or patterns on one site.
spicrawl_request_getThe full log record of one call. Any key for its own project's calls; another project's need read.id (the X-Request-Id or request_id of an error)Debugging one failed or slow scrape.
spicrawl_usageBilling-grade usage over a date window. Needs read.from, to (exclusive, YYYY-MM-DD), group_by (day, project, engine, feature, key: per API key id/name/prefix, per day), metrics (billing metrics plus requests_failed, feature_*, client_*)"How many credits did I use, and where?"
spicrawl_usage_summaryCurrent-period totals. Needs read.noneA quick "how much have I used".
spicrawl_usage_reconciliationOne day's drift between the request log and the billed rollup. Needs read.day (default yesterday)Auditing billing.

Browser

Coming soon

spicrawl_browser_connect_url will return a URL for connecting Puppeteer or Playwright to a browser hosted by Spicrawl. Remote browsers are not available yet, so agents should not rely on this tool yet. See CDP browser.

Docs

ToolWhat it doesKey inputsUse it when
spicrawl_docs_searchFull-text search over these docs. Returns up to limit pages, each with its title, matching sections (heading, snippet, URL) and md_url. An error code as the query (ERR::FAMILY::NAME) also returns a link to its entry on Errors.query (required, up to 200 characters), limit (1-20, default 8)Before guessing a parameter name or value, and to explain an error code.
spicrawl_docs_readOne docs page as Markdown. Pages over 60,000 characters are truncated, and the text says so.path: a page path (guides/anti-bot), a docs URL from spicrawl_docs_search, or either with a #anchorYou know which page answers the question.
spicrawl_docs_indexThe docs' llms.txt: every page with its title, one-line description and .md URL.noneA search finds nothing, or the agent needs an overview.

The docs tools read the public docs at https://docs.spicrawl.com and never send your API key there. A self-hosted server reads $SPICRAWL_DOCS_URL, a docs base URL such as http://<host>:8080/docs. When it is unset: the legacy $SPICRAWL_DOCS_HOST plus /docs, then SPICRAWL_PUBLIC_BASE_URL, then SPICRAWL_BASE_URL, plus /docs (a self-hosted API serves its own docs there); with none of them set, or with the public API as the base URL, the public docs site.

How errors reach the agent

A failed call comes back as an MCP tool error (isError: true) whose text is the API's error code, its detail, a retryable note when retryable is true, and the doc_url:

ERR::UPSTREAM::CHALLENGE: <the API's detail text> (retryable — the same request may succeed on a retry)
See https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE

The agent should act on the code: retry when the message says retryable, and change parameters otherwise. For ERR::UPSTREAM::CHALLENGE, retry once, then add render: true or the user's own proxy. Failed calls cost 0 credits. The full error table is on Errors.

Self-host over stdio

Run the server as a local subprocess of your MCP client instead of using the hosted endpoint. It needs Node.js 20 or later.

Build it

From the mcp/ directory of the Spicrawl source tree (provided to self-hosting customers; there is no public repository):

cd spicrawl/mcp
npm install
npm run build

This produces dist/index.js (stdio) and dist/http.js (Streamable HTTP).

Add it to your client

{
  "mcpServers": {
    "spicrawl": {
      "command": "node",
      "args": ["/absolute/path/to/spicrawl/mcp/dist/index.js"],
      "env": {
        "SPICRAWL_API_KEY": "spicrawl_live_...",
        "SPICRAWL_BASE_URL": "https://api.spicrawl.com"
      }
    }
  }
}

For Claude Code: claude mcp add spicrawl --env SPICRAWL_API_KEY=$SPICRAWL_API_KEY --env SPICRAWL_BASE_URL=https://api.spicrawl.com -- node /absolute/path/to/spicrawl/mcp/dist/index.js.

VariableRequiredDefaultNotes
SPICRAWL_API_KEYyesnoneYour API key.
SPICRAWL_BASE_URLnohttps://api.spicrawl.comThe API the tools call. Set it for a self-hosted API (http://<host>:8080).
SPICRAWL_PUBLIC_BASE_URLnoSPICRAWL_BASE_URLBase used only in URLs handed back to you (spicrawl_browser_connect_url Coming soon).

Troubleshooting

SymptomCauseFix
HTTP 401 when the client connects, or "Unauthorized: send your Spicrawl API key"No Authorization header, or the key is unknown, revoked or expired.Check the header is Authorization: Bearer spicrawl_… and that the environment variable is set in the environment the client was started from. Create a new key in the dashboard if it was revoked.
The header contains a literal ${SPICRAWL_API_KEY} or ${env:SPICRAWL_API_KEY}The client did not expand the variable (wrong syntax for that client, or the variable is not set when the client starts).Use ${VAR} in Claude Code, ${env:VAR} in Cursor, ${input:id} in VS Code, bearer_token_env_var in Codex. Start the client from a shell where the variable is set.
HTTP 503 with Retry-After on connectThe MCP server could not reach the API to check your key.Retry after the given seconds. Your key is fine.
ERR::AUTH::INSUFFICIENT_SCOPE from a spicrawl_usage* toolThe key lacks the read scope.Grant read to the key in the dashboard, or use a key that has it.
ERR::LIMIT::QUOTA_EXCEEDED (HTTP 402)Out of credits or over the monthly ceiling.Stop; top up or wait for the reset. Not retryable.
"Unknown argument" or "Unrecognized key" on a tool callThe agent used a field the tool does not accept, often an API name where the tool uses its own (js_render instead of render).Use the tool's argument name, listed above.
Tools do not appearThe client was not restarted, or the config is in the wrong file or key (servers in VS Code, mcpServers elsewhere).Restart the client and check its MCP panel or logs.

Check that the endpoint itself is up with curl -i https://mcp.spicrawl.com/mcp. Sent without a key, it answers HTTP 401 with Unauthorized: send your Spicrawl API key, which means the server is reachable; a connection error or a 5xx points at the network or the server, not your client's configuration.

On this page