MCP server
Connect any MCP client to the hosted Spicrawl MCP server at https://mcp.spicrawl.com/mcp and use its 25 spicrawl_* tools to scrape, batch, manage sessions, inspect usage and search the docs.
The Spicrawl MCP server gives an agent Spicrawl as native tools. It is hosted at:
https://mcp.spicrawl.com/mcpIt speaks MCP Streamable HTTP and authenticates with your own API key as a bearer token. Every tool call runs under that key, exactly as if you had called the API yourself: same scopes, same limits, same credits.
claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
--header "Authorization: Bearer $SPICRAWL_API_KEY"Authentication
Send Authorization: Bearer spicrawl_live_… (or a spicrawl_test_… key) on every request to the endpoint. The key is checked against the API when the MCP session opens:
- An unknown, revoked or expired key is refused with HTTP
401and aWWW-Authenticate: Bearer realm="spicrawl-mcp"header before any session exists. - If the API cannot be reached to check the key, the endpoint answers
503withRetry-After: 5. Retry; do not rotate the key. - A session is bound to the key that opened it.
Keys get the scrape, batch and sessions scopes by default. Two tool groups need a scope you grant deliberately in the dashboard:
| Tools | Scope needed | Without it |
|---|---|---|
spicrawl_usage, spicrawl_usage_summary, spicrawl_usage_reconciliation | read | ERR::AUTH::INSUFFICIENT_SCOPE (HTTP 403) |
spicrawl_browser_connect_url Coming soon | browser | The WebSocket upgrade fails with 403 |
Set up your client
Put the key in the SPICRAWL_API_KEY environment variable and reference it from the config, so the key never lands in a committed file.
One command, stored for you only (local scope):
claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
--header "Authorization: Bearer $SPICRAWL_API_KEY"Add --scope user to use it in every project. To share it with your team, commit a .mcp.json at the project root instead. Claude Code expands ${VAR} in url and headers, so each teammate's own key is used:
{
"mcpServers": {
"spicrawl": {
"type": "http",
"url": "https://mcp.spicrawl.com/mcp",
"headers": { "Authorization": "Bearer ${SPICRAWL_API_KEY}" }
}
}
}Check it with claude mcp list, or /mcp inside a session. More in Claude Code.
Tools
The server exposes 25 tools. Argument names are the API's own except on spicrawl_scrape and spicrawl_batch_submit, which use render for js_render and format for response_format. Tools reject unknown arguments: a misspelt field (for example js_render on spicrawl_scrape) fails the call with an error naming it, instead of being dropped.
Scrape
| Tool | What it does | Key inputs | Use it when |
|---|---|---|---|
spicrawl_scrape | POST /v1/scrape: retrieve one URL as markdown (default), text, html, json or pdf. | url (required), format, render, main_content_only, include_tags, exclude_tags, extract, ai_extract, autoparse, links, screenshot, mode, engine, session_id, wait_for, actions, cache, max_cost | You need one page's content or data now. |
Every other POST /v1/scrape field is accepted under its API name: impersonate, proxy, proxy_verify, method, custom_headers, headless, wait, wait_for_timeout, block_resources, network_capture, screenshot_fullpage, screenshot_selector, screenshot_format, screenshot_quality, parse_pdf, cache_ttl, original_status, allowed_status_codes, extract_preset.
What comes back:
- With
format: "markdown","text"or"html", the tool returns the document itself. Response headers such asX-Target-StatusandX-Credits-Chargedare not passed through. - With
format: "json", or wheneverextract,ai_extract,autoparse,links,network_captureorscreenshotis set, the tool returns the JSON envelope:content, the site'sstatus,credits,engine,warnings,empty_fields, anddatafor extraction. Use it when the agent must check the site's status or the cost. - Screenshots come back as MCP image blocks the model can see. Each
screenshotsentry in the JSON keeps its metadata and points to its block. An image over 5 MB of base64 is not attached; its entry says how to get a smaller one. format: "pdf"prints the page in a browser (setrenderorengine: "chromium") and returns the file as an MCP resource block (application/pdf). The JSON beside it hasengine,status,credits,request_idandpdf.size_bytes. A PDF over 10 MB of base64 is not attached; itspdf.not_attachedsays to callPOST /v1/scrapedirectly.actionsis typed for the tool: each step is{"type": "click" | "fill" | "wait_for" | "wait_for_navigation" | "scroll" | "select" | "evaluate" | "screenshot", ...fields}, with the fields of the browser actions verb plustimeout_ms,on_error,label. The tool sends them in the API's{"click": {...}}shape. Infinite scroll:[{"type":"scroll","to_bottom":true},{"type":"wait_for","ms":1000}]. Ascreenshotstep returns the JSON envelope whateverformatsays.- Error messages and warnings name the tool's arguments: the API's
js_renderreads asrender,response_formatasformat. ai_extractComing soon takespromptorschema, not both, and adds 4 credits. Until AI extraction launches it fails withERR::INTERNAL::UNAVAILABLEat 0 credits; do not retry.
Batch
| Tool | What it does | Key inputs | Use it when |
|---|---|---|---|
spicrawl_batch_submit | POST /v1/batch: queue many URLs as one async job. Returns the job (id, status, estimated_credits, progress). | urls or items (not both; up to 10,000), render, block_resources, name, open, max_attempts, credit_budget | You have tens to thousands of URLs. |
spicrawl_batch_list | List the project's jobs, newest first. | status, limit (max 200), cursor | You lost a job id. |
spicrawl_batch_status | One job's status and progress. | job_id | Polling after submit, until completed, failed or cancelled. |
spicrawl_batch_results | A page of finished items (content, or a result_url API content path per item). | job_id, status, limit (default 500, max 5000), cursor | Reading results; works while the job runs. |
spicrawl_batch_task_content | The full payload of one item by seq (0-based, submission order). | job_id, seq | You need one URL's output. ERR::REQUEST::CONFLICT means not finished yet, or failed. |
spicrawl_batch_cancel | Cancel: unstarted items never run, finished ones keep results. Cannot be undone. | job_id | The job is wrong or too expensive. |
spicrawl_batch_retry | Re-queue every failed item. | job_id | Items failed with retryable errors. |
spicrawl_batch_add_items | Append URLs to a job submitted with open: true. Not idempotent. | job_id, urls or items | You discover URLs while the job runs (crawling). |
spicrawl_batch_close | Stop an open job accepting items so it can complete. | job_id | You are done adding to an open job. An open job never finishes until closed. |
Batch items apply only js_render (render), proxy and block_resources today, and return the raw page. Other settings, including format, are not applied per item. For markdown or extraction, run spicrawl_scrape per URL.
max_cost on a batch caps one item; credit_budget caps the whole job. Results expire after the job's retention window (HTTP 410).
Sessions
| Tool | What it does | Key inputs | Use it when |
|---|---|---|---|
spicrawl_session_create | Create a persistent browser identity: cookies, storage, pinned engine. | engine, ttl_seconds (min 30, default 1800), session_context (to clone) | Multi-step flows: log in, then read pages behind the login. Pass the returned id as session_id to spicrawl_scrape. |
spicrawl_session_list | List sessions, newest first. Never includes credentials. | status, engine, limit, cursor | Finding a session. |
spicrawl_session_get | One session's metadata: status, engine, exit, usage, expiry. | session_id | Checking a session is still active. |
spicrawl_session_context | Dump a live session's cookies and storage. Contains secrets. | session_id | Cloning a session. Do not echo the result to the user or logs. |
spicrawl_session_release | End a session; the record stays as released and its context is purged. | session_id, force | You are done with it. |
spicrawl_session_delete | Permanently delete a session and its record. | session_id, force | Removing it entirely. |
A session runs one request at a time. A second concurrent scrape on it fails with ERR::SESSION::BUSY (retryable).
Usage and request history
| Tool | What it does | Key inputs | Use it when |
|---|---|---|---|
spicrawl_requests_list | Recent API calls in the key's project with status, error code, engine, proxy, latency, credits and project_id. Any key; all_projects: true lists every project in the organization and needs read. | all_projects, only_errors, status, limit (max 200), before + before_id | Finding failing calls or patterns on one site. |
spicrawl_request_get | The full log record of one call. Any key for its own project's calls; another project's need read. | id (the X-Request-Id or request_id of an error) | Debugging one failed or slow scrape. |
spicrawl_usage | Billing-grade usage over a date window. Needs read. | from, to (exclusive, YYYY-MM-DD), group_by (day, project, engine, feature, key: per API key id/name/prefix, per day), metrics (billing metrics plus requests_failed, feature_*, client_*) | "How many credits did I use, and where?" |
spicrawl_usage_summary | Current-period totals. Needs read. | none | A quick "how much have I used". |
spicrawl_usage_reconciliation | One day's drift between the request log and the billed rollup. Needs read. | day (default yesterday) | Auditing billing. |
Browser
Coming soon
spicrawl_browser_connect_url will return a URL for connecting Puppeteer or Playwright to a browser hosted by Spicrawl. Remote browsers are not available yet, so agents should not rely on this tool yet. See CDP browser.
Docs
| Tool | What it does | Key inputs | Use it when |
|---|---|---|---|
spicrawl_docs_search | Full-text search over these docs. Returns up to limit pages, each with its title, matching sections (heading, snippet, URL) and md_url. An error code as the query (ERR::FAMILY::NAME) also returns a link to its entry on Errors. | query (required, up to 200 characters), limit (1-20, default 8) | Before guessing a parameter name or value, and to explain an error code. |
spicrawl_docs_read | One docs page as Markdown. Pages over 60,000 characters are truncated, and the text says so. | path: a page path (guides/anti-bot), a docs URL from spicrawl_docs_search, or either with a #anchor | You know which page answers the question. |
spicrawl_docs_index | The docs' llms.txt: every page with its title, one-line description and .md URL. | none | A search finds nothing, or the agent needs an overview. |
The docs tools read the public docs at https://docs.spicrawl.com and never send your API key there. A self-hosted server reads $SPICRAWL_DOCS_URL, a docs base URL such as http://<host>:8080/docs. When it is unset: the legacy $SPICRAWL_DOCS_HOST plus /docs, then SPICRAWL_PUBLIC_BASE_URL, then SPICRAWL_BASE_URL, plus /docs (a self-hosted API serves its own docs there); with none of them set, or with the public API as the base URL, the public docs site.
How errors reach the agent
A failed call comes back as an MCP tool error (isError: true) whose text is the API's error code, its detail, a retryable note when retryable is true, and the doc_url:
ERR::UPSTREAM::CHALLENGE: <the API's detail text> (retryable — the same request may succeed on a retry)
See https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGEThe agent should act on the code: retry when the message says retryable, and change parameters otherwise. For ERR::UPSTREAM::CHALLENGE, retry once, then add render: true or the user's own proxy. Failed calls cost 0 credits. The full error table is on Errors.
Self-host over stdio
Run the server as a local subprocess of your MCP client instead of using the hosted endpoint. It needs Node.js 20 or later.
Build it
From the mcp/ directory of the Spicrawl source tree (provided to self-hosting customers; there is no public repository):
cd spicrawl/mcp
npm install
npm run buildThis produces dist/index.js (stdio) and dist/http.js (Streamable HTTP).
Add it to your client
{
"mcpServers": {
"spicrawl": {
"command": "node",
"args": ["/absolute/path/to/spicrawl/mcp/dist/index.js"],
"env": {
"SPICRAWL_API_KEY": "spicrawl_live_...",
"SPICRAWL_BASE_URL": "https://api.spicrawl.com"
}
}
}
}For Claude Code: claude mcp add spicrawl --env SPICRAWL_API_KEY=$SPICRAWL_API_KEY --env SPICRAWL_BASE_URL=https://api.spicrawl.com -- node /absolute/path/to/spicrawl/mcp/dist/index.js.
| Variable | Required | Default | Notes |
|---|---|---|---|
SPICRAWL_API_KEY | yes | none | Your API key. |
SPICRAWL_BASE_URL | no | https://api.spicrawl.com | The API the tools call. Set it for a self-hosted API (http://<host>:8080). |
SPICRAWL_PUBLIC_BASE_URL | no | SPICRAWL_BASE_URL | Base used only in URLs handed back to you (spicrawl_browser_connect_url Coming soon). |
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
HTTP 401 when the client connects, or "Unauthorized: send your Spicrawl API key" | No Authorization header, or the key is unknown, revoked or expired. | Check the header is Authorization: Bearer spicrawl_… and that the environment variable is set in the environment the client was started from. Create a new key in the dashboard if it was revoked. |
The header contains a literal ${SPICRAWL_API_KEY} or ${env:SPICRAWL_API_KEY} | The client did not expand the variable (wrong syntax for that client, or the variable is not set when the client starts). | Use ${VAR} in Claude Code, ${env:VAR} in Cursor, ${input:id} in VS Code, bearer_token_env_var in Codex. Start the client from a shell where the variable is set. |
HTTP 503 with Retry-After on connect | The MCP server could not reach the API to check your key. | Retry after the given seconds. Your key is fine. |
ERR::AUTH::INSUFFICIENT_SCOPE from a spicrawl_usage* tool | The key lacks the read scope. | Grant read to the key in the dashboard, or use a key that has it. |
ERR::LIMIT::QUOTA_EXCEEDED (HTTP 402) | Out of credits or over the monthly ceiling. | Stop; top up or wait for the reset. Not retryable. |
| "Unknown argument" or "Unrecognized key" on a tool call | The agent used a field the tool does not accept, often an API name where the tool uses its own (js_render instead of render). | Use the tool's argument name, listed above. |
| Tools do not appear | The client was not restarted, or the config is in the wrong file or key (servers in VS Code, mcpServers elsewhere). | Restart the client and check its MCP panel or logs. |
Check that the endpoint itself is up with curl -i https://mcp.spicrawl.com/mcp. Sent without a key, it answers HTTP 401 with Unauthorized: send your Spicrawl API key, which means the server is reachable; a connection error or a 5xx points at the network or the server, not your client's configuration.
Quickstart
Connect your coding agent to Spicrawl in one command with spicrawl init, or by pasting one self-contained setup prompt into the agent.
Agent skill
Install the Spicrawl agent skill (SKILL.md from https://app.spicrawl.com/skill.md) so a coding agent writes correct calls to the Spicrawl HTTP API.