Build with agents
The four ways an AI agent can use Spicrawl (MCP server, agent skill, CLI, raw HTTP), when to pick each, and how agents read these docs.
An agent can reach Spicrawl four ways. All four call the same API at https://api.spicrawl.com with the same key (SPICRAWL_API_KEY, formats spicrawl_live_… and spicrawl_test_…), so they return the same data, errors and credit charges. Pick by where your agent runs and what it can already do.
The fastest setup for a coding agent is one command from your project directory:
spicrawl initIt detects Claude Code, Cursor, VS Code and Codex, adds the hosted MCP server to their config and writes the agent skill. See Agent quickstart for the command and for a copy-paste prompt that does the same without the CLI. Every other agent has a setup page: see Set up your agent.
The four ways
| MCP server | Agent skill | CLI | Raw HTTP | |
|---|---|---|---|---|
| What the agent gets | 25 spicrawl_* tools it calls directly | Instructions (SKILL.md) that teach it the HTTP API | The spicrawl binary in a shell | Nothing: you write the calls |
| Where it runs | Any MCP client that supports Streamable HTTP and custom headers: Claude Code, Cursor, VS Code, Codex, OpenCode, Gemini CLI, Devin and more | Agents that load skills: Claude Code, Cursor, Codex, GitHub Copilot, Gemini CLI, OpenCode, Devin and more | Agents with a shell tool | Your own agent code |
| Setup | Add https://mcp.spicrawl.com/mcp with an Authorization: Bearer header | Save https://app.spicrawl.com/skill.md as a skill | Install spicrawl, set SPICRAWL_API_KEY | None |
| Output the agent sees | JSON tool results; screenshots as image blocks | Whatever its HTTP calls return | Document or JSON on stdout, errors on stderr, a stable exit code | Full HTTP response, including headers |
Response headers (X-Target-Status, X-Credits-Charged) | Not exposed. Use format: "json" to get the envelope's status and credits | Yes | In the JSON output when piped (status, credits_charged); --meta on a terminal | Yes |
| Best for | Chat and coding agents that should scrape with no code | Agents that write code which calls Spicrawl | Scripts, CI, agents that prefer a shell | Agents you build yourself, production pipelines |
You can combine them. A common setup is the MCP server for ad-hoc scraping in the chat plus the skill so the agent writes correct HTTP code when it adds Spicrawl to your application.
When to pick each
Pick the MCP server when you want the agent to fetch pages during a conversation: "read this pricing page", "get the product name and price from these 30 URLs". The agent calls spicrawl_scrape, spicrawl_batch_submit, spicrawl_session_create and the rest as native tools and never writes an HTTP request.
The hosted server is https://mcp.spicrawl.com/mcp (MCP Streamable HTTP). Each call runs under the API key you send as Authorization: Bearer $SPICRAWL_API_KEY. See MCP server.
Pick the skill when the agent's job is to write software that calls Spicrawl: adding a scraper to your backend, a data pipeline, a script. The skill is a Markdown file with the base URL, auth, the request fields, the error codes and a cost playbook, so the code it writes uses real field names (response_format, js_render, custom_headers) instead of guessed ones.
The skill is served at https://app.spicrawl.com/skill.md. See Agent skill.
Pick the CLI when the agent already works in a terminal (Claude Code, Codex, Cursor's agent) and you want scraping without an MCP server, or in CI. spicrawl prints JSON when stdout is not a terminal, never prompts without a TTY, writes errors to stderr as the API's problem document, and exits with a stable code: 0 ok, 2 usage, 3 auth, 4 request, 5 limit, 6 upstream site, 7 proxy, 8 engine or extraction, 9 session, 10 network, 11 a wait gave up.
spicrawl scrape https://example.com/products/42 --format markdown | jq -r .contentSee CLI overview and Exit codes.
Pick raw HTTP when you are building your own agent and expose Spicrawl as one of its tools. You control retries, escalation and cost, and you see every response header. Best practices has a complete fetch_page(url) tool in Python and TypeScript.
curl -sS -X POST https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/products/42", "response_format": "markdown"}'Set up your agent
Every agent below connects to the same hosted MCP server, https://mcp.spicrawl.com/mcp, with Authorization: Bearer $SPICRAWL_API_KEY. spicrawl init configures the four agents marked "Yes"; every other page gives the steps by hand.
| Agent | Setup page | spicrawl init |
|---|---|---|
| Claude Code | Claude Code | Yes |
| Cursor (editor, CLI and cloud agents) | Cursor | Yes |
| Codex (CLI and cloud) | Codex | Yes |
| GitHub Copilot (VS Code, Copilot CLI, cloud agent) | GitHub Copilot and VS Code | Yes (VS Code) |
| OpenCode | OpenCode | No |
| Gemini CLI, Qwen Code | Gemini CLI | No |
| Windsurf (Devin Desktop) | Windsurf | No |
| Cline | Cline | No |
| Kilo Code | Kilo Code | No |
| Zed | Zed | No |
| JetBrains AI Assistant, Junie | JetBrains | No |
| Goose | Goose | No |
| Kiro | Kiro | No |
| Factory Droid | Factory Droid | No |
| Crush, Amp, Kimi Code, Warp, Continue, Augment, Amazon Q, Antigravity | More MCP clients | No |
| Devin | Devin | No |
| OpenHands | OpenHands | No |
| Replit Agent, v0, Lovable, Bolt | AI app builders | No |
| Agent frameworks (your own code) | Agent frameworks | No |
How agents read these docs
Every page on https://docs.spicrawl.com is also available as Markdown, so an agent can read the docs without parsing HTML:
| Source | URL or command | Use it for |
|---|---|---|
| Page index | https://docs.spicrawl.com/llms.txt | Finding the right page: one line per page with its description |
| Full text | https://docs.spicrawl.com/llms-full.txt | Loading all docs into context at once |
| One page as Markdown | Append .md, e.g. https://docs.spicrawl.com/agents/best-practices.md | Reading one page |
| Docs tools in the MCP server | spicrawl_docs_search, spicrawl_docs_read, spicrawl_docs_index on https://mcp.spicrawl.com/mcp | Searching and reading the docs as tools from an MCP client |
| CLI | spicrawl docs agents/best-practices | Printing a page in a terminal; spicrawl docs --list prints the index |
See Docs for agents for details.
Rules every agent should follow
These hold for all four ways:
- Ask for
response_format: "markdown". The API default ishtml. - The HTTP status is Spicrawl's; the site's status is
X-Target-Status(orstatusin the JSON envelope). A site that returned403can still arrive inside an HTTP200. - Switch on the error
code(ERR::FAMILY::NAME) and retry only whenretryableistrue, afterretry_after_seconds. - Failed requests cost 0 credits. Retrying a successful
POST /v1/scrapebills a second scrape, because the API is not idempotent. - Set
max_costso a request that would cost more is refused before it runs (ERR::LIMIT::MAX_COST_EXCEEDED, HTTP 400, 0 credits). - Use
POST /v1/batchfor more than about 20 URLs.