spicrawlspicrawlDocs

Build with agents

The four ways an AI agent can use Spicrawl (MCP server, agent skill, CLI, raw HTTP), when to pick each, and how agents read these docs.

An agent can reach Spicrawl four ways. All four call the same API at https://api.spicrawl.com with the same key (SPICRAWL_API_KEY, formats spicrawl_live_… and spicrawl_test_…), so they return the same data, errors and credit charges. Pick by where your agent runs and what it can already do.

The fastest setup for a coding agent is one command from your project directory:

spicrawl init

It detects Claude Code, Cursor, VS Code and Codex, adds the hosted MCP server to their config and writes the agent skill. See Agent quickstart for the command and for a copy-paste prompt that does the same without the CLI. Every other agent has a setup page: see Set up your agent.

The four ways

MCP serverAgent skillCLIRaw HTTP
What the agent gets25 spicrawl_* tools it calls directlyInstructions (SKILL.md) that teach it the HTTP APIThe spicrawl binary in a shellNothing: you write the calls
Where it runsAny MCP client that supports Streamable HTTP and custom headers: Claude Code, Cursor, VS Code, Codex, OpenCode, Gemini CLI, Devin and moreAgents that load skills: Claude Code, Cursor, Codex, GitHub Copilot, Gemini CLI, OpenCode, Devin and moreAgents with a shell toolYour own agent code
SetupAdd https://mcp.spicrawl.com/mcp with an Authorization: Bearer headerSave https://app.spicrawl.com/skill.md as a skillInstall spicrawl, set SPICRAWL_API_KEYNone
Output the agent seesJSON tool results; screenshots as image blocksWhatever its HTTP calls returnDocument or JSON on stdout, errors on stderr, a stable exit codeFull HTTP response, including headers
Response headers (X-Target-Status, X-Credits-Charged)Not exposed. Use format: "json" to get the envelope's status and creditsYesIn the JSON output when piped (status, credits_charged); --meta on a terminalYes
Best forChat and coding agents that should scrape with no codeAgents that write code which calls SpicrawlScripts, CI, agents that prefer a shellAgents you build yourself, production pipelines

You can combine them. A common setup is the MCP server for ad-hoc scraping in the chat plus the skill so the agent writes correct HTTP code when it adds Spicrawl to your application.

When to pick each

Set up your agent

Every agent below connects to the same hosted MCP server, https://mcp.spicrawl.com/mcp, with Authorization: Bearer $SPICRAWL_API_KEY. spicrawl init configures the four agents marked "Yes"; every other page gives the steps by hand.

AgentSetup pagespicrawl init
Claude CodeClaude CodeYes
Cursor (editor, CLI and cloud agents)CursorYes
Codex (CLI and cloud)CodexYes
GitHub Copilot (VS Code, Copilot CLI, cloud agent)GitHub Copilot and VS CodeYes (VS Code)
OpenCodeOpenCodeNo
Gemini CLI, Qwen CodeGemini CLINo
Windsurf (Devin Desktop)WindsurfNo
ClineClineNo
Kilo CodeKilo CodeNo
ZedZedNo
JetBrains AI Assistant, JunieJetBrainsNo
GooseGooseNo
KiroKiroNo
Factory DroidFactory DroidNo
Crush, Amp, Kimi Code, Warp, Continue, Augment, Amazon Q, AntigravityMore MCP clientsNo
DevinDevinNo
OpenHandsOpenHandsNo
Replit Agent, v0, Lovable, BoltAI app buildersNo
Agent frameworks (your own code)Agent frameworksNo

How agents read these docs

Every page on https://docs.spicrawl.com is also available as Markdown, so an agent can read the docs without parsing HTML:

SourceURL or commandUse it for
Page indexhttps://docs.spicrawl.com/llms.txtFinding the right page: one line per page with its description
Full texthttps://docs.spicrawl.com/llms-full.txtLoading all docs into context at once
One page as MarkdownAppend .md, e.g. https://docs.spicrawl.com/agents/best-practices.mdReading one page
Docs tools in the MCP serverspicrawl_docs_search, spicrawl_docs_read, spicrawl_docs_index on https://mcp.spicrawl.com/mcpSearching and reading the docs as tools from an MCP client
CLIspicrawl docs agents/best-practicesPrinting a page in a terminal; spicrawl docs --list prints the index

See Docs for agents for details.

Rules every agent should follow

These hold for all four ways:

  • Ask for response_format: "markdown". The API default is html.
  • The HTTP status is Spicrawl's; the site's status is X-Target-Status (or status in the JSON envelope). A site that returned 403 can still arrive inside an HTTP 200.
  • Switch on the error code (ERR::FAMILY::NAME) and retry only when retryable is true, after retry_after_seconds.
  • Failed requests cost 0 credits. Retrying a successful POST /v1/scrape bills a second scrape, because the API is not idempotent.
  • Set max_cost so a request that would cost more is refused before it runs (ERR::LIMIT::MAX_COST_EXCEEDED, HTTP 400, 0 credits).
  • Use POST /v1/batch for more than about 20 URLs.

Next steps

On this page