spicrawlspicrawlDocs

Set up AI agents with the CLI

Connect Claude Code, Cursor, VS Code and Codex to Spicrawl with spicrawl init, mcp install and skill install, and give agents docs and request schemas offline with spicrawl docs and spicrawl schema.

Five commands prepare a project for AI coding agents: init does everything, mcp install and skill install do one part each, and docs and schema give an agent reference material from the shell.

export SPICRAWL_API_KEY=spicrawl_live_...
spicrawl init --client claude,cursor --yes

spicrawl init

Detects the agent clients in use (Claude Code, Cursor, VS Code, Codex) in the project directory and your home directory, shows the files it would create or change, and after confirmation runs mcp install and skill install for each.

spicrawl init                                   # detect, show the plan, ask
spicrawl init --client claude,cursor --yes      # no questions
spicrawl init --client all --global --yes       # user-level config for every client
FlagDefaultMeaning
--client LISTdetected clientsclaude, cursor, vscode, codex or all, comma-separated.
--globaloffConfigure user-level client settings instead of the project.
--dir DIRcurrent directoryProject directory.
--mcp-url URL$SPICRAWL_MCP_URL, then https://mcp.spicrawl.com/mcpMCP server URL.
--use-envtrueReference $SPICRAWL_API_KEY instead of embedding the key, where the client supports it.
--agents-mdoffAlso add a Spicrawl section to AGENTS.md.
-y, --yesoffApply without asking. Required without a terminal on stdin.

Without a terminal and without --yes, init prints the plan to stderr and exits 2 with init changes files and stdin is not a terminal: re-run with --yes to apply the plan above. With no client detected and no --client, it exits 2.

Output is one record per file touched:

[
  { "client": "claude", "kind": "mcp", "file": "/home/dev/shop/.mcp.json", "action": "created" },
  { "client": "claude", "kind": "skill", "file": "/home/dev/shop/.claude/skills/spicrawl/SKILL.md", "action": "created" }
]

kind is mcp, skill or agents_md. action is created, updated, unchanged or skipped (with a reason). Human mode prints the same as a CLIENT KIND ACTION FILE NOTE table. Re-running is safe.

spicrawl mcp install

Merges a spicrawl server entry for the hosted MCP server (https://mcp.spicrawl.com/mcp, streamable HTTP) into a client's MCP config. Every other server and key in the file is kept. An entry that is already correct is reported as unchanged.

spicrawl mcp install --client claude
spicrawl mcp install --client all --global
spicrawl mcp install --client cursor --print    # show the snippet, write nothing

Flags: --client, --global, --dir, --mcp-url, --use-env as for init, plus --print to print the config snippet instead of writing it.

Per-client result

By default the key is referenced, not copied into the file.

ClientProject file--global fileHow the key is supplied
Claude Code (claude).mcp.json~/.claude.json${SPICRAWL_API_KEY} from the environment. The user config cannot expand variables, so --global embeds the key and warns.
Cursor (cursor).cursor/mcp.json~/.cursor/mcp.json${env:SPICRAWL_API_KEY} from the environment.
VS Code (vscode).vscode/mcp.json<user config dir>/Code/User/mcp.jsonVS Code prompts once for the key and stores it securely.
Codex (codex).codex/config.toml~/.codex/config.toml (or $CODEX_HOME)bearer_token_env_var = "SPICRAWL_API_KEY".

--use-env=false embeds the literal key in every client's file. Do not commit a file with an embedded key.

{
  "mcpServers": {
    "spicrawl": {
      "type": "http",
      "url": "https://mcp.spicrawl.com/mcp",
      "headers": { "Authorization": "Bearer ${SPICRAWL_API_KEY}" }
    }
  }
}

With --print in JSON mode the output is [{"client", "file", "format", "snippet"}]; "secret": true marks a snippet that contains the literal key. See MCP server for the tools the server exposes.

spicrawl skill install

Downloads the Spicrawl agent skill from https://app.spicrawl.com/skill.md (override with $SPICRAWL_SKILL_URL) and writes it as spicrawl/SKILL.md in each client's skill directory. If the download fails, it prints a warning and uses the copy built into the CLI, so it works offline. That copy is taken when the CLI is released, so it can be older than the served skill.

spicrawl skill install --client claude
spicrawl skill install --client all --global
spicrawl skill install --print > SKILL.md
ClientProject pathAlso reads
Claude Code.claude/skills/spicrawl/SKILL.md
Cursor.cursor/skills/spicrawl/SKILL.md.claude/ and .agents/ skills
VS Code.github/skills/spicrawl/SKILL.md (~/.copilot/skills with --global).claude/ and .agents/ skills
Codex.agents/skills/spicrawl/SKILL.md

A client that already reads a skill written for another selected client is skipped ("action": "skipped", "reason": "already reads this skill"), so an agent never sees the skill twice. --global writes under your home directory instead of the project. --agents-md also adds a short Spicrawl section to AGENTS.md ($CODEX_HOME/AGENTS.md with --global).

Flags: --client, --global, --dir, --agents-md, --print. With --print in JSON mode the output is {"content": "...", "source": "downloaded" | "embedded"}. See agent skill.

spicrawl docs

Fetches a page of these docs as Markdown from docs.spicrawl.com and prints it, so an agent can read the docs without a browser. No API key is needed.

spicrawl docs --list                          # llms.txt: every page with its path
spicrawl docs quickstart
spicrawl docs guides/anti-bot > anti-bot.md
spicrawl docs errors --json | jq -r .markdown

A topic is the page's path without the extension (quickstart, guides/anti-bot, cli/scrape). --list, or no topic, prints the site's llms.txt index. Human mode prints the Markdown as is; JSON mode wraps it as {"topic", "url", "markdown"}. An unknown topic exits 2; an unreachable docs host exits 10. Set $SPICRAWL_DOCS_URL to a docs base URL such as https://docs.example.test (a self-hosted API's docs: http://<host>:8080/docs) to read from another host.

spicrawl schema

Prints the JSON Schema (draft 2020-12) of an API request body, so an agent can validate a body before sending it. The schemas are embedded in the CLI: no network, no key.

spicrawl schema --list
spicrawl schema scrape > scrape.schema.json
spicrawl schema batch-item | jq '.properties | keys'
NameBody
scrapePOST /v1/scrape (what spicrawl scrape --body takes)
batchPOST /v1/batch and POST /v1/batch/{id}/items
batch-itemOne element of a batch's items[], which is one line of a JSONL batch file
sessionPOST /v1/sessions

Every $ref is inlined except the recursive extract-rule definition, which lives under $defs in the same document. The output is the schema itself in both human and JSON mode.

# Validate a hand-written body, then send it
spicrawl schema scrape > scrape.schema.json
npx -y ajv-cli validate --spec=draft2020 -s scrape.schema.json -d body.json \
  && spicrawl scrape https://example.com/products/42 --body @body.json

On this page