Set up AI agents with the CLI
Connect Claude Code, Cursor, VS Code and Codex to Spicrawl with spicrawl init, mcp install and skill install, and give agents docs and request schemas offline with spicrawl docs and spicrawl schema.
Five commands prepare a project for AI coding agents: init does everything, mcp install and skill install do one part each, and docs and schema give an agent reference material from the shell.
export SPICRAWL_API_KEY=spicrawl_live_...
spicrawl init --client claude,cursor --yesspicrawl init
Detects the agent clients in use (Claude Code, Cursor, VS Code, Codex) in the project directory and your home directory, shows the files it would create or change, and after confirmation runs mcp install and skill install for each.
spicrawl init # detect, show the plan, ask
spicrawl init --client claude,cursor --yes # no questions
spicrawl init --client all --global --yes # user-level config for every client| Flag | Default | Meaning |
|---|---|---|
--client LIST | detected clients | claude, cursor, vscode, codex or all, comma-separated. |
--global | off | Configure user-level client settings instead of the project. |
--dir DIR | current directory | Project directory. |
--mcp-url URL | $SPICRAWL_MCP_URL, then https://mcp.spicrawl.com/mcp | MCP server URL. |
--use-env | true | Reference $SPICRAWL_API_KEY instead of embedding the key, where the client supports it. |
--agents-md | off | Also add a Spicrawl section to AGENTS.md. |
-y, --yes | off | Apply without asking. Required without a terminal on stdin. |
Without a terminal and without --yes, init prints the plan to stderr and exits 2 with init changes files and stdin is not a terminal: re-run with --yes to apply the plan above. With no client detected and no --client, it exits 2.
Output is one record per file touched:
[
{ "client": "claude", "kind": "mcp", "file": "/home/dev/shop/.mcp.json", "action": "created" },
{ "client": "claude", "kind": "skill", "file": "/home/dev/shop/.claude/skills/spicrawl/SKILL.md", "action": "created" }
]kind is mcp, skill or agents_md. action is created, updated, unchanged or skipped (with a reason). Human mode prints the same as a CLIENT KIND ACTION FILE NOTE table. Re-running is safe.
spicrawl mcp install
Merges a spicrawl server entry for the hosted MCP server (https://mcp.spicrawl.com/mcp, streamable HTTP) into a client's MCP config. Every other server and key in the file is kept. An entry that is already correct is reported as unchanged.
spicrawl mcp install --client claude
spicrawl mcp install --client all --global
spicrawl mcp install --client cursor --print # show the snippet, write nothingFlags: --client, --global, --dir, --mcp-url, --use-env as for init, plus --print to print the config snippet instead of writing it.
Per-client result
By default the key is referenced, not copied into the file.
| Client | Project file | --global file | How the key is supplied |
|---|---|---|---|
Claude Code (claude) | .mcp.json | ~/.claude.json | ${SPICRAWL_API_KEY} from the environment. The user config cannot expand variables, so --global embeds the key and warns. |
Cursor (cursor) | .cursor/mcp.json | ~/.cursor/mcp.json | ${env:SPICRAWL_API_KEY} from the environment. |
VS Code (vscode) | .vscode/mcp.json | <user config dir>/Code/User/mcp.json | VS Code prompts once for the key and stores it securely. |
Codex (codex) | .codex/config.toml | ~/.codex/config.toml (or $CODEX_HOME) | bearer_token_env_var = "SPICRAWL_API_KEY". |
--use-env=false embeds the literal key in every client's file. Do not commit a file with an embedded key.
{
"mcpServers": {
"spicrawl": {
"type": "http",
"url": "https://mcp.spicrawl.com/mcp",
"headers": { "Authorization": "Bearer ${SPICRAWL_API_KEY}" }
}
}
}With --print in JSON mode the output is [{"client", "file", "format", "snippet"}]; "secret": true marks a snippet that contains the literal key. See MCP server for the tools the server exposes.
spicrawl skill install
Downloads the Spicrawl agent skill from https://app.spicrawl.com/skill.md (override with $SPICRAWL_SKILL_URL) and writes it as spicrawl/SKILL.md in each client's skill directory. If the download fails, it prints a warning and uses the copy built into the CLI, so it works offline. That copy is taken when the CLI is released, so it can be older than the served skill.
spicrawl skill install --client claude
spicrawl skill install --client all --global
spicrawl skill install --print > SKILL.md| Client | Project path | Also reads |
|---|---|---|
| Claude Code | .claude/skills/spicrawl/SKILL.md | |
| Cursor | .cursor/skills/spicrawl/SKILL.md | .claude/ and .agents/ skills |
| VS Code | .github/skills/spicrawl/SKILL.md (~/.copilot/skills with --global) | .claude/ and .agents/ skills |
| Codex | .agents/skills/spicrawl/SKILL.md |
A client that already reads a skill written for another selected client is skipped ("action": "skipped", "reason": "already reads this skill"), so an agent never sees the skill twice. --global writes under your home directory instead of the project. --agents-md also adds a short Spicrawl section to AGENTS.md ($CODEX_HOME/AGENTS.md with --global).
Flags: --client, --global, --dir, --agents-md, --print. With --print in JSON mode the output is {"content": "...", "source": "downloaded" | "embedded"}. See agent skill.
spicrawl docs
Fetches a page of these docs as Markdown from docs.spicrawl.com and prints it, so an agent can read the docs without a browser. No API key is needed.
spicrawl docs --list # llms.txt: every page with its path
spicrawl docs quickstart
spicrawl docs guides/anti-bot > anti-bot.md
spicrawl docs errors --json | jq -r .markdownA topic is the page's path without the extension (quickstart, guides/anti-bot, cli/scrape). --list, or no topic, prints the site's llms.txt index. Human mode prints the Markdown as is; JSON mode wraps it as {"topic", "url", "markdown"}. An unknown topic exits 2; an unreachable docs host exits 10. Set $SPICRAWL_DOCS_URL to a docs base URL such as https://docs.example.test (a self-hosted API's docs: http://<host>:8080/docs) to read from another host.
spicrawl schema
Prints the JSON Schema (draft 2020-12) of an API request body, so an agent can validate a body before sending it. The schemas are embedded in the CLI: no network, no key.
spicrawl schema --list
spicrawl schema scrape > scrape.schema.json
spicrawl schema batch-item | jq '.properties | keys'| Name | Body |
|---|---|
scrape | POST /v1/scrape (what spicrawl scrape --body takes) |
batch | POST /v1/batch and POST /v1/batch/{id}/items |
batch-item | One element of a batch's items[], which is one line of a JSONL batch file |
session | POST /v1/sessions |
Every $ref is inlined except the recursive extract-rule definition, which lives under $defs in the same document. The output is the schema itself in both human and JSON mode.
# Validate a hand-written body, then send it
spicrawl schema scrape > scrape.schema.json
npx -y ajv-cli validate --spec=draft2020 -s scrape.schema.json -d body.json \
&& spicrawl scrape https://example.com/products/42 --body @body.json