Spicrawl CLI
The spicrawl command-line client scrapes, renders and extracts web pages, and prints JSON when piped so scripts and AI agents can use it directly.
spicrawl is the command-line client for the Spicrawl API. It wraps every public endpoint (scrape, batch, sessions, request logs, usage; cloud browsers Coming soon) in one static binary, and it is built so an AI agent driving a shell can use it as reliably as a person at a terminal.
export SPICRAWL_API_KEY=spicrawl_live_...
spicrawl scrape https://example.com/blog/launch --format markdownFive-line tour
spicrawl login --api-key "$SPICRAWL_API_KEY" # save and validate a key
spicrawl scrape https://example.com/pricing --render --wait-for '.plans' --format markdown
spicrawl scrape - --format markdown --concurrency 8 < urls.txt > pages.jsonl # many URLs, one JSON line each
spicrawl batch submit urls.txt --render --wait # large async job
spicrawl logs --errors # what failed and whyThe agent contract
These rules hold for every command. Scripts and agents can depend on them.
| Rule | What it means |
|---|---|
| stdout is data only | The document, JSON object or JSONL stream goes to stdout. Progress, warnings (warning: ...) and errors go to stderr. |
| JSON when not a terminal | When stdout is a pipe or file, output is JSON with no flag. --json forces JSON on a terminal. On a terminal, commands print tables or the document itself. |
| Never prompts without a TTY | Anything interactive has a flag form. spicrawl login without a key and without a terminal fails with exit code 2 instead of waiting. |
--yes on destructive commands | spicrawl sessions delete asks for confirmation on a terminal. Without a terminal it refuses with exit code 2 unless you pass --yes (-y). |
- reads stdin | Commands that take URLs accept - to read one URL per line. Blank lines and lines starting with # are skipped. |
| Streams are JSONL | Commands that return many records (scrape -, batch results) print one compact JSON object per line. |
| Stable exit codes | 0 success, 2 usage, 3 auth, 4 bad request, 5 limit, 6 upstream site, 7 proxy, 8 engine or extraction, 9 session, 10 network, 11 a wait gave up. See exit codes. |
| Errors are the API's problem document | In JSON mode a failed command writes the API's RFC 7807 body (with code, retryable, request_id) to stderr unchanged. |
A minimal agent loop needs nothing else:
if out=$(spicrawl scrape https://example.com/products/42 --format markdown 2>err.json); then
printf '%s\n' "$out" | jq -r .content
else
code=$?
jq -r '.code, .detail' err.json # e.g. ERR::UPSTREAM::CHALLENGE
exit "$code"
fiGlobal flags
Every command accepts these flags.
| Flag | Default | Meaning |
|---|---|---|
--json | off (on when stdout is not a terminal) | Force JSON output. |
-q, --quiet | off | Suppress progress messages on stderr. Warnings and errors still print. |
--api-key KEY | $SPICRAWL_API_KEY, then the config file | API key for this call. |
--base-url URL | $SPICRAWL_BASE_URL, then the config file, then https://api.spicrawl.com | API base URL. |
--timeout DURATION | 3m0s | HTTP timeout per API call, as a Go duration (90s, 5m). |
Commands
| Command | API | Page |
|---|---|---|
spicrawl login, logout, auth status, config | none (GET /v1/requests?limit=1 to validate) | Authentication |
spicrawl scrape <url|-> | POST /v1/scrape | Scrape |
spicrawl batch submit|list|get|results|wait|cancel|retry|close|append | /v1/batch | Batch |
spicrawl sessions create|list|get|context|release|delete | /v1/sessions | Sessions and browser |
spicrawl browser url | Coming soon Builds the GET /v1/browser WebSocket URL | Sessions and browser |
spicrawl logs, logs get, usage, status | /v1/requests, /v1/usage, /readyz, /v1/workers | Logs and usage |
spicrawl init, mcp install, skill install, docs, schema | none | Agent setup |
spicrawl exit-codes, spicrawl version | none | Exit codes |
Run spicrawl <command> --help for the flags and examples of any command.