spicrawlspicrawlDocs

Spicrawl CLI

The spicrawl command-line client scrapes, renders and extracts web pages, and prints JSON when piped so scripts and AI agents can use it directly.

spicrawl is the command-line client for the Spicrawl API. It wraps every public endpoint (scrape, batch, sessions, request logs, usage; cloud browsers Coming soon) in one static binary, and it is built so an AI agent driving a shell can use it as reliably as a person at a terminal.

export SPICRAWL_API_KEY=spicrawl_live_...
spicrawl scrape https://example.com/blog/launch --format markdown

Five-line tour

spicrawl login --api-key "$SPICRAWL_API_KEY"                                         # save and validate a key
spicrawl scrape https://example.com/pricing --render --wait-for '.plans' --format markdown
spicrawl scrape - --format markdown --concurrency 8 < urls.txt > pages.jsonl      # many URLs, one JSON line each
spicrawl batch submit urls.txt --render --wait                                     # large async job
spicrawl logs --errors                                                             # what failed and why

The agent contract

These rules hold for every command. Scripts and agents can depend on them.

RuleWhat it means
stdout is data onlyThe document, JSON object or JSONL stream goes to stdout. Progress, warnings (warning: ...) and errors go to stderr.
JSON when not a terminalWhen stdout is a pipe or file, output is JSON with no flag. --json forces JSON on a terminal. On a terminal, commands print tables or the document itself.
Never prompts without a TTYAnything interactive has a flag form. spicrawl login without a key and without a terminal fails with exit code 2 instead of waiting.
--yes on destructive commandsspicrawl sessions delete asks for confirmation on a terminal. Without a terminal it refuses with exit code 2 unless you pass --yes (-y).
- reads stdinCommands that take URLs accept - to read one URL per line. Blank lines and lines starting with # are skipped.
Streams are JSONLCommands that return many records (scrape -, batch results) print one compact JSON object per line.
Stable exit codes0 success, 2 usage, 3 auth, 4 bad request, 5 limit, 6 upstream site, 7 proxy, 8 engine or extraction, 9 session, 10 network, 11 a wait gave up. See exit codes.
Errors are the API's problem documentIn JSON mode a failed command writes the API's RFC 7807 body (with code, retryable, request_id) to stderr unchanged.

A minimal agent loop needs nothing else:

if out=$(spicrawl scrape https://example.com/products/42 --format markdown 2>err.json); then
  printf '%s\n' "$out" | jq -r .content
else
  code=$?
  jq -r '.code, .detail' err.json   # e.g. ERR::UPSTREAM::CHALLENGE
  exit "$code"
fi

Global flags

Every command accepts these flags.

FlagDefaultMeaning
--jsonoff (on when stdout is not a terminal)Force JSON output.
-q, --quietoffSuppress progress messages on stderr. Warnings and errors still print.
--api-key KEY$SPICRAWL_API_KEY, then the config fileAPI key for this call.
--base-url URL$SPICRAWL_BASE_URL, then the config file, then https://api.spicrawl.comAPI base URL.
--timeout DURATION3m0sHTTP timeout per API call, as a Go duration (90s, 5m).

Commands

CommandAPIPage
spicrawl login, logout, auth status, confignone (GET /v1/requests?limit=1 to validate)Authentication
spicrawl scrape <url|->POST /v1/scrapeScrape
spicrawl batch submit|list|get|results|wait|cancel|retry|close|append/v1/batchBatch
spicrawl sessions create|list|get|context|release|delete/v1/sessionsSessions and browser
spicrawl browser urlComing soon Builds the GET /v1/browser WebSocket URLSessions and browser
spicrawl logs, logs get, usage, status/v1/requests, /v1/usage, /readyz, /v1/workersLogs and usage
spicrawl init, mcp install, skill install, docs, schemanoneAgent setup
spicrawl exit-codes, spicrawl versionnoneExit codes

Run spicrawl <command> --help for the flags and examples of any command.

Next steps

On this page