Scripting with the CLI
Recipes for using spicrawl in shell scripts, agent tool calls and CI: jq pipelines, URL files, exit-code branching, retries and per-URL files.
Every recipe here relies on three rules: stdout is data only, output is JSON whenever stdout is not a terminal, and the exit code says what went wrong. Inside $(...) or a pipe you always get JSON, so jq always works.
spicrawl scrape https://example.com/blog/launch --format markdown | jq -r .contentPipe to jq
# Just the markdown
spicrawl scrape https://example.com/blog/launch --format markdown | jq -r .content
# Credits charged and the engine that served
spicrawl scrape https://example.com/pricing --render | jq '{engine, credits_charged, cache_state}'
# Extracted fields only
spicrawl scrape https://example.com/products/42 \
--extract '{"title":"h1","price":".price"}' | jq .data
# All links on a page, one per line
spicrawl scrape https://example.com --links | jq -r '.links[]'With --extract, --ai, --autoparse, --links, --screenshot, --network-capture or --format json, the JSON is the API's envelope (fields content, data, links, credits, engine, ...). Otherwise it is the CLI's wrapper (fields content, status, credits_charged, request_id, ...). See scrape output.
Loop over a file of URLs
Let the CLI do the loop: spicrawl scrape - reads one URL per line, runs them in parallel, and prints one JSON line per URL.
spicrawl scrape - --format markdown --concurrency 8 < urls.txt > pages.jsonl
# Which failed, and why
jq -r 'select(.error) | [.index, .url, .error.code] | @tsv' pages.jsonl
# Retry only the failures, with a stronger engine
jq -r 'select(.error) | .url' pages.jsonl \
| spicrawl scrape - --format markdown --render > retried.jsonlLines arrive in completion order; sort by .index to restore input order: jq -s -c 'sort_by(.index)[]' pages.jsonl. The exit code is 0 only if every URL succeeded.
For thousands of URLs that should keep running after your shell exits, use spicrawl batch instead. Batch returns raw HTML only today.
Save markdown per URL
mkdir -p pages
spicrawl scrape - --format markdown --concurrency 8 < urls.txt \
| jq -c 'select(.error | not)' \
| while IFS= read -r line; do
name=$(printf '%s' "$line" | jq -r '.url | sub("^https?://"; "") | gsub("[^A-Za-z0-9._-]"; "_")')
printf '%s' "$line" | jq -r .content > "pages/$name.md"
doneFor a handful of URLs, a plain loop with -o is clearer:
while IFS= read -r url; do
name=$(printf '%s' "$url" | sed -E 's#^https?://##; s#[^A-Za-z0-9._-]#_#g')
spicrawl scrape "$url" --format markdown -o "pages/$name.md" -q || echo "failed: $url" >&2
done < urls.txtBranch on exit codes
spicrawl scrape "$url" --format markdown --json -o page.md >/dev/null 2>err.json
case $? in
0) ;; # page.md written
2) echo "bad flags" >&2; exit 2 ;;
3) echo "set SPICRAWL_API_KEY" >&2; exit 3 ;;
4) jq -r .detail err.json >&2 ;; # fix the request; do not retry
5) sleep "$(jq -r '.retry_after_seconds // 30' err.json)" ;;
6) spicrawl scrape "$url" --format markdown --render -o page.md ;;
7|8) echo "retryable: $(jq -r .retryable err.json)" >&2 ;;
10) echo "cannot reach the API" >&2 ;;
*) jq -r '.code // .error' err.json >&2 ;;
esac--json keeps err.json machine-readable when you run the script from a terminal; without a terminal it is JSON anyway.
Retry wrapper
For transient failures, prefer the built-in --retry N: it retries only errors the API marks retryable, and waits retry_after_seconds when the API gives one (otherwise 1 s, doubling up to 30 s).
spicrawl scrape https://example.com/pricing --render --format markdown --retry 4To escalate between attempts (a cheaper engine first, a stronger one on a bot challenge), write the loop yourself:
scrape_with_escalation() {
url=$1
for flags in "" "--render" "--render --proxy $MY_PROXY_URL"; do
# shellcheck disable=SC2086
out=$(spicrawl scrape "$url" --format markdown $flags 2>err.json)
code=$?
if [ "$code" -eq 0 ]; then
printf '%s\n' "$out" | jq -r .content
return 0
fi
case $code in
6|7|8) continue ;; # site, proxy or engine: try the next tier
5) sleep "$(jq -r '.retry_after_seconds // 10' err.json)" ;;
*) jq -r '.code // .error' err.json >&2; return "$code" ;;
esac
done
return 6
}
scrape_with_escalation https://example.com/products/42 > product.md--mode auto does the first three rungs of this server-side and bills only the rung that succeeded.
Agent tool calls
An agent that drives a shell should:
- Read the key from the environment. Do not run
spicrawl logininside the agent. - Use
-qto keep stderr free of progress lines, and read stderr only when the exit code is non-zero. - Cap spend with
--max-cost N: a request that could cost more fails withERR::LIMIT::MAX_COST_EXCEEDED(exit5) at 0 credits. - Pass
--yesto destructive commands; without a terminal they refuse with exit2. - Validate hand-built bodies against
spicrawl schema scrapeand send them with--body @file.
spicrawl scrape https://example.com/products/42 --format markdown --max-cost 5 -qCI
Pass the key as a secret environment variable. Nothing is written to disk and nothing prompts.
name: Nightly pricing snapshot
on:
schedule:
- cron: "0 3 * * *"
jobs:
snapshot:
runs-on: ubuntu-latest
env:
SPICRAWL_API_KEY: ${{ secrets.SPICRAWL_API_KEY }}
steps:
- uses: actions/checkout@v4
- name: Install spicrawl
run: npm install -g @spicrawl/cli
- name: Check the key
run: spicrawl auth status
- name: Scrape
run: |
spicrawl scrape - --format markdown --concurrency 4 < urls.txt > pages.jsonl \
|| { jq -c 'select(.error) | {url, code: .error.code}' pages.jsonl; exit 1; }
- uses: actions/upload-artifact@v4
with:
name: pages
path: pages.jsonlspicrawl auth status exits 3 on a missing or revoked key, so the job fails before any scraping. spicrawl scrape - exits non-zero if any URL failed; the jq fallback prints each failed URL and its error code to the log before failing the step.
Agent setup
Connect Claude Code, Cursor, VS Code and Codex to Spicrawl with spicrawl init, mcp install and skill install, and give agents docs and request schemas offline with spicrawl docs and spicrawl schema.
Exit codes
The stable exit codes of the spicrawl CLI, which API error codes map to each, and the JSON error written to stderr on failure.