spicrawlspicrawlDocs

Scripting with the CLI

Recipes for using spicrawl in shell scripts, agent tool calls and CI: jq pipelines, URL files, exit-code branching, retries and per-URL files.

Every recipe here relies on three rules: stdout is data only, output is JSON whenever stdout is not a terminal, and the exit code says what went wrong. Inside $(...) or a pipe you always get JSON, so jq always works.

spicrawl scrape https://example.com/blog/launch --format markdown | jq -r .content

Pipe to jq

# Just the markdown
spicrawl scrape https://example.com/blog/launch --format markdown | jq -r .content

# Credits charged and the engine that served
spicrawl scrape https://example.com/pricing --render | jq '{engine, credits_charged, cache_state}'

# Extracted fields only
spicrawl scrape https://example.com/products/42 \
  --extract '{"title":"h1","price":".price"}' | jq .data

# All links on a page, one per line
spicrawl scrape https://example.com --links | jq -r '.links[]'

With --extract, --ai, --autoparse, --links, --screenshot, --network-capture or --format json, the JSON is the API's envelope (fields content, data, links, credits, engine, ...). Otherwise it is the CLI's wrapper (fields content, status, credits_charged, request_id, ...). See scrape output.

Loop over a file of URLs

Let the CLI do the loop: spicrawl scrape - reads one URL per line, runs them in parallel, and prints one JSON line per URL.

spicrawl scrape - --format markdown --concurrency 8 < urls.txt > pages.jsonl

# Which failed, and why
jq -r 'select(.error) | [.index, .url, .error.code] | @tsv' pages.jsonl

# Retry only the failures, with a stronger engine
jq -r 'select(.error) | .url' pages.jsonl \
  | spicrawl scrape - --format markdown --render > retried.jsonl

Lines arrive in completion order; sort by .index to restore input order: jq -s -c 'sort_by(.index)[]' pages.jsonl. The exit code is 0 only if every URL succeeded.

For thousands of URLs that should keep running after your shell exits, use spicrawl batch instead. Batch returns raw HTML only today.

Save markdown per URL

mkdir -p pages
spicrawl scrape - --format markdown --concurrency 8 < urls.txt \
  | jq -c 'select(.error | not)' \
  | while IFS= read -r line; do
      name=$(printf '%s' "$line" | jq -r '.url | sub("^https?://"; "") | gsub("[^A-Za-z0-9._-]"; "_")')
      printf '%s' "$line" | jq -r .content > "pages/$name.md"
    done

For a handful of URLs, a plain loop with -o is clearer:

while IFS= read -r url; do
  name=$(printf '%s' "$url" | sed -E 's#^https?://##; s#[^A-Za-z0-9._-]#_#g')
  spicrawl scrape "$url" --format markdown -o "pages/$name.md" -q || echo "failed: $url" >&2
done < urls.txt

Branch on exit codes

spicrawl scrape "$url" --format markdown --json -o page.md >/dev/null 2>err.json
case $? in
  0)  ;;                                            # page.md written
  2)  echo "bad flags" >&2; exit 2 ;;
  3)  echo "set SPICRAWL_API_KEY" >&2; exit 3 ;;
  4)  jq -r .detail err.json >&2 ;;                 # fix the request; do not retry
  5)  sleep "$(jq -r '.retry_after_seconds // 30' err.json)" ;;
  6)  spicrawl scrape "$url" --format markdown --render -o page.md ;;
  7|8) echo "retryable: $(jq -r .retryable err.json)" >&2 ;;
  10) echo "cannot reach the API" >&2 ;;
  *)  jq -r '.code // .error' err.json >&2 ;;
esac

--json keeps err.json machine-readable when you run the script from a terminal; without a terminal it is JSON anyway.

Retry wrapper

For transient failures, prefer the built-in --retry N: it retries only errors the API marks retryable, and waits retry_after_seconds when the API gives one (otherwise 1 s, doubling up to 30 s).

spicrawl scrape https://example.com/pricing --render --format markdown --retry 4

To escalate between attempts (a cheaper engine first, a stronger one on a bot challenge), write the loop yourself:

scrape_with_escalation() {
  url=$1
  for flags in "" "--render" "--render --proxy $MY_PROXY_URL"; do
    # shellcheck disable=SC2086
    out=$(spicrawl scrape "$url" --format markdown $flags 2>err.json)
    code=$?
    if [ "$code" -eq 0 ]; then
      printf '%s\n' "$out" | jq -r .content
      return 0
    fi
    case $code in
      6|7|8) continue ;;          # site, proxy or engine: try the next tier
      5) sleep "$(jq -r '.retry_after_seconds // 10' err.json)" ;;
      *) jq -r '.code // .error' err.json >&2; return "$code" ;;
    esac
  done
  return 6
}

scrape_with_escalation https://example.com/products/42 > product.md

--mode auto does the first three rungs of this server-side and bills only the rung that succeeded.

Agent tool calls

An agent that drives a shell should:

  1. Read the key from the environment. Do not run spicrawl login inside the agent.
  2. Use -q to keep stderr free of progress lines, and read stderr only when the exit code is non-zero.
  3. Cap spend with --max-cost N: a request that could cost more fails with ERR::LIMIT::MAX_COST_EXCEEDED (exit 5) at 0 credits.
  4. Pass --yes to destructive commands; without a terminal they refuse with exit 2.
  5. Validate hand-built bodies against spicrawl schema scrape and send them with --body @file.
spicrawl scrape https://example.com/products/42 --format markdown --max-cost 5 -q

CI

Pass the key as a secret environment variable. Nothing is written to disk and nothing prompts.

.github/workflows/snapshot.yml
name: Nightly pricing snapshot
on:
  schedule:
    - cron: "0 3 * * *"
jobs:
  snapshot:
    runs-on: ubuntu-latest
    env:
      SPICRAWL_API_KEY: ${{ secrets.SPICRAWL_API_KEY }}
    steps:
      - uses: actions/checkout@v4
      - name: Install spicrawl
        run: npm install -g @spicrawl/cli
      - name: Check the key
        run: spicrawl auth status
      - name: Scrape
        run: |
          spicrawl scrape - --format markdown --concurrency 4 < urls.txt > pages.jsonl \
            || { jq -c 'select(.error) | {url, code: .error.code}' pages.jsonl; exit 1; }
      - uses: actions/upload-artifact@v4
        with:
          name: pages
          path: pages.jsonl

spicrawl auth status exits 3 on a missing or revoked key, so the job fails before any scraping. spicrawl scrape - exits non-zero if any URL failed; the jq fallback prints each failed URL and its error code to the log before failing the step.

On this page