# Scripting with the CLI

> Recipes for using spicrawl in shell scripts, agent tool calls and CI: jq pipelines, URL files, exit-code branching, retries and per-URL files.

Source: https://docs.spicrawl.com/cli/scripting

Every recipe here relies on three rules: stdout is data only, output is JSON whenever stdout is not a terminal, and the [exit code](https://docs.spicrawl.com/cli/exit-codes.md) says what went wrong. Inside `$(...)` or a pipe you always get JSON, so `jq` always works.

```bash
spicrawl scrape https://example.com/blog/launch --format markdown | jq -r .content
```

## Pipe to jq

```bash
# Just the markdown
spicrawl scrape https://example.com/blog/launch --format markdown | jq -r .content

# Credits charged and the engine that served
spicrawl scrape https://example.com/pricing --render | jq '{engine, credits_charged, cache_state}'

# Extracted fields only
spicrawl scrape https://example.com/products/42 \
  --extract '{"title":"h1","price":".price"}' | jq .data

# All links on a page, one per line
spicrawl scrape https://example.com --links | jq -r '.links[]'
```

With `--extract`, `--ai`, `--autoparse`, `--links`, `--screenshot`, `--network-capture` or `--format json`, the JSON is the API's envelope (fields `content`, `data`, `links`, `credits`, `engine`, ...). Otherwise it is the CLI's wrapper (fields `content`, `status`, `credits_charged`, `request_id`, ...). See [scrape output](https://docs.spicrawl.com/cli/scrape.md#output).

## Loop over a file of URLs

Let the CLI do the loop: `spicrawl scrape -` reads one URL per line, runs them in parallel, and prints one JSON line per URL.

```bash
spicrawl scrape - --format markdown --concurrency 8 < urls.txt > pages.jsonl

# Which failed, and why
jq -r 'select(.error) | [.index, .url, .error.code] | @tsv' pages.jsonl

# Retry only the failures, with a stronger engine
jq -r 'select(.error) | .url' pages.jsonl \
  | spicrawl scrape - --format markdown --render > retried.jsonl
```

Lines arrive in completion order; sort by `.index` to restore input order: `jq -s -c 'sort_by(.index)[]' pages.jsonl`. The exit code is `0` only if every URL succeeded.

For thousands of URLs that should keep running after your shell exits, use [`spicrawl batch`](https://docs.spicrawl.com/cli/batch.md) instead. Batch returns raw HTML only today.

## Save markdown per URL

```bash
mkdir -p pages
spicrawl scrape - --format markdown --concurrency 8 < urls.txt \
  | jq -c 'select(.error | not)' \
  | while IFS= read -r line; do
      name=$(printf '%s' "$line" | jq -r '.url | sub("^https?://"; "") | gsub("[^A-Za-z0-9._-]"; "_")')
      printf '%s' "$line" | jq -r .content > "pages/$name.md"
    done
```

For a handful of URLs, a plain loop with `-o` is clearer:

```bash
while IFS= read -r url; do
  name=$(printf '%s' "$url" | sed -E 's#^https?://##; s#[^A-Za-z0-9._-]#_#g')
  spicrawl scrape "$url" --format markdown -o "pages/$name.md" -q || echo "failed: $url" >&2
done < urls.txt
```

## Branch on exit codes

```bash
spicrawl scrape "$url" --format markdown --json -o page.md >/dev/null 2>err.json
case $? in
  0)  ;;                                            # page.md written
  2)  echo "bad flags" >&2; exit 2 ;;
  3)  echo "set SPICRAWL_API_KEY" >&2; exit 3 ;;
  4)  jq -r .detail err.json >&2 ;;                 # fix the request; do not retry
  5)  sleep "$(jq -r '.retry_after_seconds // 30' err.json)" ;;
  6)  spicrawl scrape "$url" --format markdown --render -o page.md ;;
  7|8) echo "retryable: $(jq -r .retryable err.json)" >&2 ;;
  10) echo "cannot reach the API" >&2 ;;
  *)  jq -r '.code // .error' err.json >&2 ;;
esac
```

`--json` keeps `err.json` machine-readable when you run the script from a terminal; without a terminal it is JSON anyway.

## Retry wrapper

For transient failures, prefer the built-in `--retry N`: it retries only errors the API marks `retryable`, and waits `retry_after_seconds` when the API gives one (otherwise 1 s, doubling up to 30 s).

```bash
spicrawl scrape https://example.com/pricing --render --format markdown --retry 4
```

To escalate between attempts (a cheaper engine first, a stronger one on a bot challenge), write the loop yourself:

```bash
scrape_with_escalation() {
  url=$1
  for flags in "" "--render" "--render --proxy $MY_PROXY_URL"; do
    # shellcheck disable=SC2086
    out=$(spicrawl scrape "$url" --format markdown $flags 2>err.json)
    code=$?
    if [ "$code" -eq 0 ]; then
      printf '%s\n' "$out" | jq -r .content
      return 0
    fi
    case $code in
      6|7|8) continue ;;          # site, proxy or engine: try the next tier
      5) sleep "$(jq -r '.retry_after_seconds // 10' err.json)" ;;
      *) jq -r '.code // .error' err.json >&2; return "$code" ;;
    esac
  done
  return 6
}

scrape_with_escalation https://example.com/products/42 > product.md
```

`--mode auto` does the first three rungs of this server-side and bills only the rung that succeeded.

## Agent tool calls

An agent that drives a shell should:

1. Read the key from the environment. Do not run `spicrawl login` inside the agent.
2. Use `-q` to keep stderr free of progress lines, and read stderr only when the exit code is non-zero.
3. Cap spend with `--max-cost N`: a request that could cost more fails with `ERR::LIMIT::MAX_COST_EXCEEDED` (exit `5`) at 0 credits.
4. Pass `--yes` to destructive commands; without a terminal they refuse with exit `2`.
5. Validate hand-built bodies against `spicrawl schema scrape` and send them with `--body @file`.

```bash
spicrawl scrape https://example.com/products/42 --format markdown --max-cost 5 -q
```

## CI

Pass the key as a secret environment variable. Nothing is written to disk and nothing prompts.

```yaml title=".github/workflows/snapshot.yml"
name: Nightly pricing snapshot
on:
  schedule:
    - cron: "0 3 * * *"
jobs:
  snapshot:
    runs-on: ubuntu-latest
    env:
      SPICRAWL_API_KEY: ${{ secrets.SPICRAWL_API_KEY }}
    steps:
      - uses: actions/checkout@v4
      - name: Install spicrawl
        run: npm install -g @spicrawl/cli
      - name: Check the key
        run: spicrawl auth status
      - name: Scrape
        run: |
          spicrawl scrape - --format markdown --concurrency 4 < urls.txt > pages.jsonl \
            || { jq -c 'select(.error) | {url, code: .error.code}' pages.jsonl; exit 1; }
      - uses: actions/upload-artifact@v4
        with:
          name: pages
          path: pages.jsonl
```

`spicrawl auth status` exits `3` on a missing or revoked key, so the job fails before any scraping. `spicrawl scrape -` exits non-zero if any URL failed; the `jq` fallback prints each failed URL and its error code to the log before failing the step.
