spicrawlspicrawlDocs

CLI exit codes

The stable exit codes of the spicrawl CLI, which API error codes map to each, and the JSON error written to stderr on failure.

Every spicrawl command exits with one of these codes. They are stable: numbers are never reassigned, only added. Scripts and agents can branch on them without parsing any text. spicrawl exit-codes prints the same table.

CodeMeaningTypical next step
0success
1internal error (CLI bug or ERR::INTERNAL::*)Retry once; if it repeats, report it with the request_id.
2usage error: bad flags or arguments, nothing was sentFix the command. No request was made and no credits were spent.
3ERR::AUTH::*, or no API key configuredSet SPICRAWL_API_KEY or run spicrawl login; check the key's scopes.
4ERR::REQUEST::*, ERR::SECURITY::*, ERR::EXTRACT::INVALID_RULES: fix the requestRead detail and change the flags. Retrying the same call fails the same way.
5ERR::LIMIT::*: rate, concurrency, quota or max_costWait retry_after_seconds for rate and concurrency; for quota, stop; for max_cost, raise --max-cost or use a cheaper engine.
6ERR::UPSTREAM::*: the target site failed or served a bot challengeRetry if retryable; for ERR::UPSTREAM::CHALLENGE add --render, then your own --proxy.
7ERR::PROXY::*Retry if retryable; check your --proxy.
8ERR::ENGINE::*, ERR::EXTRACT::FAILEDRetry if retryable; for extraction, change the prompt or schema.
9ERR::SESSION::*Create a new session, or wait for a busy one.
10network: the Spicrawl API could not be reachedCheck connectivity and --base-url; retry.
11a --wait gave up before the job finishedThe job is still running. Run spicrawl batch wait <id> again.

The mapping uses the second segment of the API's code, so a code added to the API later still lands in the right bucket (a new ERR::PROXY::* code exits 7). Any other ERR::EXTRACT::* code besides INVALID_RULES exits 8. See errors for every API code.

What is written on failure

The error goes to stderr, in the same mode as the output would have been. stdout holds no partial result, with three exceptions: spicrawl scrape - keeps printing the lines of URLs that did succeed; spicrawl auth status and spicrawl status print their report even when they exit non-zero; and spicrawl batch wait (or batch submit --wait) prints the last job it saw when it exits 11.

JSON mode (stdout not a terminal, or --json)

An API error writes the API's RFC 7807 problem document to stderr unchanged:

{
  "type": "https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE",
  "title": "Target served a bot challenge",
  "status": 502,
  "code": "ERR::UPSTREAM::CHALLENGE",
  "detail": "the target answered with a Cloudflare challenge page",
  "retryable": true,
  "doc_url": "https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE",
  "request_id": "01J9ZQ5B2C7D9EXAMPLE00000",
  "target_status": 403
}

Switch on code, retry only when retryable is true (after retry_after_seconds when present), and pass request_id to spicrawl logs get for the full trace.

A failure that did not come from the API (bad flags, no key, network, wait timeout) writes a smaller object:

{ "error": "no API key: run \"spicrawl login\", set SPICRAWL_API_KEY, or pass --api-key", "exit_code": 3 }

Both shapes are distinguishable by the presence of code (API) or exit_code (CLI).

Human mode (terminal)

error: ERR::UPSTREAM::CHALLENGE: the target answered with a Cloudflare challenge page
target status: 403
retryable: yes
request id: 01J9ZQ5B2C7D9EXAMPLE00000 (spicrawl logs get 01J9ZQ5B2C7D9EXAMPLE00000)

A hint: line follows error: when the API sent diagnostics.hint; target status:, retryable: and request id: appear only when known. Non-API failures print one error: ... line.

Streams

spicrawl scrape - reports each failed URL on its own stdout line as {"index", "url", "error"}, where error is the problem document, and keeps going. The process exits 0 if every URL succeeded, else with the code of the first failure. See scrape.

Examples

Branch on the code in bash. --json makes err.json JSON even when you run this at a terminal:

spicrawl scrape https://example.com/products/42 --format markdown -o page.md --json >/dev/null 2>err.json
case $? in
  0)  echo "saved page.md" ;;
  3)  echo "fix the API key" >&2; exit 1 ;;
  5)  sleep "$(jq -r '.retry_after_seconds // 30' err.json)" ;;
  6)  spicrawl scrape https://example.com/products/42 --format markdown --render -o page.md ;;
  10) echo "network down" >&2 ;;
  *)  jq -r '.code // .error' err.json >&2 ;;
esac

Resume a long batch wait:

until spicrawl batch wait "$id" --max-wait 10m; do
  [ $? -eq 11 ] || exit 1   # anything but "still running" is a real failure
done

Fail a CI job early on a bad key:

spicrawl auth status >/dev/null || { echo "SPICRAWL_API_KEY missing or rejected" >&2; exit 1; }

On this page