spicrawlspicrawlDocs

spicrawl scrape

Fetch one URL, or many from stdin, as markdown, html, text, PDF or extracted JSON with spicrawl scrape; every flag maps to a POST /v1/scrape field.

spicrawl scrape <url> calls POST /v1/scrape and prints the result. On a terminal it prints the document itself; piped, it prints one JSON object.

spicrawl scrape https://example.com/blog/launch --format markdown

Only the flags you set are sent. Everything else takes the API's default, or your project's stored default. Each flag below lists the API field it sets; see POST /v1/scrape for the full field rules.

Examples

# A page as clean markdown, for an LLM context window
spicrawl scrape https://example.com/blog/launch --format markdown

# A JS-rendered page, waiting for the element that holds the data
spicrawl scrape https://example.com/pricing --render --wait-for '.plans' --format markdown

# Through your own proxy, with the metadata on stderr
spicrawl scrape https://example.de/produkt/42 --proxy "$MY_PROXY_URL" --meta

# Selector extraction; prints only the extracted data
spicrawl scrape https://example.com/products/1 --extract '{"title":"h1","price":".price"}'

# Render JavaScript, retrying transient failures
spicrawl scrape https://news.example.com --render --retry 3

# Many URLs from a file, 8 at a time, one JSON line each; keep the failures
spicrawl scrape - --format markdown --concurrency 8 < urls.txt | jq -c 'select(.error)'

Flags

Target

FlagAPI fieldMeaning
<url>urlTarget URL, http or https. - reads URLs from stdin (see many URLs).
--method MmethodHTTP method sent to the target, uppercased. Default GET. No request body is sent; only GET results are cached.
--header 'Name: value'custom_headersHeader sent to the target. Repeatable. A value without : exits 2. Hop-by-hop headers (Host, Connection, Content-Length, ...) are refused with ERR::REQUEST::INVALID_PARAMETER.
--body JSON|@filewhole bodyRaw request body as inline JSON, @path, or @- for stdin. Flags you also set override its fields.

Engine

FlagAPI fieldMeaning
--renderjs_renderRender JavaScript in a browser (obscura).
--impersonateimpersonateFetch tier only: present a current Chrome TLS/HTTP2 fingerprint. The API does this by default; --impersonate=false turns it off. No extra cost.
--engine EenginePin fetch, obscura or chromium. Disables escalation.
--mode automodeEscalate fetch, then obscura until one returns a usable page. Cannot be combined with --render or --engine.
--headless=falseheadlessChromium only: run with a real display. Any other engine fails with ERR::REQUEST::CAPABILITY_UNSUPPORTED.
--session IDsession_idReuse a session's cookies and storage. See sessions.

Proxy

FlagAPI fieldMeaning
--proxy URLproxyYour own proxy (http, https, socks5, socks5h).
--proxy-verifyproxy_verifyProbe --proxy before spending engine time. Requires --proxy.

Two more options are on the way: stealth mode Coming soon (the Camoufox browser), and the managed proxy flags --premium-proxy, --country and --sticky-key Coming soon.

Waiting and page control

These need a browser engine (--render, a browser --engine, or --mode auto); otherwise the API returns ERR::REQUEST::INCOMPATIBLE_FLAGS.

FlagAPI fieldMeaning
--wait MSwaitFixed delay after load, in milliseconds.
--wait-for SELwait_forCSS selector to wait for before capture. If it never appears the page is returned with the warning RENDER_DEGRADED.
--wait-for-timeout MSwait_for_timeoutCap on --wait-for, in milliseconds. Requires --wait-for.
--block LISTblock_resourcesComma-separated: images, fonts, media, stylesheets, scripts, or none. Replaces the default (images and fonts blocked).
--actions JSON|@fileactionsOrdered browser actions (click, fill, scroll, ...). See browser actions.
--network-capture JSON|@filenetwork_captureRecord the XHR/fetch responses the page made. Returned under network. See network capture.

Extraction

Each of these makes the API return its JSON envelope instead of the raw document.

FlagAPI fieldMeaning
--format Fresponse_formathtml (API default), markdown, text, json (the envelope) or pdf (needs chromium).
--extract JSON|@fileextractSelector map or JSON Schema. Result under data.
--ai PROMPTai_extract.promptComing soon Model extraction from a plain-language prompt. Adds 4 credits.
--ai-schema JSON|@fileai_extract.schemaComing soon Model extraction to a JSON Schema. Cannot be combined with --ai (exit 2): the API takes a prompt or a schema, not both.
--autoparseautoparseThe page's own JSON-LD, OpenGraph and embedded state, under data.
--linkslinksThe page's absolute, de-duplicated links, under links.

Cleaning

FlagAPI fieldMeaning
--main-contentmain_content_only: trueKeep only the main article.
--no-main-contentmain_content_only: falseKeep the whole document. Markdown defaults to main content. Setting both flags exits 2.
--include SELinclude_tagsKeep only matching subtrees. Repeatable.
--exclude SELexclude_tagsRemove matching elements. Repeatable.
--no-parse-pdfparse_pdf: falseRefuse PDF targets instead of parsing them to text (ERR::EXTRACT::FAILED, 0 credits).

Screenshots

Screenshots need chromium (--engine chromium); otherwise ERR::REQUEST::CAPABILITY_UNSUPPORTED.

FlagAPI fieldMeaning
--screenshotscreenshotCapture a screenshot.
--screenshot-full-pagescreenshot_fullpageFull scrollable page.
--screenshot-selector SELscreenshot_selectorOne element. Mutually exclusive with --screenshot-full-page.
--screenshot-format Fscreenshot_formatpng, jpeg or webp.
--screenshot-quality Nscreenshot_qualityjpeg/webp quality, 1-100.
--screenshot-dir DIRnoneSave screenshots into DIR as <label>.<format>.

Cache and cost

FlagAPI fieldMeaning
--no-cachecache: falseForce a fresh fetch.
--cache-ttl Scache_ttlAccept a cached copy no older than this many seconds. Default 172800 (48 hours).
--max-cost Nmax_costRefuse with ERR::LIMIT::MAX_COST_EXCEEDED (0 credits, exit 5) if the request could cost more.
--allowed-status 403,451allowed_status_codesExtra target statuses to treat as billable success (200, 404 and 410 always are). A non-number exits 2.
--original-statusoriginal_statusUse the target's status as the HTTP status.

Output and retries (CLI only)

FlagDefaultMeaning
-o, --output FILEstdoutWrite the document to FILE (raw bytes for PDF).
--metaoffHuman mode: print engine, credits, cache state, target status, proxy, final URL and request id to stderr. A no-op under --json (and when stdout is not a terminal): the JSON output already carries all of it.
--retry N1Attempts for errors the API marks retryable. 1 means no retry. Waits retry_after_seconds when given, else backs off from 1 s, doubling to at most 30 s.
--concurrency N4With -: parallel requests.
--jsonlon for -With -: one compact JSON result per line. Always on for -.

Output

Human mode (terminal)

Prints the document itself: markdown, html or text. With --extract, --ai, --ai-schema or --autoparse it prints only the extracted data, indented. API warnings go to stderr as warning: ....

--format pdf is binary and is never written to a terminal: pass -o FILE, or --json to get it base64-encoded. Without either, the command exits 2.

spicrawl scrape https://example.com/pricing --format markdown --meta
stderr
engine: fetch
credits: 1 (request cost 1)
cache: miss
target status: 200
proxy: direct
request id: 01J9ZQ4M7R3T8VX2K5N6P0B1CD

JSON mode (piped, or --json)

When the API answers with its JSON envelope (--format json, or any of --extract, --ai, --ai-schema, --autoparse, --links, --network-capture, --screenshot, or --actions containing a screenshot step or an evaluate with return_value), the envelope is printed as-is: url, final_url, status, content, credits, engine, proxy_source, warnings, data, links, screenshots, network.

When the API answers with the raw document (html, markdown, text), the CLI wraps it with the response headers:

{
  "url": "https://example.com/blog/launch",
  "final_url": "https://example.com/blog/launch",
  "status": 200,
  "engine": "fetch",
  "proxy_source": "direct",
  "credits_charged": 1,
  "request_cost": 1,
  "cache_state": "miss",
  "request_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD",
  "warnings": [],
  "content_type": "text/markdown; charset=utf-8",
  "content": "# Launch\n\nToday we ..."
}

status is the target site's status (X-Target-Status), not the API's. Binary documents (--format pdf) carry content_base64 instead of content.

spicrawl scrape https://example.com/blog/launch --format markdown | jq -r .content

Writing to a file with -o

-o FILE writes the document (the extracted data when extraction is on, raw bytes for PDF) to FILE and prints wrote N bytes to FILE on stderr. In JSON mode stdout still carries the metadata object, without the written field and with "output": "FILE" added.

spicrawl scrape https://example.com/report --engine chromium --format pdf -o report.pdf

Screenshots in the envelope are saved next to FILE as <FILE stem>.<label>.<format>, or into --screenshot-dir. Their base64 data is replaced by "file": "<path>" in the printed JSON. With neither -o nor --screenshot-dir, human mode leaves screenshots unsaved and says so on stderr; JSON mode keeps them base64 in the output.

spicrawl scrape https://example.com --engine chromium --screenshot --screenshot-full-page -o home.html
# writes home.html and home.<label>.png

Many URLs from stdin

spicrawl scrape - reads URLs from stdin, one per line. Blank lines and lines starting with # are skipped. It runs --concurrency requests at once (default 4) with the same flags for every URL, and prints one compact JSON line per URL in completion order, in every mode.

spicrawl scrape - --format markdown --concurrency 8 < urls.txt > pages.jsonl

Each success line is the result object above plus index, the 1-based position of the URL in the input. Each failure line is:

{"index":3,"url":"https://example.com/products/43","error":{"type":"https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE","title":"Target served a bot challenge","status":502,"code":"ERR::UPSTREAM::CHALLENGE","retryable":true,"doc_url":"https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE","target_status":403,"request_id":"01J9ZQ5B2C7D9EXAMPLE00000"}}

error is the API's problem document, or {"detail": "...", "exit_code": N} for a network or local failure.

  • Exit code is 0 when every URL succeeded, else the exit code of the first failure. In human mode a N ok, M failed summary goes to stderr.
  • -o is refused with - (exit 2); use --screenshot-dir for screenshots, which are saved as <index>.<label>.<format>.
  • Empty stdin exits 2 with no URLs on stdin.

Errors

A failed scrape writes the API's problem document to stderr (JSON mode) or error:, hint:, target status:, retryable: and request id: lines (human mode), and exits with the code for its code. See exit codes.

You seeExitDo next
ERR::UPSTREAM::CHALLENGE6Retry, then try --render or --mode auto, and your own --proxy.
ERR::UPSTREAM::TIMEOUT with --wait-for6Raise --wait-for-timeout or fix the selector.
ERR::REQUEST::INCOMPATIBLE_FLAGS4Remove the conflicting flag named in detail (for example --proxy-verify without --proxy).
ERR::LIMIT::MAX_COST_EXCEEDED5Raise --max-cost or pick a cheaper engine.
ERR::PROXY::UNREACHABLE7Your --proxy did not accept a connection. Check the URL.
ERR::EXTRACT::FAILED8Check --ai prompt or schema; the call cost 0 credits.
ERR::ENGINE::RENDER_FAILED8A browser action failed (no element matched, a script threw); detail names the step. Fix the selector, add a wait_for before it, or mark the step "on_error":"skip". 0 credits.

On this page