spicrawl scrape
Fetch one URL, or many from stdin, as markdown, html, text, PDF or extracted JSON with spicrawl scrape; every flag maps to a POST /v1/scrape field.
spicrawl scrape <url> calls POST /v1/scrape and prints the result. On a terminal it prints the document itself; piped, it prints one JSON object.
spicrawl scrape https://example.com/blog/launch --format markdownOnly the flags you set are sent. Everything else takes the API's default, or your project's stored default. Each flag below lists the API field it sets; see POST /v1/scrape for the full field rules.
Examples
# A page as clean markdown, for an LLM context window
spicrawl scrape https://example.com/blog/launch --format markdown
# A JS-rendered page, waiting for the element that holds the data
spicrawl scrape https://example.com/pricing --render --wait-for '.plans' --format markdown
# Through your own proxy, with the metadata on stderr
spicrawl scrape https://example.de/produkt/42 --proxy "$MY_PROXY_URL" --meta
# Selector extraction; prints only the extracted data
spicrawl scrape https://example.com/products/1 --extract '{"title":"h1","price":".price"}'
# Render JavaScript, retrying transient failures
spicrawl scrape https://news.example.com --render --retry 3
# Many URLs from a file, 8 at a time, one JSON line each; keep the failures
spicrawl scrape - --format markdown --concurrency 8 < urls.txt | jq -c 'select(.error)'Flags
Target
| Flag | API field | Meaning |
|---|---|---|
<url> | url | Target URL, http or https. - reads URLs from stdin (see many URLs). |
--method M | method | HTTP method sent to the target, uppercased. Default GET. No request body is sent; only GET results are cached. |
--header 'Name: value' | custom_headers | Header sent to the target. Repeatable. A value without : exits 2. Hop-by-hop headers (Host, Connection, Content-Length, ...) are refused with ERR::REQUEST::INVALID_PARAMETER. |
--body JSON|@file | whole body | Raw request body as inline JSON, @path, or @- for stdin. Flags you also set override its fields. |
Engine
| Flag | API field | Meaning |
|---|---|---|
--render | js_render | Render JavaScript in a browser (obscura). |
--impersonate | impersonate | Fetch tier only: present a current Chrome TLS/HTTP2 fingerprint. The API does this by default; --impersonate=false turns it off. No extra cost. |
--engine E | engine | Pin fetch, obscura or chromium. Disables escalation. |
--mode auto | mode | Escalate fetch, then obscura until one returns a usable page. Cannot be combined with --render or --engine. |
--headless=false | headless | Chromium only: run with a real display. Any other engine fails with ERR::REQUEST::CAPABILITY_UNSUPPORTED. |
--session ID | session_id | Reuse a session's cookies and storage. See sessions. |
Proxy
| Flag | API field | Meaning |
|---|---|---|
--proxy URL | proxy | Your own proxy (http, https, socks5, socks5h). |
--proxy-verify | proxy_verify | Probe --proxy before spending engine time. Requires --proxy. |
Two more options are on the way: stealth mode Coming soon (the Camoufox browser), and the managed proxy flags --premium-proxy, --country and --sticky-key Coming soon.
Waiting and page control
These need a browser engine (--render, a browser --engine, or --mode auto); otherwise the API returns ERR::REQUEST::INCOMPATIBLE_FLAGS.
| Flag | API field | Meaning |
|---|---|---|
--wait MS | wait | Fixed delay after load, in milliseconds. |
--wait-for SEL | wait_for | CSS selector to wait for before capture. If it never appears the page is returned with the warning RENDER_DEGRADED. |
--wait-for-timeout MS | wait_for_timeout | Cap on --wait-for, in milliseconds. Requires --wait-for. |
--block LIST | block_resources | Comma-separated: images, fonts, media, stylesheets, scripts, or none. Replaces the default (images and fonts blocked). |
--actions JSON|@file | actions | Ordered browser actions (click, fill, scroll, ...). See browser actions. |
--network-capture JSON|@file | network_capture | Record the XHR/fetch responses the page made. Returned under network. See network capture. |
Extraction
Each of these makes the API return its JSON envelope instead of the raw document.
| Flag | API field | Meaning |
|---|---|---|
--format F | response_format | html (API default), markdown, text, json (the envelope) or pdf (needs chromium). |
--extract JSON|@file | extract | Selector map or JSON Schema. Result under data. |
--ai PROMPT | ai_extract.prompt | Coming soon Model extraction from a plain-language prompt. Adds 4 credits. |
--ai-schema JSON|@file | ai_extract.schema | Coming soon Model extraction to a JSON Schema. Cannot be combined with --ai (exit 2): the API takes a prompt or a schema, not both. |
--autoparse | autoparse | The page's own JSON-LD, OpenGraph and embedded state, under data. |
--links | links | The page's absolute, de-duplicated links, under links. |
Cleaning
| Flag | API field | Meaning |
|---|---|---|
--main-content | main_content_only: true | Keep only the main article. |
--no-main-content | main_content_only: false | Keep the whole document. Markdown defaults to main content. Setting both flags exits 2. |
--include SEL | include_tags | Keep only matching subtrees. Repeatable. |
--exclude SEL | exclude_tags | Remove matching elements. Repeatable. |
--no-parse-pdf | parse_pdf: false | Refuse PDF targets instead of parsing them to text (ERR::EXTRACT::FAILED, 0 credits). |
Screenshots
Screenshots need chromium (--engine chromium); otherwise ERR::REQUEST::CAPABILITY_UNSUPPORTED.
| Flag | API field | Meaning |
|---|---|---|
--screenshot | screenshot | Capture a screenshot. |
--screenshot-full-page | screenshot_fullpage | Full scrollable page. |
--screenshot-selector SEL | screenshot_selector | One element. Mutually exclusive with --screenshot-full-page. |
--screenshot-format F | screenshot_format | png, jpeg or webp. |
--screenshot-quality N | screenshot_quality | jpeg/webp quality, 1-100. |
--screenshot-dir DIR | none | Save screenshots into DIR as <label>.<format>. |
Cache and cost
| Flag | API field | Meaning |
|---|---|---|
--no-cache | cache: false | Force a fresh fetch. |
--cache-ttl S | cache_ttl | Accept a cached copy no older than this many seconds. Default 172800 (48 hours). |
--max-cost N | max_cost | Refuse with ERR::LIMIT::MAX_COST_EXCEEDED (0 credits, exit 5) if the request could cost more. |
--allowed-status 403,451 | allowed_status_codes | Extra target statuses to treat as billable success (200, 404 and 410 always are). A non-number exits 2. |
--original-status | original_status | Use the target's status as the HTTP status. |
Output and retries (CLI only)
| Flag | Default | Meaning |
|---|---|---|
-o, --output FILE | stdout | Write the document to FILE (raw bytes for PDF). |
--meta | off | Human mode: print engine, credits, cache state, target status, proxy, final URL and request id to stderr. A no-op under --json (and when stdout is not a terminal): the JSON output already carries all of it. |
--retry N | 1 | Attempts for errors the API marks retryable. 1 means no retry. Waits retry_after_seconds when given, else backs off from 1 s, doubling to at most 30 s. |
--concurrency N | 4 | With -: parallel requests. |
--jsonl | on for - | With -: one compact JSON result per line. Always on for -. |
Output
Human mode (terminal)
Prints the document itself: markdown, html or text. With --extract, --ai, --ai-schema or --autoparse it prints only the extracted data, indented. API warnings go to stderr as warning: ....
--format pdf is binary and is never written to a terminal: pass -o FILE, or --json to get it base64-encoded. Without either, the command exits 2.
spicrawl scrape https://example.com/pricing --format markdown --metaengine: fetch
credits: 1 (request cost 1)
cache: miss
target status: 200
proxy: direct
request id: 01J9ZQ4M7R3T8VX2K5N6P0B1CDJSON mode (piped, or --json)
When the API answers with its JSON envelope (--format json, or any of --extract, --ai, --ai-schema, --autoparse, --links, --network-capture, --screenshot, or --actions containing a screenshot step or an evaluate with return_value), the envelope is printed as-is: url, final_url, status, content, credits, engine, proxy_source, warnings, data, links, screenshots, network.
When the API answers with the raw document (html, markdown, text), the CLI wraps it with the response headers:
{
"url": "https://example.com/blog/launch",
"final_url": "https://example.com/blog/launch",
"status": 200,
"engine": "fetch",
"proxy_source": "direct",
"credits_charged": 1,
"request_cost": 1,
"cache_state": "miss",
"request_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD",
"warnings": [],
"content_type": "text/markdown; charset=utf-8",
"content": "# Launch\n\nToday we ..."
}status is the target site's status (X-Target-Status), not the API's. Binary documents (--format pdf) carry content_base64 instead of content.
spicrawl scrape https://example.com/blog/launch --format markdown | jq -r .contentWriting to a file with -o
-o FILE writes the document (the extracted data when extraction is on, raw bytes for PDF) to FILE and prints wrote N bytes to FILE on stderr. In JSON mode stdout still carries the metadata object, without the written field and with "output": "FILE" added.
spicrawl scrape https://example.com/report --engine chromium --format pdf -o report.pdfScreenshots in the envelope are saved next to FILE as <FILE stem>.<label>.<format>, or into --screenshot-dir. Their base64 data is replaced by "file": "<path>" in the printed JSON. With neither -o nor --screenshot-dir, human mode leaves screenshots unsaved and says so on stderr; JSON mode keeps them base64 in the output.
spicrawl scrape https://example.com --engine chromium --screenshot --screenshot-full-page -o home.html
# writes home.html and home.<label>.pngMany URLs from stdin
spicrawl scrape - reads URLs from stdin, one per line. Blank lines and lines starting with # are skipped. It runs --concurrency requests at once (default 4) with the same flags for every URL, and prints one compact JSON line per URL in completion order, in every mode.
spicrawl scrape - --format markdown --concurrency 8 < urls.txt > pages.jsonlEach success line is the result object above plus index, the 1-based position of the URL in the input. Each failure line is:
{"index":3,"url":"https://example.com/products/43","error":{"type":"https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE","title":"Target served a bot challenge","status":502,"code":"ERR::UPSTREAM::CHALLENGE","retryable":true,"doc_url":"https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE","target_status":403,"request_id":"01J9ZQ5B2C7D9EXAMPLE00000"}}error is the API's problem document, or {"detail": "...", "exit_code": N} for a network or local failure.
- Exit code is
0when every URL succeeded, else the exit code of the first failure. In human mode aN ok, M failedsummary goes to stderr. -ois refused with-(exit2); use--screenshot-dirfor screenshots, which are saved as<index>.<label>.<format>.- Empty stdin exits
2withno URLs on stdin.
Errors
A failed scrape writes the API's problem document to stderr (JSON mode) or error:, hint:, target status:, retryable: and request id: lines (human mode), and exits with the code for its code. See exit codes.
| You see | Exit | Do next |
|---|---|---|
ERR::UPSTREAM::CHALLENGE | 6 | Retry, then try --render or --mode auto, and your own --proxy. |
ERR::UPSTREAM::TIMEOUT with --wait-for | 6 | Raise --wait-for-timeout or fix the selector. |
ERR::REQUEST::INCOMPATIBLE_FLAGS | 4 | Remove the conflicting flag named in detail (for example --proxy-verify without --proxy). |
ERR::LIMIT::MAX_COST_EXCEEDED | 5 | Raise --max-cost or pick a cheaper engine. |
ERR::PROXY::UNREACHABLE | 7 | Your --proxy did not accept a connection. Check the URL. |
ERR::EXTRACT::FAILED | 8 | Check --ai prompt or schema; the call cost 0 credits. |
ERR::ENGINE::RENDER_FAILED | 8 | A browser action failed (no element matched, a script threw); detail names the step. Fix the selector, add a wait_for before it, or mark the step "on_error":"skip". 0 credits. |