# spicrawl scrape

> Fetch one URL, or many from stdin, as markdown, html, text, PDF or extracted JSON with spicrawl scrape; every flag maps to a POST /v1/scrape field.

Source: https://docs.spicrawl.com/cli/scrape

`spicrawl scrape <url>` calls `POST /v1/scrape` and prints the result. On a terminal it prints the document itself; piped, it prints one JSON object.

```bash
spicrawl scrape https://example.com/blog/launch --format markdown
```

Only the flags you set are sent. Everything else takes the API's default, or your project's stored default. Each flag below lists the API field it sets; see [`POST /v1/scrape`](https://docs.spicrawl.com/api-reference/introduction.md) for the full field rules.

## Examples

```bash
# A page as clean markdown, for an LLM context window
spicrawl scrape https://example.com/blog/launch --format markdown

# A JS-rendered page, waiting for the element that holds the data
spicrawl scrape https://example.com/pricing --render --wait-for '.plans' --format markdown

# Through your own proxy, with the metadata on stderr
spicrawl scrape https://example.de/produkt/42 --proxy "$MY_PROXY_URL" --meta

# Selector extraction; prints only the extracted data
spicrawl scrape https://example.com/products/1 --extract '{"title":"h1","price":".price"}'

# Render JavaScript, retrying transient failures
spicrawl scrape https://news.example.com --render --retry 3

# Many URLs from a file, 8 at a time, one JSON line each; keep the failures
spicrawl scrape - --format markdown --concurrency 8 < urls.txt | jq -c 'select(.error)'
```

## Flags

### Target

| Flag                     | API field        | Meaning                                                                                                                                                                                    |
| ------------------------ | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `<url>`                  | `url`            | Target URL, `http` or `https`. `-` reads URLs from stdin (see [many URLs](#many-urls-from-stdin)).                                                                                         |
| `--method M`             | `method`         | HTTP method sent to the target, uppercased. Default `GET`. No request body is sent; only `GET` results are cached.                                                                         |
| `--header 'Name: value'` | `custom_headers` | Header sent to the target. Repeatable. A value without `:` exits `2`. Hop-by-hop headers (`Host`, `Connection`, `Content-Length`, ...) are refused with `ERR::REQUEST::INVALID_PARAMETER`. |
| `--body JSON\|@file`     | whole body       | Raw request body as inline JSON, `@path`, or `@-` for stdin. Flags you also set override its fields.                                                                                       |

### Engine

| Flag               | API field     | Meaning                                                                                                                                           |
| ------------------ | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--render`         | `js_render`   | Render JavaScript in a browser (obscura).                                                                                                         |
| `--impersonate`    | `impersonate` | Fetch tier only: present a current Chrome TLS/HTTP2 fingerprint. The API does this by default; `--impersonate=false` turns it off. No extra cost. |
| `--engine E`       | `engine`      | Pin `fetch`, `obscura` or `chromium`. Disables escalation.                                                                                        |
| `--mode auto`      | `mode`        | Escalate fetch, then obscura until one returns a usable page. Cannot be combined with `--render` or `--engine`.                                   |
| `--headless=false` | `headless`    | Chromium only: run with a real display. Any other engine fails with `ERR::REQUEST::CAPABILITY_UNSUPPORTED`.                                       |
| `--session ID`     | `session_id`  | Reuse a session's cookies and storage. See [sessions](https://docs.spicrawl.com/cli/sessions-and-browser.md).                                                                 |

### Proxy

| Flag             | API field      | Meaning                                                          |
| ---------------- | -------------- | ---------------------------------------------------------------- |
| `--proxy URL`    | `proxy`        | Your own proxy (`http`, `https`, `socks5`, `socks5h`).           |
| `--proxy-verify` | `proxy_verify` | Probe `--proxy` before spending engine time. Requires `--proxy`. |

Two more options are on the way: stealth mode (coming soon) (the Camoufox browser), and the [managed proxy](https://docs.spicrawl.com/guides/proxies-and-geo.md#managed-proxy-pool-coming-soon) flags `--premium-proxy`, `--country` and `--sticky-key` (coming soon).

### Waiting and page control

These need a browser engine (`--render`, a browser `--engine`, or `--mode auto`); otherwise the API returns `ERR::REQUEST::INCOMPATIBLE_FLAGS`.

| Flag                            | API field          | Meaning                                                                                                                            |
| ------------------------------- | ------------------ | ---------------------------------------------------------------------------------------------------------------------------------- |
| `--wait MS`                     | `wait`             | Fixed delay after load, in milliseconds.                                                                                           |
| `--wait-for SEL`                | `wait_for`         | CSS selector to wait for before capture. If it never appears the page is returned with the warning `RENDER_DEGRADED`.              |
| `--wait-for-timeout MS`         | `wait_for_timeout` | Cap on `--wait-for`, in milliseconds. Requires `--wait-for`.                                                                       |
| `--block LIST`                  | `block_resources`  | Comma-separated: `images`, `fonts`, `media`, `stylesheets`, `scripts`, or `none`. Replaces the default (images and fonts blocked). |
| `--actions JSON\|@file`         | `actions`          | Ordered browser actions (click, fill, scroll, ...). See [browser actions](https://docs.spicrawl.com/guides/browser-actions.md).                                |
| `--network-capture JSON\|@file` | `network_capture`  | Record the XHR/fetch responses the page made. Returned under `network`. See [network capture](https://docs.spicrawl.com/guides/network-capture.md).            |

### Extraction

Each of these makes the API return its JSON envelope instead of the raw document.

| Flag                      | API field           | Meaning                                                                                                                                   |
| ------------------------- | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- |
| `--format F`              | `response_format`   | `html` (API default), `markdown`, `text`, `json` (the envelope) or `pdf` (needs chromium).                                                |
| `--extract JSON\|@file`   | `extract`           | Selector map or JSON Schema. Result under `data`.                                                                                         |
| `--ai PROMPT`             | `ai_extract.prompt` | (coming soon) Model extraction from a plain-language prompt. Adds 4 credits.                                                              |
| `--ai-schema JSON\|@file` | `ai_extract.schema` | (coming soon) Model extraction to a JSON Schema. Cannot be combined with `--ai` (exit `2`): the API takes a prompt or a schema, not both. |
| `--autoparse`             | `autoparse`         | The page's own JSON-LD, OpenGraph and embedded state, under `data`.                                                                       |
| `--links`                 | `links`             | The page's absolute, de-duplicated links, under `links`.                                                                                  |

### Cleaning

| Flag                | API field                  | Meaning                                                                                   |
| ------------------- | -------------------------- | ----------------------------------------------------------------------------------------- |
| `--main-content`    | `main_content_only: true`  | Keep only the main article.                                                               |
| `--no-main-content` | `main_content_only: false` | Keep the whole document. Markdown defaults to main content. Setting both flags exits `2`. |
| `--include SEL`     | `include_tags`             | Keep only matching subtrees. Repeatable.                                                  |
| `--exclude SEL`     | `exclude_tags`             | Remove matching elements. Repeatable.                                                     |
| `--no-parse-pdf`    | `parse_pdf: false`         | Refuse PDF targets instead of parsing them to text (`ERR::EXTRACT::FAILED`, 0 credits).   |

### Screenshots

Screenshots need chromium (`--engine chromium`); otherwise `ERR::REQUEST::CAPABILITY_UNSUPPORTED`.

| Flag                        | API field             | Meaning                                                        |
| --------------------------- | --------------------- | -------------------------------------------------------------- |
| `--screenshot`              | `screenshot`          | Capture a screenshot.                                          |
| `--screenshot-full-page`    | `screenshot_fullpage` | Full scrollable page.                                          |
| `--screenshot-selector SEL` | `screenshot_selector` | One element. Mutually exclusive with `--screenshot-full-page`. |
| `--screenshot-format F`     | `screenshot_format`   | `png`, `jpeg` or `webp`.                                       |
| `--screenshot-quality N`    | `screenshot_quality`  | jpeg/webp quality, 1-100.                                      |
| `--screenshot-dir DIR`      | none                  | Save screenshots into `DIR` as `<label>.<format>`.             |

### Cache and cost

| Flag                       | API field              | Meaning                                                                                                   |
| -------------------------- | ---------------------- | --------------------------------------------------------------------------------------------------------- |
| `--no-cache`               | `cache: false`         | Force a fresh fetch.                                                                                      |
| `--cache-ttl S`            | `cache_ttl`            | Accept a cached copy no older than this many seconds. Default 172800 (48 hours).                          |
| `--max-cost N`             | `max_cost`             | Refuse with `ERR::LIMIT::MAX_COST_EXCEEDED` (0 credits, exit `5`) if the request could cost more.         |
| `--allowed-status 403,451` | `allowed_status_codes` | Extra target statuses to treat as billable success (200, 404 and 410 always are). A non-number exits `2`. |
| `--original-status`        | `original_status`      | Use the target's status as the HTTP status.                                                               |

### Output and retries (CLI only)

| Flag                  | Default    | Meaning                                                                                                                                                                                                          |
| --------------------- | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `-o`, `--output FILE` | stdout     | Write the document to `FILE` (raw bytes for PDF).                                                                                                                                                                |
| `--meta`              | off        | Human mode: print engine, credits, cache state, target status, proxy, final URL and request id to stderr. A no-op under `--json` (and when stdout is not a terminal): the JSON output already carries all of it. |
| `--retry N`           | `1`        | Attempts for errors the API marks `retryable`. `1` means no retry. Waits `retry_after_seconds` when given, else backs off from 1 s, doubling to at most 30 s.                                                    |
| `--concurrency N`     | `4`        | With `-`: parallel requests.                                                                                                                                                                                     |
| `--jsonl`             | on for `-` | With `-`: one compact JSON result per line. Always on for `-`.                                                                                                                                                   |

## Output

### Human mode (terminal)

Prints the document itself: markdown, html or text. With `--extract`, `--ai`, `--ai-schema` or `--autoparse` it prints only the extracted `data`, indented. API warnings go to stderr as `warning: ...`.

`--format pdf` is binary and is never written to a terminal: pass `-o FILE`, or `--json` to get it base64-encoded. Without either, the command exits `2`.

```bash
spicrawl scrape https://example.com/pricing --format markdown --meta
```

```text title="stderr"
engine: fetch
credits: 1 (request cost 1)
cache: miss
target status: 200
proxy: direct
request id: 01J9ZQ4M7R3T8VX2K5N6P0B1CD
```

### JSON mode (piped, or `--json`)

When the API answers with its JSON envelope (`--format json`, or any of `--extract`, `--ai`, `--ai-schema`, `--autoparse`, `--links`, `--network-capture`, `--screenshot`, or `--actions` containing a `screenshot` step or an `evaluate` with `return_value`), the envelope is printed as-is: `url`, `final_url`, `status`, `content`, `credits`, `engine`, `proxy_source`, `warnings`, `data`, `links`, `screenshots`, `network`.

When the API answers with the raw document (`html`, `markdown`, `text`), the CLI wraps it with the response headers:

```json
{
  "url": "https://example.com/blog/launch",
  "final_url": "https://example.com/blog/launch",
  "status": 200,
  "engine": "fetch",
  "proxy_source": "direct",
  "credits_charged": 1,
  "request_cost": 1,
  "cache_state": "miss",
  "request_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD",
  "warnings": [],
  "content_type": "text/markdown; charset=utf-8",
  "content": "# Launch\n\nToday we ..."
}
```

`status` is the target site's status (`X-Target-Status`), not the API's. Binary documents (`--format pdf`) carry `content_base64` instead of `content`.

```bash
spicrawl scrape https://example.com/blog/launch --format markdown | jq -r .content
```

### Writing to a file with `-o`

`-o FILE` writes the document (the extracted `data` when extraction is on, raw bytes for PDF) to `FILE` and prints `wrote N bytes to FILE` on stderr. In JSON mode stdout still carries the metadata object, without the written field and with `"output": "FILE"` added.

```bash
spicrawl scrape https://example.com/report --engine chromium --format pdf -o report.pdf
```

Screenshots in the envelope are saved next to `FILE` as `<FILE stem>.<label>.<format>`, or into `--screenshot-dir`. Their base64 `data` is replaced by `"file": "<path>"` in the printed JSON. With neither `-o` nor `--screenshot-dir`, human mode leaves screenshots unsaved and says so on stderr; JSON mode keeps them base64 in the output.

```bash
spicrawl scrape https://example.com --engine chromium --screenshot --screenshot-full-page -o home.html
# writes home.html and home.<label>.png
```

## Many URLs from stdin

`spicrawl scrape -` reads URLs from stdin, one per line. Blank lines and lines starting with `#` are skipped. It runs `--concurrency` requests at once (default 4) with the same flags for every URL, and prints one compact JSON line per URL in completion order, in every mode.

```bash
spicrawl scrape - --format markdown --concurrency 8 < urls.txt > pages.jsonl
```

Each success line is the result object above plus `index`, the 1-based position of the URL in the input. Each failure line is:

```json
{"index":3,"url":"https://example.com/products/43","error":{"type":"https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE","title":"Target served a bot challenge","status":502,"code":"ERR::UPSTREAM::CHALLENGE","retryable":true,"doc_url":"https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE","target_status":403,"request_id":"01J9ZQ5B2C7D9EXAMPLE00000"}}
```

`error` is the API's problem document, or `{"detail": "...", "exit_code": N}` for a network or local failure.

* Exit code is `0` when every URL succeeded, else the exit code of the first failure. In human mode a `N ok, M failed` summary goes to stderr.
* `-o` is refused with `-` (exit `2`); use `--screenshot-dir` for screenshots, which are saved as `<index>.<label>.<format>`.
* Empty stdin exits `2` with `no URLs on stdin`.

## Errors

A failed scrape writes the API's problem document to stderr (JSON mode) or `error:`, `hint:`, `target status:`, `retryable:` and `request id:` lines (human mode), and exits with the code for its `code`. See [exit codes](https://docs.spicrawl.com/cli/exit-codes.md).

| You see                                    | Exit | Do next                                                                                                                                                                               |
| ------------------------------------------ | ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `ERR::UPSTREAM::CHALLENGE`                 | `6`  | Retry, then try `--render` or `--mode auto`, and your own `--proxy`.                                                                                                                  |
| `ERR::UPSTREAM::TIMEOUT` with `--wait-for` | `6`  | Raise `--wait-for-timeout` or fix the selector.                                                                                                                                       |
| `ERR::REQUEST::INCOMPATIBLE_FLAGS`         | `4`  | Remove the conflicting flag named in `detail` (for example `--proxy-verify` without `--proxy`).                                                                                       |
| `ERR::LIMIT::MAX_COST_EXCEEDED`            | `5`  | Raise `--max-cost` or pick a cheaper engine.                                                                                                                                          |
| `ERR::PROXY::UNREACHABLE`                  | `7`  | Your `--proxy` did not accept a connection. Check the URL.                                                                                                                            |
| `ERR::EXTRACT::FAILED`                     | `8`  | Check `--ai` prompt or schema; the call cost 0 credits.                                                                                                                               |
| `ERR::ENGINE::RENDER_FAILED`               | `8`  | A browser action failed (no element matched, a script threw); `detail` names the step. Fix the selector, add a `wait_for` before it, or mark the step `"on_error":"skip"`. 0 credits. |
