# Extract structured data

> Get JSON out of a page three ways: autoparse for the page's own metadata, a selector map for exact fields, or a JSON Schema with selectors for typed, validated output.

Source: https://docs.spicrawl.com/guides/structured-data

Use this when you want fields (a title, a price, a list of reviews) rather than the page text. There are three tiers, cheapest to write first, and none of them costs extra credits:

1. **`autoparse: true`** returns what the page already says about itself: JSON-LD, OpenGraph, Twitter Card, microdata, RDFa, `<head>` metadata and embedded SPA state. No selectors.
2. **`extract` as a selector map** returns exactly the fields you name, as strings.
3. **`extract` as a JSON Schema** whose properties carry a `selector`, which also coerces values to the declared types (`"$1,234.56"` becomes `1234.56`) and validates the result.

All three write to `data` in the JSON envelope. For pages where selectors are impractical, [AI extraction](https://docs.spicrawl.com/guides/ai-extraction.md) (coming soon) will let a model read the page.

## Minimal request

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/products/42",
    "extract": {
      "title": "h1",
      "price": ".price",
      "images": {"selector": ".gallery img", "kind": "list", "output": "attr", "attribute": "src"},
      "next_page": "a.next@href"
    }
  }'
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={
        "url": "https://example.com/products/42",
        "extract": {
            "title": "h1",
            "price": ".price",
            "images": {"selector": ".gallery img", "kind": "list", "output": "attr", "attribute": "src"},
            "next_page": "a.next@href",
        },
    },
    timeout=120,
)
r.raise_for_status()
env = r.json()
print(env["data"], env.get("empty_fields"))
```

```typescript title="TypeScript"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    url: "https://example.com/products/42",
    extract: {
      title: "h1",
      price: ".price",
      images: { selector: ".gallery img", kind: "list", output: "attr", attribute: "src" },
      next_page: "a.next@href",
    },
  }),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
const { data, empty_fields } = await r.json();
```

```bash title="CLI"
spicrawl scrape https://example.com/products/42 \
  --extract '{"title":"h1","price":".price","next_page":"a.next@href"}'
```

## What comes back

The JSON envelope, with your fields under `data`:

```json
{
  "url": "https://example.com/products/42",
  "final_url": "https://example.com/products/42",
  "status": 200,
  "content": "<html>...</html>",
  "credits": 1,
  "engine": "fetch",
  "proxy_source": "pool",
  "warnings": [],
  "data": {
    "title": "Walnut Desk",
    "price": "$349.00",
    "images": ["https://example.com/img/desk-1.jpg", "https://example.com/img/desk-2.jpg"]
  },
  "empty_fields": ["next_page"]
}
```

`empty_fields` lists rules that matched nothing. It is the first sign that a site changed its markup: alert on it rather than on an empty value. `extract` forces the envelope; if you also asked for `markdown`, `content` holds the markdown and the response carries `X-Warning: FORMAT_COERCED`.

## Tier 1: autoparse

Send `"autoparse": true` (CLI: `--autoparse`). `data` then holds only the keys whose source exists on the page:

| Key              | Contents                                                                                                                                        |
| ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| `json_ld`        | Every JSON-LD block, `@graph` flattened, as an array.                                                                                           |
| `open_graph`     | `og:*` properties, with `og:image` as structured objects.                                                                                       |
| `twitter`        | `twitter:*` card properties.                                                                                                                    |
| `microdata`      | `itemscope` items with `type` and `properties`.                                                                                                 |
| `rdfa`           | RDFa items.                                                                                                                                     |
| `meta`           | `title`, `description`, `canonical`, `favicon`, `language`.                                                                                     |
| `embedded_state` | Hydration payloads, keyed by source: `next` (`__NEXT_DATA__`), `nuxt`, `angular`, `apollo`, `redux`, `remix`, `run_params`, `data_layer` (GA4). |

URLs are resolved to absolute. `embedded_state` is often the richest source (variants, stock, regional prices), but **Next.js App Router pages are not covered**: they ship their data as `self.__next_f` flight chunks, not `__NEXT_DATA__`, so they yield no `embedded_state`. Use a selector map or [network capture](https://docs.spicrawl.com/guides/network-capture.md) for those.

## Tier 2: selector map

Each field is a selector string or a rule object.

* `"title": "h1"` returns the cleaned text of the first match.
* `"next_page": "a.next@href"` returns the `href` attribute of the first match.
* A rule object takes `selector` (required), `type` (`css` or `xpath`), `kind` (`item` for the first match, `list` for all), `output` (`text`, `html`, `outer_html`, `attr`, `table_json`, `table_array`, `exists`, `count`), `attribute` (required with `output: "attr"`, forbidden otherwise), `clean` (default `true`) and `nested`.
* `nested` is a selector map evaluated inside each match. It requires `output` `html` or `outer_html`.

```json
{
  "reviews": {
    "selector": ".review", "kind": "list", "output": "outer_html",
    "nested": { "author": ".author", "rating": ".stars@data-rating", "body": "p" }
  }
}
```

## Tier 3: JSON Schema with selectors

When `extract` has `"type": "object"` and a `properties` object, it is read as a JSON Schema. Every property (or its `items`, or its nested `properties`) needs a `selector`; `selector_type` (`css` or `xpath`) and `attribute` are also accepted.

```json
{
  "type": "object",
  "properties": {
    "title":    { "type": "string",  "selector": "h1" },
    "price":    { "type": "number",  "selector": ".price" },
    "sku":      { "type": "string",  "selector": ".sku", "attribute": "data-sku" },
    "tags":     { "type": "array",   "selector": ".tag", "items": { "type": "string" } }
  },
  "required": ["title", "price"],
  "strict": false
}
```

Values are coerced to the declared type: `"$1,234.56"` becomes `1234.56` and `"1.234,56"` becomes `1234.56`. Properties that match nothing are dropped and listed in `empty_fields`. With `strict: false` (the default) a field that fails validation is dropped with a warning; with `strict: true` the whole extraction fails.

## Limits

| Limit                                           | Value                                  |
| ----------------------------------------------- | -------------------------------------- |
| Fields per level (selector map or `properties`) | 1–100                                  |
| `nested` depth                                  | 5                                      |
| `extract: {}`                                   | Refused (400). Omit the field instead. |

## Failure modes

| Code                               | HTTP | When                                                                                                                                                              | What to do                                                                           |
| ---------------------------------- | ---- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ |
| `ERR::REQUEST::INVALID_PARAMETER`  | 400  | A malformed rule: over 100 fields, depth above 5, `attribute` without `output: "attr"`, a schema property with no `selector`. Checked before fetching; 0 credits. | Fix the rule at the path the `detail` names.                                         |
| `ERR::EXTRACT::INVALID_RULES`      | 400  | `extract_preset` was set. Extraction presets (coming soon): until they launch, every value fails (after the fetch, at 0 credits).                                 | Put the rules inline in `extract`.                                                   |
| `ERR::REQUEST::INCOMPATIBLE_FLAGS` | 400  | Both `extract` and `extract_preset`.                                                                                                                              | Send only `extract`.                                                                 |
| `ERR::EXTRACT::FAILED`             | 502  | The page could not be extracted: the target is a PDF, an image or other non-HTML body, or a `strict` schema failed validation. 0 credits.                         | For a PDF use `response_format=markdown`; otherwise check the target and the schema. |

A broken selector is not an error: it returns HTTP 200 with the field in `empty_fields`.

## Cost

`extract` and `autoparse` add 0 credits. You pay the engine and proxy tier (1 credit on fetch, 3 with `js_render`). Failures cost 0. See [Credits](https://docs.spicrawl.com/credits.md).

## Related

* [AI extraction](https://docs.spicrawl.com/guides/ai-extraction.md) (coming soon)
* [Network capture](https://docs.spicrawl.com/guides/network-capture.md) for data the page loads over XHR
* [JavaScript rendering](https://docs.spicrawl.com/guides/javascript-rendering.md) when selectors match nothing on a client-rendered page
* [Batch](https://docs.spicrawl.com/guides/batch.md). Batch items return the raw page today; run extraction per URL with `/v1/scrape`.
