# Extract data with a model (coming soon)

> Coming soon: describe the data in plain language or as a JSON Schema with ai_extract and get it back under data, for 4 credits on top of the engine price.

Source: https://docs.spicrawl.com/guides/ai-extraction

> **Coming soon:** AI extraction is not available yet. Until it launches, a request with `ai_extract` is refused before the fetch, at 0 credits. To get structured data today, use [selectors or autoparse](https://docs.spicrawl.com/guides/structured-data.md). This page describes how `ai_extract` will work.

Use this when writing selectors is impractical: layouts differ from page to page, the value is buried in prose ("ships in 3–5 business days"), or you need the same fields from many unrelated sites. `ai_extract` gives the page to a model as cleaned markdown and returns what it extracted under `data`. It costs 4 credits on top of the engine price, so prefer [selectors or autoparse](https://docs.spicrawl.com/guides/structured-data.md) for a site you scrape repeatedly with stable markup.

## Minimal request

Send exactly one of `prompt` or `schema`.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/products/42",
    "ai_extract": {"prompt": "Product name, price with currency, and delivery time in days"}
  }'
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={
        "url": "https://example.com/products/42",
        "ai_extract": {"prompt": "Product name, price with currency, and delivery time in days"},
    },
    timeout=180,
)
r.raise_for_status()
print(r.json()["data"])
```

```typescript title="TypeScript"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    url: "https://example.com/products/42",
    ai_extract: { prompt: "Product name, price with currency, and delivery time in days" },
  }),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
const { data } = await r.json();
```

```bash title="CLI"
spicrawl scrape https://example.com/products/42 --ai "Product name, price with currency, and delivery time in days"
```

To fix the output shape, send a JSON Schema instead of a prompt (CLI: `--ai-schema @schema.json`):

```json
{
  "url": "https://example.com/products/42",
  "ai_extract": {
    "schema": {
      "type": "object",
      "properties": {
        "name": { "type": "string" },
        "price": { "type": "number" },
        "currency": { "type": "string" },
        "delivery_days_max": { "type": "integer" }
      },
      "required": ["name", "price"]
    }
  }
}
```

The schema needs no `selector` keywords; that is the difference from the schema form of `extract`.

## What comes back

The JSON envelope, with the model's output under `data`:

```json
{
  "url": "https://example.com/products/42",
  "final_url": "https://example.com/products/42",
  "status": 200,
  "content": "<html>...</html>",
  "credits": 5,
  "engine": "fetch",
  "proxy_source": "direct",
  "warnings": [],
  "data": {
    "name": "Walnut Desk",
    "price": 349,
    "currency": "USD",
    "delivery_days_max": 5
  }
}
```

`credits` is the engine price plus 4 (here 1 + 4). `ai_extract` always forces the envelope; if you asked for `markdown`, `content` holds it and the response carries `X-Warning: FORMAT_COERCED`.

## Options that matter

| Field                                               | Rule                                                                                                                                     |
| --------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| `ai_extract.prompt`                                 | Non-empty plain-language description. Name units and formats you want ("price as a number, currency as ISO 4217").                       |
| `ai_extract.schema`                                 | JSON Schema object with at least one property. Use it when downstream code parses the result.                                            |
| `main_content_only`, `include_tags`, `exclude_tags` | Shape the markdown the model reads. Scoping to the relevant container improves accuracy on busy pages. See [Markdown](https://docs.spicrawl.com/guides/markdown.md). |
| `js_render`                                         | Needed when the data is built by JavaScript; the model only sees what the engine retrieved.                                              |

Both `prompt` and `schema` is 400 `ERR::REQUEST::INCOMPATIBLE_FLAGS`; neither is 400 `ERR::REQUEST::INVALID_PARAMETER`.

## When to prefer selectors

| Situation                              | Use                                                                  |
| -------------------------------------- | -------------------------------------------------------------------- |
| One site, stable markup, many pages    | `extract` selector map or schema: 0 extra credits and deterministic. |
| The page has JSON-LD or embedded state | `autoparse`: 0 extra credits.                                        |
| Many unrelated sites, one target shape | `ai_extract` with a `schema`.                                        |
| Values inside free text                | `ai_extract`.                                                        |

You can combine them: run `ai_extract` once to find where the data lives, then write selectors for the site.

## Failure modes

| Code                               | HTTP | When                                                                                                         | What to do                                                                                                                 |
| ---------------------------------- | ---- | ------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------- |
| `ERR::INTERNAL::UNAVAILABLE`       | 503  | AI extraction is not available yet (no extraction model is configured). Refused before the fetch; 0 credits. | Use `extract` or `autoparse`. Retrying will not help until it launches.                                                    |
| `ERR::EXTRACT::FAILED`             | 502  | The model could not produce data from this page, or the target is not HTML. 0 credits.                       | Check the page has the data (try `response_format=markdown`), add `js_render=true` if it was empty, or tighten the prompt. |
| `ERR::REQUEST::INCOMPATIBLE_FLAGS` | 400  | Both `prompt` and `schema`.                                                                                  | Send one.                                                                                                                  |
| `ERR::REQUEST::INVALID_PARAMETER`  | 400  | Neither `prompt` nor `schema`, or an empty one.                                                              | Add a prompt or schema.                                                                                                    |
| `ERR::LIMIT::MAX_COST_EXCEEDED`    | 400  | Engine price plus 4 exceeds `max_cost`.                                                                      | Raise `max_cost`.                                                                                                          |

`ERR::INTERNAL::UNAVAILABLE` is marked `retryable: true` in the error table because it is also used for transient outages. For `ai_extract`, read the `detail`: "no extraction model is configured" means AI extraction has not launched, so do not retry.

## Cost

Engine price plus 4 credits per success: 5 on fetch, 7 with `js_render`, 12 on chromium. The 4 credits are the only additive surcharge in the price table. Failures, including a model failure, cost 0. See [Credits](https://docs.spicrawl.com/credits.md).

## Related

* [Structured data](https://docs.spicrawl.com/guides/structured-data.md) for selectors and autoparse
* [Markdown](https://docs.spicrawl.com/guides/markdown.md) to see what the model reads
* [JavaScript rendering](https://docs.spicrawl.com/guides/javascript-rendering.md)
