Extract data with a model Coming soon
Coming soon: describe the data in plain language or as a JSON Schema with ai_extract and get it back under data, for 4 credits on top of the engine price.
Coming soon
AI extraction is not available yet. Until it launches, a request with ai_extract is refused before the fetch, at 0 credits. To get structured data today, use selectors or autoparse. This page describes how ai_extract will work.
Use this when writing selectors is impractical: layouts differ from page to page, the value is buried in prose ("ships in 3–5 business days"), or you need the same fields from many unrelated sites. ai_extract gives the page to a model as cleaned markdown and returns what it extracted under data. It costs 4 credits on top of the engine price, so prefer selectors or autoparse for a site you scrape repeatedly with stable markup.
Minimal request
Send exactly one of prompt or schema.
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/products/42",
"ai_extract": {"prompt": "Product name, price with currency, and delivery time in days"}
}'To fix the output shape, send a JSON Schema instead of a prompt (CLI: --ai-schema @schema.json):
{
"url": "https://example.com/products/42",
"ai_extract": {
"schema": {
"type": "object",
"properties": {
"name": { "type": "string" },
"price": { "type": "number" },
"currency": { "type": "string" },
"delivery_days_max": { "type": "integer" }
},
"required": ["name", "price"]
}
}
}The schema needs no selector keywords; that is the difference from the schema form of extract.
What comes back
The JSON envelope, with the model's output under data:
{
"url": "https://example.com/products/42",
"final_url": "https://example.com/products/42",
"status": 200,
"content": "<html>...</html>",
"credits": 5,
"engine": "fetch",
"proxy_source": "direct",
"warnings": [],
"data": {
"name": "Walnut Desk",
"price": 349,
"currency": "USD",
"delivery_days_max": 5
}
}credits is the engine price plus 4 (here 1 + 4). ai_extract always forces the envelope; if you asked for markdown, content holds it and the response carries X-Warning: FORMAT_COERCED.
Options that matter
| Field | Rule |
|---|---|
ai_extract.prompt | Non-empty plain-language description. Name units and formats you want ("price as a number, currency as ISO 4217"). |
ai_extract.schema | JSON Schema object with at least one property. Use it when downstream code parses the result. |
main_content_only, include_tags, exclude_tags | Shape the markdown the model reads. Scoping to the relevant container improves accuracy on busy pages. See Markdown. |
js_render | Needed when the data is built by JavaScript; the model only sees what the engine retrieved. |
Both prompt and schema is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS; neither is 400 ERR::REQUEST::INVALID_PARAMETER.
When to prefer selectors
| Situation | Use |
|---|---|
| One site, stable markup, many pages | extract selector map or schema: 0 extra credits and deterministic. |
| The page has JSON-LD or embedded state | autoparse: 0 extra credits. |
| Many unrelated sites, one target shape | ai_extract with a schema. |
| Values inside free text | ai_extract. |
You can combine them: run ai_extract once to find where the data lives, then write selectors for the site.
Failure modes
| Code | HTTP | When | What to do |
|---|---|---|---|
ERR::INTERNAL::UNAVAILABLE | 503 | AI extraction is not available yet (no extraction model is configured). Refused before the fetch; 0 credits. | Use extract or autoparse. Retrying will not help until it launches. |
ERR::EXTRACT::FAILED | 502 | The model could not produce data from this page, or the target is not HTML. 0 credits. | Check the page has the data (try response_format=markdown), add js_render=true if it was empty, or tighten the prompt. |
ERR::REQUEST::INCOMPATIBLE_FLAGS | 400 | Both prompt and schema. | Send one. |
ERR::REQUEST::INVALID_PARAMETER | 400 | Neither prompt nor schema, or an empty one. | Add a prompt or schema. |
ERR::LIMIT::MAX_COST_EXCEEDED | 400 | Engine price plus 4 exceeds max_cost. | Raise max_cost. |
ERR::INTERNAL::UNAVAILABLE is marked retryable: true in the error table because it is also used for transient outages. For ai_extract, read the detail: "no extraction model is configured" means AI extraction has not launched, so do not retry.
Cost
Engine price plus 4 credits per success: 5 on fetch, 7 with js_render, 12 on chromium. The 4 credits are the only additive surcharge in the price table. Failures, including a model failure, cost 0. See Credits.
Related
- Structured data for selectors and autoparse
- Markdown to see what the model reads
- JavaScript rendering
Structured data
Get JSON out of a page three ways: autoparse for the page's own metadata, a selector map for exact fields, or a JSON Schema with selectors for typed, validated output.
Network capture
Record the XHR and fetch responses a page makes while it renders with network_capture, and read their JSON bodies from the network array.