Extract structured data
Get JSON out of a page three ways: autoparse for the page's own metadata, a selector map for exact fields, or a JSON Schema with selectors for typed, validated output.
Use this when you want fields (a title, a price, a list of reviews) rather than the page text. There are three tiers, cheapest to write first, and none of them costs extra credits:
autoparse: truereturns what the page already says about itself: JSON-LD, OpenGraph, Twitter Card, microdata, RDFa,<head>metadata and embedded SPA state. No selectors.extractas a selector map returns exactly the fields you name, as strings.extractas a JSON Schema whose properties carry aselector, which also coerces values to the declared types ("$1,234.56"becomes1234.56) and validates the result.
All three write to data in the JSON envelope. For pages where selectors are impractical, AI extraction Coming soon will let a model read the page.
Minimal request
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/products/42",
"extract": {
"title": "h1",
"price": ".price",
"images": {"selector": ".gallery img", "kind": "list", "output": "attr", "attribute": "src"},
"next_page": "a.next@href"
}
}'What comes back
The JSON envelope, with your fields under data:
{
"url": "https://example.com/products/42",
"final_url": "https://example.com/products/42",
"status": 200,
"content": "<html>...</html>",
"credits": 1,
"engine": "fetch",
"proxy_source": "pool",
"warnings": [],
"data": {
"title": "Walnut Desk",
"price": "$349.00",
"images": ["https://example.com/img/desk-1.jpg", "https://example.com/img/desk-2.jpg"]
},
"empty_fields": ["next_page"]
}empty_fields lists rules that matched nothing. It is the first sign that a site changed its markup: alert on it rather than on an empty value. extract forces the envelope; if you also asked for markdown, content holds the markdown and the response carries X-Warning: FORMAT_COERCED.
Tier 1: autoparse
Send "autoparse": true (CLI: --autoparse). data then holds only the keys whose source exists on the page:
| Key | Contents |
|---|---|
json_ld | Every JSON-LD block, @graph flattened, as an array. |
open_graph | og:* properties, with og:image as structured objects. |
twitter | twitter:* card properties. |
microdata | itemscope items with type and properties. |
rdfa | RDFa items. |
meta | title, description, canonical, favicon, language. |
embedded_state | Hydration payloads, keyed by source: next (__NEXT_DATA__), nuxt, angular, apollo, redux, remix, run_params, data_layer (GA4). |
URLs are resolved to absolute. embedded_state is often the richest source (variants, stock, regional prices), but Next.js App Router pages are not covered: they ship their data as self.__next_f flight chunks, not __NEXT_DATA__, so they yield no embedded_state. Use a selector map or network capture for those.
Tier 2: selector map
Each field is a selector string or a rule object.
"title": "h1"returns the cleaned text of the first match."next_page": "a.next@href"returns thehrefattribute of the first match.- A rule object takes
selector(required),type(cssorxpath),kind(itemfor the first match,listfor all),output(text,html,outer_html,attr,table_json,table_array,exists,count),attribute(required withoutput: "attr", forbidden otherwise),clean(defaulttrue) andnested. nestedis a selector map evaluated inside each match. It requiresoutputhtmlorouter_html.
{
"reviews": {
"selector": ".review", "kind": "list", "output": "outer_html",
"nested": { "author": ".author", "rating": ".stars@data-rating", "body": "p" }
}
}Tier 3: JSON Schema with selectors
When extract has "type": "object" and a properties object, it is read as a JSON Schema. Every property (or its items, or its nested properties) needs a selector; selector_type (css or xpath) and attribute are also accepted.
{
"type": "object",
"properties": {
"title": { "type": "string", "selector": "h1" },
"price": { "type": "number", "selector": ".price" },
"sku": { "type": "string", "selector": ".sku", "attribute": "data-sku" },
"tags": { "type": "array", "selector": ".tag", "items": { "type": "string" } }
},
"required": ["title", "price"],
"strict": false
}Values are coerced to the declared type: "$1,234.56" becomes 1234.56 and "1.234,56" becomes 1234.56. Properties that match nothing are dropped and listed in empty_fields. With strict: false (the default) a field that fails validation is dropped with a warning; with strict: true the whole extraction fails.
Limits
| Limit | Value |
|---|---|
Fields per level (selector map or properties) | 1–100 |
nested depth | 5 |
extract: {} | Refused (400). Omit the field instead. |
Failure modes
| Code | HTTP | When | What to do |
|---|---|---|---|
ERR::REQUEST::INVALID_PARAMETER | 400 | A malformed rule: over 100 fields, depth above 5, attribute without output: "attr", a schema property with no selector. Checked before fetching; 0 credits. | Fix the rule at the path the detail names. |
ERR::EXTRACT::INVALID_RULES | 400 | extract_preset was set. Extraction presets Coming soon: until they launch, every value fails (after the fetch, at 0 credits). | Put the rules inline in extract. |
ERR::REQUEST::INCOMPATIBLE_FLAGS | 400 | Both extract and extract_preset. | Send only extract. |
ERR::EXTRACT::FAILED | 502 | The page could not be extracted: the target is a PDF, an image or other non-HTML body, or a strict schema failed validation. 0 credits. | For a PDF use response_format=markdown; otherwise check the target and the schema. |
A broken selector is not an error: it returns HTTP 200 with the field in empty_fields.
Cost
extract and autoparse add 0 credits. You pay the engine and proxy tier (1 credit on fetch, 3 with js_render). Failures cost 0. See Credits.
Related
- AI extraction Coming soon
- Network capture for data the page loads over XHR
- JavaScript rendering when selectors match nothing on a client-rendered page
- Batch. Batch items return the raw page today; run extraction per URL with
/v1/scrape.
Proxies and geo
Route requests through your own proxy with proxy and proxy_verify, and read how a request was routed. Spicrawl's managed proxy pool is coming soon.
AI extraction Coming Soon
Coming soon: describe the data in plain language or as a JSON Schema with ai_extract and get it back under data, for 4 credits on top of the engine price.