spicrawlspicrawlDocs

Extract structured data

Get JSON out of a page three ways: autoparse for the page's own metadata, a selector map for exact fields, or a JSON Schema with selectors for typed, validated output.

Use this when you want fields (a title, a price, a list of reviews) rather than the page text. There are three tiers, cheapest to write first, and none of them costs extra credits:

  1. autoparse: true returns what the page already says about itself: JSON-LD, OpenGraph, Twitter Card, microdata, RDFa, <head> metadata and embedded SPA state. No selectors.
  2. extract as a selector map returns exactly the fields you name, as strings.
  3. extract as a JSON Schema whose properties carry a selector, which also coerces values to the declared types ("$1,234.56" becomes 1234.56) and validates the result.

All three write to data in the JSON envelope. For pages where selectors are impractical, AI extraction Coming soon will let a model read the page.

Minimal request

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/products/42",
    "extract": {
      "title": "h1",
      "price": ".price",
      "images": {"selector": ".gallery img", "kind": "list", "output": "attr", "attribute": "src"},
      "next_page": "a.next@href"
    }
  }'

What comes back

The JSON envelope, with your fields under data:

{
  "url": "https://example.com/products/42",
  "final_url": "https://example.com/products/42",
  "status": 200,
  "content": "<html>...</html>",
  "credits": 1,
  "engine": "fetch",
  "proxy_source": "pool",
  "warnings": [],
  "data": {
    "title": "Walnut Desk",
    "price": "$349.00",
    "images": ["https://example.com/img/desk-1.jpg", "https://example.com/img/desk-2.jpg"]
  },
  "empty_fields": ["next_page"]
}

empty_fields lists rules that matched nothing. It is the first sign that a site changed its markup: alert on it rather than on an empty value. extract forces the envelope; if you also asked for markdown, content holds the markdown and the response carries X-Warning: FORMAT_COERCED.

Tier 1: autoparse

Send "autoparse": true (CLI: --autoparse). data then holds only the keys whose source exists on the page:

KeyContents
json_ldEvery JSON-LD block, @graph flattened, as an array.
open_graphog:* properties, with og:image as structured objects.
twittertwitter:* card properties.
microdataitemscope items with type and properties.
rdfaRDFa items.
metatitle, description, canonical, favicon, language.
embedded_stateHydration payloads, keyed by source: next (__NEXT_DATA__), nuxt, angular, apollo, redux, remix, run_params, data_layer (GA4).

URLs are resolved to absolute. embedded_state is often the richest source (variants, stock, regional prices), but Next.js App Router pages are not covered: they ship their data as self.__next_f flight chunks, not __NEXT_DATA__, so they yield no embedded_state. Use a selector map or network capture for those.

Tier 2: selector map

Each field is a selector string or a rule object.

  • "title": "h1" returns the cleaned text of the first match.
  • "next_page": "a.next@href" returns the href attribute of the first match.
  • A rule object takes selector (required), type (css or xpath), kind (item for the first match, list for all), output (text, html, outer_html, attr, table_json, table_array, exists, count), attribute (required with output: "attr", forbidden otherwise), clean (default true) and nested.
  • nested is a selector map evaluated inside each match. It requires output html or outer_html.
{
  "reviews": {
    "selector": ".review", "kind": "list", "output": "outer_html",
    "nested": { "author": ".author", "rating": ".stars@data-rating", "body": "p" }
  }
}

Tier 3: JSON Schema with selectors

When extract has "type": "object" and a properties object, it is read as a JSON Schema. Every property (or its items, or its nested properties) needs a selector; selector_type (css or xpath) and attribute are also accepted.

{
  "type": "object",
  "properties": {
    "title":    { "type": "string",  "selector": "h1" },
    "price":    { "type": "number",  "selector": ".price" },
    "sku":      { "type": "string",  "selector": ".sku", "attribute": "data-sku" },
    "tags":     { "type": "array",   "selector": ".tag", "items": { "type": "string" } }
  },
  "required": ["title", "price"],
  "strict": false
}

Values are coerced to the declared type: "$1,234.56" becomes 1234.56 and "1.234,56" becomes 1234.56. Properties that match nothing are dropped and listed in empty_fields. With strict: false (the default) a field that fails validation is dropped with a warning; with strict: true the whole extraction fails.

Limits

LimitValue
Fields per level (selector map or properties)1–100
nested depth5
extract: {}Refused (400). Omit the field instead.

Failure modes

CodeHTTPWhenWhat to do
ERR::REQUEST::INVALID_PARAMETER400A malformed rule: over 100 fields, depth above 5, attribute without output: "attr", a schema property with no selector. Checked before fetching; 0 credits.Fix the rule at the path the detail names.
ERR::EXTRACT::INVALID_RULES400extract_preset was set. Extraction presets Coming soon: until they launch, every value fails (after the fetch, at 0 credits).Put the rules inline in extract.
ERR::REQUEST::INCOMPATIBLE_FLAGS400Both extract and extract_preset.Send only extract.
ERR::EXTRACT::FAILED502The page could not be extracted: the target is a PDF, an image or other non-HTML body, or a strict schema failed validation. 0 credits.For a PDF use response_format=markdown; otherwise check the target and the schema.

A broken selector is not an error: it returns HTTP 200 with the field in empty_fields.

Cost

extract and autoparse add 0 credits. You pay the engine and proxy tier (1 credit on fetch, 3 with js_render). Failures cost 0. See Credits.

On this page