spicrawlspicrawlDocs

E-commerce price monitoring and product data

Price monitoring API cookbook: extract product title, price, availability and rating, re-scrape with cache false, batch catalogs, compare country prices.

Send each product URL to POST /v1/scrape with an extract schema and cache: false, and trigger it from your own scheduler. A plain-fetch check costs 1 credit, a rendered one 3, and a failure or bot challenge 0.

How do I scrape a product page into title, price, availability and rating?

Send the URL with extract set to a JSON Schema whose properties each carry a CSS selector; Spicrawl types each value ("$1,234.56" becomes 1234.56) and returns the record under data.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://shop.example.com/products/walnut-desk",
    "cache": false,
    "extract": {
      "type": "object",
      "properties": {
        "title":        {"type": "string", "selector": "h1"},
        "price":        {"type": "number", "selector": "[itemprop=price]", "attribute": "content"},
        "availability": {"type": "string", "selector": "[itemprop=availability]", "attribute": "href"},
        "rating":       {"type": "number", "selector": "[itemprop=ratingValue]", "attribute": "content"}
      },
      "required": ["title", "price"],
      "strict": true
    }
  }'
{
  "status": 200,
  "credits": 1,
  "data": {
    "title": "Walnut Desk",
    "price": 349,
    "availability": "https://schema.org/InStock",
    "rating": 4.6
  },
  "empty_fields": []
}

The selectors read schema.org microdata; swap in your store's. List must-have fields in required and set strict: true: a redesign then fails with ERR::EXTRACT::FAILED (not charged) instead of writing an empty price, and empty_fields flags optional fields that stopped matching. To skip selectors, set autoparse: true (0 credits) and read the page's JSON-LD or embedded app state from data.json_ld and data.embedded_state. See Structured data.

How do I re-scrape prices on a schedule?

Call POST /v1/scrape from your own scheduler (cron, CI or a queue worker) and store each record with a timestamp. This script keeps history in SQLite and prints changes:

import json, os, sqlite3, time, requests

API = "https://api.spicrawl.com/v1/scrape"
HEADERS = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
SCHEMA = json.load(open("product.json"))  # the schema from the first request

db = sqlite3.connect("prices.db")
db.execute("create table if not exists price (url text, at integer, price real, availability text, rating real)")

def check(url):
    for attempt in range(3):
        r = requests.post(API, headers=HEADERS, timeout=180,
                          json={"url": url, "cache": False, "max_cost": 3, "extract": SCHEMA})
        if r.status_code == 200:
            env = r.json()
            return env["data"] if env["status"] == 200 else None  # env["status"] is the store's status
        if r.status_code == 402:
            raise SystemExit("monthly credits used up")
        if not r.json().get("retryable"):
            return None  # for example ERR::EXTRACT::FAILED: the layout changed
        time.sleep(float(r.headers.get("Retry-After", 2 ** attempt)))
    return None

for url in (line.strip() for line in open("urls.txt") if line.strip()):
    d = check(url)
    if d is None:
        continue  # no observation is not a price of 0
    last = db.execute("select price from price where url = ? order by at desc limit 1", (url,)).fetchone()
    db.execute("insert into price values (?, ?, ?, ?, ?)",
               (url, int(time.time()), d["price"], d.get("availability"), d.get("rating")))
    if last and last[0] != d["price"]:
        print(f"{url}: {last[0]} -> {d['price']}")
    time.sleep(2)  # be gentle with the store
db.commit()

Record no observation, never a price of 0, when a request fails. Failures cost 0 credits, so retrying is free; a product that now returns 404 or 410 is billed as a scrape of a missing page, so track it as delisted. "max_cost": 3 stops a challenged page climbing to the 8-credit tier. Keep requests in flight at or below the Concurrency-Limit header, or excess requests get 429 at 0 credits. See Errors, Rate limits and Scripting with the CLI.

How do I get a fresh price, not a cached one?

Send "cache": false (CLI --no-cache) on every scheduled request. Otherwise a daily job can be handed yesterday's stored result (Cache-State: hit), billed like a live fetch; a bypass (Cache-State: bypass) costs nothing extra. See Caching.

How do I monitor a whole product catalog?

POST /v1/batch takes up to 10,000 URLs, runs them server-side with its own retries and keeps JSON Lines results for 72 hours; put your SKU in external_id. A batch item returns the page (HTML, markdown, text or JSON), so parse the price yourself: extract and autoparse are refused with a 400. Use /v1/scrape per URL for tens of SKUs or when you want extract.

curl https://api.spicrawl.com/v1/batch \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "catalog-2026-10-09",
    "credit_budget": 500,
    "items": [
      {"url": "https://shop.example.com/products/1", "external_id": "sku-1"},
      {"url": "https://shop.example.com/products/2", "external_id": "sku-2"}
    ]
  }'

Batch never reads the cache and never climbs through a bot challenge, so run protected stores through /v1/scrape with your own proxy. Add "js_render": true to the job for rendered stores (3 credits per item); retry failures with POST /v1/batch/{id}/retry. See Batch jobs.

How do I handle rendered stores and per-country prices?

How do I scrape stores that load prices with JavaScript?

When a plain fetch lacks the price (it appears in empty_fields, or a strict schema fails), add js_render: true and wait_for set to the price element. The page is captured once that element appears, for 3 credits instead of 1.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://shop.example.com/products/walnut-desk",
    "js_render": true,
    "wait_for": "[data-testid=price]",
    "wait_for_timeout": 15000,
    "cache": false,
    "extract": {"title": "h1", "price": "[data-testid=price]", "stock": ".stock-status"}
  }'
  • Cheaper than the page: many stores load the price from their own JSON endpoint; record it with "network_capture": {"urls": ["/api/price"]} (3 credits) and read it under network. See Network capture.
  • Clicks first: actions (a size swatch, "Load more") always run in full-browser rendering, 8 credits. See Browser actions.
  • Two traps: an explicit block_resources list also moves a request to 8 credits, and mode: "auto" checks max_cost against its top rung (25 credits). Use js_render when you know the store needs it. See JavaScript rendering.

How do I see a store's prices in another country?

Send the request through your own proxy that exits in that country (proxy) and add proxy_country with its two-letter code. The proxy sets the IP the store sees; on a rendered request proxy_country also matches the browser's timezone and language to the exit.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://shop.example.com/products/walnut-desk",
    "proxy": "http://user:pass@de.proxy.example.net:8000",
    "proxy_country": "de",
    "js_render": true,
    "cache": false,
    "extract": {
      "price": {"selector": "[itemprop=price]", "output": "attr", "attribute": "content"},
      "currency": {"selector": "[itemprop=priceCurrency]", "output": "attr", "attribute": "content"}
    }
  }'
  • Capture the currency with the price, and note whether it includes tax (VAT-inclusive in the EU, often tax-exclusive in the US).
  • Hold one exit per country while you compare, and add proxy_verify: true to fail fast on a dead proxy. Your own proxy adds no surcharge.
  • Managed country selection (premium_proxy with proxy_country) Coming soon is not available yet. See Proxies and geo.

Which retailers block, and what does it cost?

It depends on the store's bot protection and your exit. What the October 2026 tests showed:

  • Some retail product and search pages returned the full page on a plain fetch for 1 credit, through a proxy passed in proxy.
  • Some product pages returned price and an Add to Cart button on the default shared pool, with no proxy of your own.
  • Pages protected by Akamai or DataDome may return 502 ERR::UPSTREAM::CHALLENGE for 0 credits; detected block pages are not charged.

Verdicts change by exit and by day, so re-test a store before you depend on it. Spicrawl does not promise any site returns data on every request and does not solve CAPTCHAs.

What do I do when a retailer serves a challenge?

Treat it as no observation; a challenge is never billed.

  1. Retry once or twice, later. ERR::UPSTREAM::CHALLENGE is retryable (--retry 3 on the CLI).
  2. Use your own residential proxy in proxy. A challenge that names its vendor moves the request to full-browser rendering on the same exit: 8 credits if it serves the page, about 20 seconds, so allow a 180-second timeout and max_cost of 8 or more. This improves the odds without guaranteeing a page.
  3. Read the block: spicrawl logs --status blocked, then spicrawl logs get <request-id>, names the vendor and signal.
  4. Fall back to an official source (a product feed, affiliate feed or marketplace API) for a store that keeps blocking.

See Anti-bot and Cloudflare-protected pages.

How much does price monitoring cost?

Checks multiplied by credits multiplied by days. The beta allowance is 1,000 credits a month per organization:

CheckCreditsProducts checked daily on 1,000 credits a month
Plain fetch with extract or autoparse133
js_render with extract311
Full-browser rendering (actions, or a challenge climb that served the page)84

extract, autoparse, network_capture and your own proxy add 0, and a cache hit is billed like a live fetch, so the cache does not lower the bill. Cap spend with max_cost per request and credit_budget per batch job; read X-Credits-Charged and X-Credits-Remaining on each response. See Credits.

Frequently asked questions

Start scraping in minutes

1,000 free credits every month. One API key, one request.

For educational purposes

The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and robots.txt, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.

On this page