# E-commerce price monitoring and product data

> Price monitoring API cookbook: extract product title, price, availability and rating, re-scrape with cache false, batch catalogs, compare country prices.

Source: https://docs.spicrawl.com/use-cases/price-monitoring

Send each product URL to `POST /v1/scrape` with an `extract` schema and `cache: false`, and trigger it from your own scheduler. A plain-fetch check costs 1 credit, a rendered one 3, and a failure or bot challenge 0.

 

## How do I scrape a product page into title, price, availability and rating?

Send the URL with `extract` set to a JSON Schema whose properties each carry a CSS `selector`; Spicrawl types each value (`"$1,234.56"` becomes `1234.56`) and returns the record under `data`.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://shop.example.com/products/walnut-desk",
    "cache": false,
    "extract": {
      "type": "object",
      "properties": {
        "title":        {"type": "string", "selector": "h1"},
        "price":        {"type": "number", "selector": "[itemprop=price]", "attribute": "content"},
        "availability": {"type": "string", "selector": "[itemprop=availability]", "attribute": "href"},
        "rating":       {"type": "number", "selector": "[itemprop=ratingValue]", "attribute": "content"}
      },
      "required": ["title", "price"],
      "strict": true
    }
  }'
```

```python title="Python"
import os, requests

schema = {
    "type": "object",
    "properties": {
        "title": {"type": "string", "selector": "h1"},
        "price": {"type": "number", "selector": "[itemprop=price]", "attribute": "content"},
        "availability": {"type": "string", "selector": "[itemprop=availability]", "attribute": "href"},
        "rating": {"type": "number", "selector": "[itemprop=ratingValue]", "attribute": "content"},
    },
    "required": ["title", "price"],
    "strict": True,
}

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={"url": "https://shop.example.com/products/walnut-desk", "cache": False, "extract": schema},
    timeout=120,
)
r.raise_for_status()
env = r.json()
print(env["data"], env.get("empty_fields"))
```

```typescript title="TypeScript"
const schema = {
  type: "object",
  properties: {
    title: { type: "string", selector: "h1" },
    price: { type: "number", selector: "[itemprop=price]", attribute: "content" },
    availability: { type: "string", selector: "[itemprop=availability]", attribute: "href" },
    rating: { type: "number", selector: "[itemprop=ratingValue]", attribute: "content" },
  },
  required: ["title", "price"],
  strict: true,
};

const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url: "https://shop.example.com/products/walnut-desk", cache: false, extract: schema }),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
const { data, empty_fields } = await r.json();
```

```bash title="CLI"
# product.json holds the same schema as the curl tab
spicrawl scrape https://shop.example.com/products/walnut-desk --extract @product.json --no-cache
```

```json
{
  "status": 200,
  "credits": 1,
  "data": {
    "title": "Walnut Desk",
    "price": 349,
    "availability": "https://schema.org/InStock",
    "rating": 4.6
  },
  "empty_fields": []
}
```

The selectors read schema.org microdata; swap in your store's. List must-have fields in `required` and set `strict: true`: a redesign then fails with `ERR::EXTRACT::FAILED` (not charged) instead of writing an empty price, and `empty_fields` flags optional fields that stopped matching. To skip selectors, set `autoparse: true` (0 credits) and read the page's JSON-LD or embedded app state from `data.json_ld` and `data.embedded_state`. See [Structured data](https://docs.spicrawl.com/guides/structured-data.md).

## How do I re-scrape prices on a schedule?

Call `POST /v1/scrape` from your own scheduler (cron, CI or a queue worker) and store each record with a timestamp. This script keeps history in SQLite and prints changes:

```python title="Python"
import json, os, sqlite3, time, requests

API = "https://api.spicrawl.com/v1/scrape"
HEADERS = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
SCHEMA = json.load(open("product.json"))  # the schema from the first request

db = sqlite3.connect("prices.db")
db.execute("create table if not exists price (url text, at integer, price real, availability text, rating real)")

def check(url):
    for attempt in range(3):
        r = requests.post(API, headers=HEADERS, timeout=180,
                          json={"url": url, "cache": False, "max_cost": 3, "extract": SCHEMA})
        if r.status_code == 200:
            env = r.json()
            return env["data"] if env["status"] == 200 else None  # env["status"] is the store's status
        if r.status_code == 402:
            raise SystemExit("monthly credits used up")
        if not r.json().get("retryable"):
            return None  # for example ERR::EXTRACT::FAILED: the layout changed
        time.sleep(float(r.headers.get("Retry-After", 2 ** attempt)))
    return None

for url in (line.strip() for line in open("urls.txt") if line.strip()):
    d = check(url)
    if d is None:
        continue  # no observation is not a price of 0
    last = db.execute("select price from price where url = ? order by at desc limit 1", (url,)).fetchone()
    db.execute("insert into price values (?, ?, ?, ?, ?)",
               (url, int(time.time()), d["price"], d.get("availability"), d.get("rating")))
    if last and last[0] != d["price"]:
        print(f"{url}: {last[0]} -> {d['price']}")
    time.sleep(2)  # be gentle with the store
db.commit()
```

```bash title="CLI + cron"
# urls.txt: one product URL per line. product.json: the schema. Export SPICRAWL_API_KEY in the cron environment.
# Every 6 hours, 4 at a time, one JSON line per URL appended to prices.jsonl:
0 */6 * * * cd /srv/pricewatch && spicrawl scrape - --extract @product.json --no-cache --max-cost 3 --concurrency 4 < urls.txt >> prices.jsonl
```

Record no observation, never a price of 0, when a request fails. Failures cost 0 credits, so retrying is free; a product that now returns 404 or 410 is billed as a scrape of a missing page, so track it as delisted. `"max_cost": 3` stops a challenged page climbing to the 8-credit tier. Keep requests in flight at or below the `Concurrency-Limit` header, or excess requests get `429` at 0 credits. See [Errors](https://docs.spicrawl.com/errors.md), [Rate limits](https://docs.spicrawl.com/rate-limits.md) and [Scripting with the CLI](https://docs.spicrawl.com/cli/scripting.md).

### How do I get a fresh price, not a cached one?

Send `"cache": false` (CLI `--no-cache`) on every scheduled request. Otherwise a daily job can be handed yesterday's stored result (`Cache-State: hit`), billed like a live fetch; a bypass (`Cache-State: bypass`) costs nothing extra. See [Caching](https://docs.spicrawl.com/guides/caching.md).

### How do I monitor a whole product catalog?

`POST /v1/batch` takes up to 10,000 URLs, runs them server-side with its own retries and keeps JSON Lines results for 72 hours; put your SKU in `external_id`. A batch item returns the page (HTML, markdown, text or JSON), so parse the price yourself: `extract` and `autoparse` are refused with a `400`. Use `/v1/scrape` per URL for tens of SKUs or when you want `extract`.

```bash title="curl"
curl https://api.spicrawl.com/v1/batch \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "catalog-2026-10-09",
    "credit_budget": 500,
    "items": [
      {"url": "https://shop.example.com/products/1", "external_id": "sku-1"},
      {"url": "https://shop.example.com/products/2", "external_id": "sku-2"}
    ]
  }'
```

```python title="Python"
import json, os, re, time, requests

API = "https://api.spicrawl.com"
H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
LD = re.compile(r'<script[^>]+application/ld\+json[^>]*>(.*?)</script>', re.S)

def offer(html):
    for block in LD.findall(html):
        try:
            node = json.loads(block)
        except ValueError:
            continue
        for n in node if isinstance(node, list) else [node]:
            if n.get("@type") == "Product":  # top-level Product only; handle @graph if your store uses it
                o = n.get("offers") or {}
                o = o[0] if isinstance(o, list) else o
                return {"name": n.get("name"), "price": o.get("price"), "currency": o.get("priceCurrency"),
                        "availability": o.get("availability"),
                        "rating": (n.get("aggregateRating") or {}).get("ratingValue")}

catalog = {"sku-1": "https://shop.example.com/products/1", "sku-2": "https://shop.example.com/products/2"}
job = requests.post(f"{API}/v1/batch", headers=H, timeout=60, json={
    "name": "catalog-2026-10-09",
    "credit_budget": 500,
    "items": [{"url": url, "external_id": sku} for sku, url in catalog.items()],
}).json()

while job["status"] not in ("completed", "failed", "cancelled"):
    time.sleep(10)
    job = requests.get(f"{API}/v1/batch/{job['id']}", headers=H, timeout=30).json()

cursor = None
while True:
    r = requests.get(f"{API}/v1/batch/{job['id']}/results", headers=H, timeout=120,
                     params={"cursor": cursor} if cursor else {})
    r.raise_for_status()
    for line in r.text.splitlines():
        item = json.loads(line)
        content = item.get("result", {}).get("content")
        if item["status"] == "succeeded" and item["http_status"] == 200 and content:
            print(item["external_id"], offer(json.loads(content)))  # content is a JSON string literal
    cursor = r.headers.get("X-Next-Cursor")
    if not cursor:
        break
```

```bash title="CLI"
# items.jsonl: one {"url": "...", "external_id": "sku-1"} object per line
spicrawl batch submit items.jsonl --name catalog-2026-10-09 --credit-budget 500 --wait --max-wait 30m
spicrawl batch results 01J9Z6T3W9E21T5TZARVJRVN5C --all -o results.jsonl
```

Batch never reads the cache and never climbs through a bot challenge, so run protected stores through `/v1/scrape` with your own proxy. Add `"js_render": true` to the job for rendered stores (3 credits per item); retry failures with `POST /v1/batch/{id}/retry`. See [Batch jobs](https://docs.spicrawl.com/guides/batch.md).

## How do I handle rendered stores and per-country prices?

### How do I scrape stores that load prices with JavaScript?

When a plain fetch lacks the price (it appears in `empty_fields`, or a `strict` schema fails), add `js_render: true` and `wait_for` set to the price element. The page is captured once that element appears, for 3 credits instead of 1.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://shop.example.com/products/walnut-desk",
    "js_render": true,
    "wait_for": "[data-testid=price]",
    "wait_for_timeout": 15000,
    "cache": false,
    "extract": {"title": "h1", "price": "[data-testid=price]", "stock": ".stock-status"}
  }'
```

```bash title="CLI"
spicrawl scrape https://shop.example.com/products/walnut-desk --render \
  --wait-for '[data-testid=price]' --wait-for-timeout 15000 --no-cache \
  --extract '{"title":"h1","price":"[data-testid=price]","stock":".stock-status"}'
```

* **Cheaper than the page:** many stores load the price from their own JSON endpoint; record it with `"network_capture": {"urls": ["/api/price"]}` (3 credits) and read it under `network`. See [Network capture](https://docs.spicrawl.com/guides/network-capture.md).
* **Clicks first:** `actions` (a size swatch, "Load more") always run in full-browser rendering, 8 credits. See [Browser actions](https://docs.spicrawl.com/guides/browser-actions.md).
* **Two traps:** an explicit `block_resources` list also moves a request to 8 credits, and `mode: "auto"` checks `max_cost` against its top rung (25 credits). Use `js_render` when you know the store needs it. See [JavaScript rendering](https://docs.spicrawl.com/guides/javascript-rendering.md).

### How do I see a store's prices in another country?

Send the request through your own proxy that exits in that country (`proxy`) and add `proxy_country` with its two-letter code. The proxy sets the IP the store sees; on a rendered request `proxy_country` also matches the browser's timezone and language to the exit.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://shop.example.com/products/walnut-desk",
    "proxy": "http://user:pass@de.proxy.example.net:8000",
    "proxy_country": "de",
    "js_render": true,
    "cache": false,
    "extract": {
      "price": {"selector": "[itemprop=price]", "output": "attr", "attribute": "content"},
      "currency": {"selector": "[itemprop=priceCurrency]", "output": "attr", "attribute": "content"}
    }
  }'
```

```python title="Python"
import os, requests

PROXIES = {"us": os.environ["PROXY_US"], "de": os.environ["PROXY_DE"], "gb": os.environ["PROXY_GB"]}
RULES = {
    "price": {"selector": "[itemprop=price]", "output": "attr", "attribute": "content"},
    "currency": {"selector": "[itemprop=priceCurrency]", "output": "attr", "attribute": "content"},
}

for country, proxy in PROXIES.items():
    r = requests.post(
        "https://api.spicrawl.com/v1/scrape",
        headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
        json={"url": "https://shop.example.com/products/walnut-desk", "proxy": proxy, "proxy_country": country,
              "js_render": True, "cache": False, "extract": RULES},
        timeout=180,
    )
    r.raise_for_status()
    print(country, r.json()["data"])
```

```bash title="CLI"
spicrawl scrape https://shop.example.com/products/walnut-desk --render --no-cache \
  --proxy "$PROXY_DE" --body '{"proxy_country":"de"}' \
  --extract '{"price":{"selector":"[itemprop=price]","output":"attr","attribute":"content"}}'
```

* **Capture the currency** with the price, and note whether it includes tax (VAT-inclusive in the EU, often tax-exclusive in the US).
* **Hold one exit per country** while you compare, and add `proxy_verify: true` to fail fast on a dead proxy. Your own proxy adds no surcharge.
* **Managed country selection** (`premium_proxy` with `proxy_country`) (coming soon) is not available yet. See [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md).

## Which retailers block, and what does it cost?

It depends on the store's bot protection and your exit. What the October 2026 tests showed:

* Some retail product and search pages returned the full page on a plain fetch for 1 credit, through a proxy passed in `proxy`.
* Some product pages returned price and an Add to Cart button on the default shared pool, with no proxy of your own.
* Pages protected by Akamai or DataDome may return `502 ERR::UPSTREAM::CHALLENGE` for 0 credits; detected block pages are not charged.

Verdicts change by exit and by day, so re-test a store before you depend on it. Spicrawl does not promise any site returns data on every request and does not solve CAPTCHAs.

### What do I do when a retailer serves a challenge?

Treat it as no observation; a challenge is never billed.

1. **Retry once or twice, later.** `ERR::UPSTREAM::CHALLENGE` is retryable (`--retry 3` on the CLI).
2. **Use your own residential proxy** in `proxy`. A challenge that names its vendor moves the request to full-browser rendering on the same exit: 8 credits if it serves the page, about 20 seconds, so allow a 180-second timeout and `max_cost` of 8 or more. This improves the odds without guaranteeing a page.
3. **Read the block:** `spicrawl logs --status blocked`, then `spicrawl logs get <request-id>`, names the vendor and signal.
4. **Fall back to an official source** (a product feed, affiliate feed or marketplace API) for a store that keeps blocking.

See [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md) and [Cloudflare-protected pages](https://docs.spicrawl.com/guides/cloudflare.md).

### How much does price monitoring cost?

Checks multiplied by credits multiplied by days. The beta allowance is 1,000 credits a month per organization:

| Check                                                                         | Credits | Products checked daily on 1,000 credits a month |
| ----------------------------------------------------------------------------- | ------- | ----------------------------------------------- |
| Plain fetch with `extract` or `autoparse`                                     | 1       | 33                                              |
| `js_render` with `extract`                                                    | 3       | 11                                              |
| Full-browser rendering (`actions`, or a challenge climb that served the page) | 8       | 4                                               |

`extract`, `autoparse`, `network_capture` and your own proxy add 0, and a cache hit is billed like a live fetch, so the cache does not lower the bill. Cap spend with `max_cost` per request and `credit_budget` per batch job; read `X-Credits-Charged` and `X-Credits-Remaining` on each response. See [Credits](https://docs.spicrawl.com/credits.md).

## Frequently asked questions

**How do I monitor competitor prices with an API?**

Send each product URL to `POST /v1/scrape` with an `extract` schema and `cache: false` from your own scheduler, store every returned `data` record with a timestamp, and compare it with the last one. A plain-fetch check costs 1 credit; see [the schedule recipe](#how-do-i-re-scrape-prices-on-a-schedule).

**Can I scrape product data from large retailers?**

Sometimes. In October 2026 tests some retail product pages returned their price on the default shared pool with no proxy of your own, but results vary by store, category, region and moment, and a challenge returns `ERR::UPSTREAM::CHALLENGE` at 0 credits. Check the store's terms and prefer an official API or product feed where it covers your need; see [the retail results](#which-retailers-block-and-what-does-it-cost).

**Why does my scheduled scrape return yesterday's price?**

The result cache is on by default with a 48-hour window, so a repeat request can return a stored result. Send `"cache": false` (CLI `--no-cache`) on every scheduled request; it costs nothing extra. See [fresh prices](#how-do-i-get-a-fresh-price-not-a-cached-one).

**Does Spicrawl send price-drop alerts or keep price history?**

No. Spicrawl returns a fresh record per request; you run the schedule, store the records and compare them. Built-in scheduled scrapes with snapshot diffing are not available yet.

**Can I extract prices from a batch job?**

Not with `extract`: a batch item returns the page as HTML, markdown, text or JSON, and `extract` and `autoparse` are refused with a `400` before anything is charged. Parse the price from the returned HTML; see [the catalog recipe](#how-do-i-monitor-a-whole-product-catalog).

**Which retailers block scraping?**

Some pages protected by Akamai or DataDome may block; a detected block page comes back as `ERR::UPSTREAM::CHALLENGE` and is not charged. In October 2026 tests some retail pages returned real content on a plain fetch; see [what to do about a challenge](#what-do-i-do-when-a-retailer-serves-a-challenge).

## Related

* [Structured data](https://docs.spicrawl.com/guides/structured-data.md), [Caching](https://docs.spicrawl.com/guides/caching.md) and [Batch jobs](https://docs.spicrawl.com/guides/batch.md)
* [JavaScript rendering](https://docs.spicrawl.com/guides/javascript-rendering.md) and [Network capture](https://docs.spicrawl.com/guides/network-capture.md)
* [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md) and [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md)
* [Scripting with the CLI](https://docs.spicrawl.com/cli/scripting.md), [Credits](https://docs.spicrawl.com/credits.md), [Rate limits](https://docs.spicrawl.com/rate-limits.md) and [Errors](https://docs.spicrawl.com/errors.md)

 

> **For educational purposes:** The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and `robots.txt`, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.
