E-commerce price monitoring and product data
Price monitoring API cookbook: extract product title, price, availability and rating, re-scrape with cache false, batch catalogs, compare country prices.
Send each product URL to POST /v1/scrape with an extract schema and cache: false, and trigger it from your own scheduler. A plain-fetch check costs 1 credit, a rendered one 3, and a failure or bot challenge 0.
How do I scrape a product page into title, price, availability and rating?
Send the URL with extract set to a JSON Schema whose properties each carry a CSS selector; Spicrawl types each value ("$1,234.56" becomes 1234.56) and returns the record under data.
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://shop.example.com/products/walnut-desk",
"cache": false,
"extract": {
"type": "object",
"properties": {
"title": {"type": "string", "selector": "h1"},
"price": {"type": "number", "selector": "[itemprop=price]", "attribute": "content"},
"availability": {"type": "string", "selector": "[itemprop=availability]", "attribute": "href"},
"rating": {"type": "number", "selector": "[itemprop=ratingValue]", "attribute": "content"}
},
"required": ["title", "price"],
"strict": true
}
}'{
"status": 200,
"credits": 1,
"data": {
"title": "Walnut Desk",
"price": 349,
"availability": "https://schema.org/InStock",
"rating": 4.6
},
"empty_fields": []
}The selectors read schema.org microdata; swap in your store's. List must-have fields in required and set strict: true: a redesign then fails with ERR::EXTRACT::FAILED (not charged) instead of writing an empty price, and empty_fields flags optional fields that stopped matching. To skip selectors, set autoparse: true (0 credits) and read the page's JSON-LD or embedded app state from data.json_ld and data.embedded_state. See Structured data.
How do I re-scrape prices on a schedule?
Call POST /v1/scrape from your own scheduler (cron, CI or a queue worker) and store each record with a timestamp. This script keeps history in SQLite and prints changes:
import json, os, sqlite3, time, requests
API = "https://api.spicrawl.com/v1/scrape"
HEADERS = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
SCHEMA = json.load(open("product.json")) # the schema from the first request
db = sqlite3.connect("prices.db")
db.execute("create table if not exists price (url text, at integer, price real, availability text, rating real)")
def check(url):
for attempt in range(3):
r = requests.post(API, headers=HEADERS, timeout=180,
json={"url": url, "cache": False, "max_cost": 3, "extract": SCHEMA})
if r.status_code == 200:
env = r.json()
return env["data"] if env["status"] == 200 else None # env["status"] is the store's status
if r.status_code == 402:
raise SystemExit("monthly credits used up")
if not r.json().get("retryable"):
return None # for example ERR::EXTRACT::FAILED: the layout changed
time.sleep(float(r.headers.get("Retry-After", 2 ** attempt)))
return None
for url in (line.strip() for line in open("urls.txt") if line.strip()):
d = check(url)
if d is None:
continue # no observation is not a price of 0
last = db.execute("select price from price where url = ? order by at desc limit 1", (url,)).fetchone()
db.execute("insert into price values (?, ?, ?, ?, ?)",
(url, int(time.time()), d["price"], d.get("availability"), d.get("rating")))
if last and last[0] != d["price"]:
print(f"{url}: {last[0]} -> {d['price']}")
time.sleep(2) # be gentle with the store
db.commit()Record no observation, never a price of 0, when a request fails. Failures cost 0 credits, so retrying is free; a product that now returns 404 or 410 is billed as a scrape of a missing page, so track it as delisted. "max_cost": 3 stops a challenged page climbing to the 8-credit tier. Keep requests in flight at or below the Concurrency-Limit header, or excess requests get 429 at 0 credits. See Errors, Rate limits and Scripting with the CLI.
How do I get a fresh price, not a cached one?
Send "cache": false (CLI --no-cache) on every scheduled request. Otherwise a daily job can be handed yesterday's stored result (Cache-State: hit), billed like a live fetch; a bypass (Cache-State: bypass) costs nothing extra. See Caching.
How do I monitor a whole product catalog?
POST /v1/batch takes up to 10,000 URLs, runs them server-side with its own retries and keeps JSON Lines results for 72 hours; put your SKU in external_id. A batch item returns the page (HTML, markdown, text or JSON), so parse the price yourself: extract and autoparse are refused with a 400. Use /v1/scrape per URL for tens of SKUs or when you want extract.
curl https://api.spicrawl.com/v1/batch \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "catalog-2026-10-09",
"credit_budget": 500,
"items": [
{"url": "https://shop.example.com/products/1", "external_id": "sku-1"},
{"url": "https://shop.example.com/products/2", "external_id": "sku-2"}
]
}'Batch never reads the cache and never climbs through a bot challenge, so run protected stores through /v1/scrape with your own proxy. Add "js_render": true to the job for rendered stores (3 credits per item); retry failures with POST /v1/batch/{id}/retry. See Batch jobs.
How do I handle rendered stores and per-country prices?
How do I scrape stores that load prices with JavaScript?
When a plain fetch lacks the price (it appears in empty_fields, or a strict schema fails), add js_render: true and wait_for set to the price element. The page is captured once that element appears, for 3 credits instead of 1.
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://shop.example.com/products/walnut-desk",
"js_render": true,
"wait_for": "[data-testid=price]",
"wait_for_timeout": 15000,
"cache": false,
"extract": {"title": "h1", "price": "[data-testid=price]", "stock": ".stock-status"}
}'- Cheaper than the page: many stores load the price from their own JSON endpoint; record it with
"network_capture": {"urls": ["/api/price"]}(3 credits) and read it undernetwork. See Network capture. - Clicks first:
actions(a size swatch, "Load more") always run in full-browser rendering, 8 credits. See Browser actions. - Two traps: an explicit
block_resourceslist also moves a request to 8 credits, andmode: "auto"checksmax_costagainst its top rung (25 credits). Usejs_renderwhen you know the store needs it. See JavaScript rendering.
How do I see a store's prices in another country?
Send the request through your own proxy that exits in that country (proxy) and add proxy_country with its two-letter code. The proxy sets the IP the store sees; on a rendered request proxy_country also matches the browser's timezone and language to the exit.
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://shop.example.com/products/walnut-desk",
"proxy": "http://user:pass@de.proxy.example.net:8000",
"proxy_country": "de",
"js_render": true,
"cache": false,
"extract": {
"price": {"selector": "[itemprop=price]", "output": "attr", "attribute": "content"},
"currency": {"selector": "[itemprop=priceCurrency]", "output": "attr", "attribute": "content"}
}
}'- Capture the currency with the price, and note whether it includes tax (VAT-inclusive in the EU, often tax-exclusive in the US).
- Hold one exit per country while you compare, and add
proxy_verify: trueto fail fast on a dead proxy. Your own proxy adds no surcharge. - Managed country selection (
premium_proxywithproxy_country) Coming soon is not available yet. See Proxies and geo.
Which retailers block, and what does it cost?
It depends on the store's bot protection and your exit. What the October 2026 tests showed:
- Some retail product and search pages returned the full page on a plain fetch for 1 credit, through a proxy passed in
proxy. - Some product pages returned price and an Add to Cart button on the default shared pool, with no proxy of your own.
- Pages protected by Akamai or DataDome may return
502 ERR::UPSTREAM::CHALLENGEfor 0 credits; detected block pages are not charged.
Verdicts change by exit and by day, so re-test a store before you depend on it. Spicrawl does not promise any site returns data on every request and does not solve CAPTCHAs.
What do I do when a retailer serves a challenge?
Treat it as no observation; a challenge is never billed.
- Retry once or twice, later.
ERR::UPSTREAM::CHALLENGEis retryable (--retry 3on the CLI). - Use your own residential proxy in
proxy. A challenge that names its vendor moves the request to full-browser rendering on the same exit: 8 credits if it serves the page, about 20 seconds, so allow a 180-second timeout andmax_costof 8 or more. This improves the odds without guaranteeing a page. - Read the block:
spicrawl logs --status blocked, thenspicrawl logs get <request-id>, names the vendor and signal. - Fall back to an official source (a product feed, affiliate feed or marketplace API) for a store that keeps blocking.
See Anti-bot and Cloudflare-protected pages.
How much does price monitoring cost?
Checks multiplied by credits multiplied by days. The beta allowance is 1,000 credits a month per organization:
| Check | Credits | Products checked daily on 1,000 credits a month |
|---|---|---|
Plain fetch with extract or autoparse | 1 | 33 |
js_render with extract | 3 | 11 |
Full-browser rendering (actions, or a challenge climb that served the page) | 8 | 4 |
extract, autoparse, network_capture and your own proxy add 0, and a cache hit is billed like a live fetch, so the cache does not lower the bill. Cap spend with max_cost per request and credit_budget per batch job; read X-Credits-Charged and X-Credits-Remaining on each response. See Credits.
Frequently asked questions
Send each product URL to POST /v1/scrape with an extract schema and cache: false from your own scheduler, store every returned data record with a timestamp, and compare it with the last one. A plain-fetch check costs 1 credit; see the schedule recipe.
Sometimes. In October 2026 tests some retail product pages returned their price on the default shared pool with no proxy of your own, but results vary by store, category, region and moment, and a challenge returns ERR::UPSTREAM::CHALLENGE at 0 credits. Check the store's terms and prefer an official API or product feed where it covers your need; see the retail results.
The result cache is on by default with a 48-hour window, so a repeat request can return a stored result. Send "cache": false (CLI --no-cache) on every scheduled request; it costs nothing extra. See fresh prices.
No. Spicrawl returns a fresh record per request; you run the schedule, store the records and compare them. Built-in scheduled scrapes with snapshot diffing are not available yet.
Not with extract: a batch item returns the page as HTML, markdown, text or JSON, and extract and autoparse are refused with a 400 before anything is charged. Parse the price from the returned HTML; see the catalog recipe.
Some pages protected by Akamai or DataDome may block; a detected block page comes back as ERR::UPSTREAM::CHALLENGE and is not charged. In October 2026 tests some retail pages returned real content on a plain fetch; see what to do about a challenge.
Related
- Structured data, Caching and Batch jobs
- JavaScript rendering and Network capture
- Proxies and geo and Anti-bot
- Scripting with the CLI, Credits, Rate limits and Errors
Start scraping in minutes
1,000 free credits every month. One API key, one request.
For educational purposes
The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and robots.txt, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.