Cache results and skip repeat charges
How the /v1/scrape result cache works: on by default with a 48-hour window, cache hits cost 0 credits, and cache=false or cache_ttl=0 forces a fresh fetch.
The result cache is on by default for /v1/scrape. When you repeat a request and a stored result is younger than cache_ttl seconds, you get that result back immediately, with Cache-State: hit and 0 credits charged. Use this page when you need fresher data than that, want to know why a request was or was not cached, or want to raise your hit rate.
Minimal requests
Accept a copy up to one hour old:
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/products/42", "response_format": "markdown", "cache_ttl": 3600}' \
-D - -o page.mdForce a fresh fetch: send "cache": false (CLI --no-cache) or "cache_ttl": 0. Either one skips the lookup and does not store the result.
What comes back
Every /v1/scrape response has a Cache-State header:
| Value | Meaning | Credits |
|---|---|---|
hit | Served from a stored result younger than your cache_ttl. A fresh X-Request-Id is issued. | 0 |
miss | Eligible, but nothing fresh enough was stored. Fetched now and stored if it succeeded. | Normal price |
bypass | Not looked up and not stored: caching was off for this request, or the request is never cacheable. | Normal price |
The body of a hit is the same document or envelope a fresh fetch would return, screenshots and extracted data included.
Options
| Field | Default | Effect |
|---|---|---|
cache | true | false turns the cache off for this request. cache: false with cache_ttl above 0 is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. |
cache_ttl | 172800 (48 h) | Maximum age, in seconds, of a stored result you will accept. 0 turns the cache off, whatever cache says. |
The default and the cap are both 48 hours on this deployment. A larger cache_ttl is clamped to the cap, not refused, and the response carries X-Warning: CACHE_TTL_CLAMPED. Freshness is decided per request: two callers with different cache_ttl values share one stored copy and each accepts it only if it is young enough for them.
What is never cached
A result is stored only when all of these hold:
- The request uses
method: GET(the default). Other methods can change the target. - There is no
session_id. A session's pages depend on its cookies. - There is no
actionslist. Actions have side effects and vary from run to run. - The result was a billable success: target status 200, 404, 410 or one of your
allowed_status_codes, with no error and no bot challenge. Failures, challenges and other target statuses are never stored, so a cached block cannot be served back to you. - The stored result fits under the deployment's size cap (1 MiB serialised by default). Larger results are served but not stored.
Requests excluded by the first three rules always report Cache-State: bypass.
What the cache key contains
A stored result is reused only for a request that matches it on every field that can change the output:
- Tenant: your organization and project. Results are never shared across projects or organizations. The API key is not part of the key, so every key in the same project shares the cache.
- Target:
urlandmethod. - Output: the resolved
response_format,main_content_only,include_tags,exclude_tags,links,autoparse,extract,extract_preset,ai_extract,network_capture, the screenshot fields andparse_pdf. - Engine and rendering:
js_render,impersonate,engine,mode,wait,wait_for,wait_for_timeout,block_resources,headless. - Target request:
custom_headers,original_status,allowed_status_codes. - Proxy intent: your own
proxy(its scheme, user, host and port; never the password).
cache and cache_ttl themselves are not in the key. To raise your hit rate, send identical parameters: {"js_render": true} and {"js_render": false} are different entries. An omitted impersonate is keyed as its default, true (or your project's default), so a request without it shares an entry with {"impersonate": true}.
Batch jobs
Batch jobs do not use the cache. On POST /v1/batch, cache defaults to false and the batch worker never reads stored results, whatever you send; every item is fetched. See Batch jobs.
Failure modes
| Code | HTTP | What to do |
|---|---|---|
ERR::REQUEST::INCOMPATIBLE_FLAGS | 400 | cache: false together with cache_ttl above 0. Send one or the other. |
ERR::REQUEST::INVALID_PARAMETER | 400 | cache_ttl is negative or not an integer. |
The cache itself cannot fail a request: if the cache store is unreachable, the request is fetched as a miss.
Cost
A hit costs 0 credits and still appears in your request history and usage, marked as a cache hit. A miss or bypass costs the normal engine price. Setting cache: false on every request removes the cheapest saving the API offers, so reserve it for data that must be current, such as stock levels or prices at checkout time. See Credits.
Related
- Markdown for LLMs
- Sessions and logins, which are never cached
- Browser actions, which are never cached
- Response headers