# Cache results and skip repeat charges

> How the /v1/scrape result cache works: on by default with a 48-hour window, cache hits cost 0 credits, and cache=false or cache_ttl=0 forces a fresh fetch.

Source: https://docs.spicrawl.com/guides/caching

The result cache is on by default for `/v1/scrape`. When you repeat a request and a stored result is younger than `cache_ttl` seconds, you get that result back immediately, with `Cache-State: hit` and 0 credits charged. Use this page when you need fresher data than that, want to know why a request was or was not cached, or want to raise your hit rate.

## Minimal requests

Accept a copy up to one hour old:

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/products/42", "response_format": "markdown", "cache_ttl": 3600}' \
  -D - -o page.md
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={"url": "https://example.com/products/42", "response_format": "markdown", "cache_ttl": 3600},
    timeout=120,
)
r.raise_for_status()
print(r.headers["Cache-State"], r.headers["X-Credits-Charged"])
```

```typescript title="TypeScript"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url: "https://example.com/products/42", response_format: "markdown", cache_ttl: 3600 }),
});
console.log(r.headers.get("Cache-State"), r.headers.get("X-Credits-Charged"));
```

```bash title="CLI"
spicrawl scrape https://example.com/products/42 --format markdown --cache-ttl 3600 --meta
```

Force a fresh fetch: send `"cache": false` (CLI `--no-cache`) or `"cache_ttl": 0`. Either one skips the lookup and does not store the result.

## What comes back

Every `/v1/scrape` response has a `Cache-State` header:

| Value    | Meaning                                                                                            | Credits      |
| -------- | -------------------------------------------------------------------------------------------------- | ------------ |
| `hit`    | Served from a stored result younger than your `cache_ttl`. A fresh `X-Request-Id` is issued.       | 0            |
| `miss`   | Eligible, but nothing fresh enough was stored. Fetched now and stored if it succeeded.             | Normal price |
| `bypass` | Not looked up and not stored: caching was off for this request, or the request is never cacheable. | Normal price |

The body of a hit is the same document or envelope a fresh fetch would return, screenshots and extracted `data` included.

## Options

| Field       | Default         | Effect                                                                                                                           |
| ----------- | --------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| `cache`     | `true`          | `false` turns the cache off for this request. `cache: false` with `cache_ttl` above 0 is `400 ERR::REQUEST::INCOMPATIBLE_FLAGS`. |
| `cache_ttl` | `172800` (48 h) | Maximum age, in seconds, of a stored result you will accept. `0` turns the cache off, whatever `cache` says.                     |

The default and the cap are both 48 hours on this deployment. A larger `cache_ttl` is clamped to the cap, not refused, and the response carries `X-Warning: CACHE_TTL_CLAMPED`. Freshness is decided per request: two callers with different `cache_ttl` values share one stored copy and each accepts it only if it is young enough for them.

## What is never cached

A result is stored only when all of these hold:

* The request uses `method: GET` (the default). Other methods can change the target.
* There is no `session_id`. A session's pages depend on its cookies.
* There is no `actions` list. Actions have side effects and vary from run to run.
* The result was a billable success: target status 200, 404, 410 or one of your `allowed_status_codes`, with no error and no bot challenge. Failures, challenges and other target statuses are never stored, so a cached block cannot be served back to you.
* The stored result fits under the deployment's size cap (1 MiB serialised by default). Larger results are served but not stored.

Requests excluded by the first three rules always report `Cache-State: bypass`.

## What the cache key contains

A stored result is reused only for a request that matches it on every field that can change the output:

* **Tenant:** your organization and project. Results are never shared across projects or organizations. The API key is not part of the key, so every key in the same project shares the cache.
* **Target:** `url` and `method`.
* **Output:** the resolved `response_format`, `main_content_only`, `include_tags`, `exclude_tags`, `links`, `autoparse`, `extract`, `extract_preset`, `ai_extract`, `network_capture`, the screenshot fields and `parse_pdf`.
* **Engine and rendering:** `js_render`, `impersonate`, `engine`, `mode`, `wait`, `wait_for`, `wait_for_timeout`, `block_resources`, `headless`.
* **Target request:** `custom_headers`, `original_status`, `allowed_status_codes`.
* **Proxy intent:** your own `proxy` (its scheme, user, host and port; never the password).

`cache` and `cache_ttl` themselves are not in the key. To raise your hit rate, send identical parameters: `{"js_render": true}` and `{"js_render": false}` are different entries. An omitted `impersonate` is keyed as its default, `true` (or your project's default), so a request without it shares an entry with `{"impersonate": true}`.

## Batch jobs

Batch jobs do not use the cache. On `POST /v1/batch`, `cache` defaults to `false` and the batch worker never reads stored results, whatever you send; every item is fetched. See [Batch jobs](https://docs.spicrawl.com/guides/batch.md).

## Failure modes

| Code                               | HTTP | What to do                                                               |
| ---------------------------------- | ---- | ------------------------------------------------------------------------ |
| `ERR::REQUEST::INCOMPATIBLE_FLAGS` | 400  | `cache: false` together with `cache_ttl` above 0. Send one or the other. |
| `ERR::REQUEST::INVALID_PARAMETER`  | 400  | `cache_ttl` is negative or not an integer.                               |

The cache itself cannot fail a request: if the cache store is unreachable, the request is fetched as a miss.

## Cost

A hit costs 0 credits and still appears in your request history and usage, marked as a cache hit. A miss or bypass costs the normal engine price. Setting `cache: false` on every request removes the cheapest saving the API offers, so reserve it for data that must be current, such as stock levels or prices at checkout time. See [Credits](https://docs.spicrawl.com/credits.md).

## Related

* [Markdown for LLMs](https://docs.spicrawl.com/guides/markdown.md)
* [Sessions and logins](https://docs.spicrawl.com/guides/sessions-and-logins.md), which are never cached
* [Browser actions](https://docs.spicrawl.com/guides/browser-actions.md), which are never cached
* [Response headers](https://docs.spicrawl.com/response-headers.md)
