# How to scrape Cloudflare-protected pages with Spicrawl

> Bypass Cloudflare's managed challenge with one API call. Spicrawl escalates to full-browser rendering on your proxy and bills only the tier that served.

Source: https://docs.spicrawl.com/guides/cloudflare

To scrape a Cloudflare-protected page, send `POST /v1/scrape` with your own proxy in `proxy`. When the site answers with a Cloudflare "Just a moment…" managed challenge (Turnstile), Spicrawl detects it, moves the request to full-browser rendering on the same proxy exit, and returns the real page. You set no browser and no rendering flag. On a live Cloudflare-protected page this took about 18 to 20 seconds and cost 8 credits.

This works with a plain request and with `js_render: true`. Residential proxies work best. Some sites with stricter rules can still block, so treat success as likely, not guaranteed.

## Why Cloudflare blocks scrapers

Cloudflare scores each request on its IP reputation, its TLS fingerprint and its browser behaviour. A low score gets an interstitial ("Just a moment…", a managed challenge, or "Access denied") instead of the page. A plain HTTP client cannot run the checks the interstitial needs, and datacenter IP ranges are challenged on sight by most vendors. A challenged request returns no content, which is why scrapers see a 403 or a challenge page where a browser sees the site.

## Scrape a Cloudflare-protected page

Pass your proxy in `proxy` as an `http://` URL and request the page as usual. Replace `http://user:pass@your-proxy:8000` with your own proxy.

```bash title="curl"
curl --max-time 180 https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/protected-page", "proxy": "http://user:pass@your-proxy:8000", "response_format": "json"}'
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={"url": "https://example.com/protected-page", "proxy": os.environ["MY_PROXY_URL"], "response_format": "json"},
    timeout=180,
)
r.raise_for_status()
page = r.json()
print(page["status"], page["credits"], r.headers.get("X-Warning"))
```

```typescript title="TypeScript"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url: "https://example.com/protected-page", proxy: process.env.MY_PROXY_URL, response_format: "json" }),
  signal: AbortSignal.timeout(180_000),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
const page = await r.json();
console.log(page.status, page.credits, r.headers.get("X-Warning"));
```

```bash title="CLI"
spicrawl scrape https://example.com/protected-page --proxy "$MY_PROXY_URL" --format markdown --meta --retry 3
```

The JSON response carries the real page, not the challenge:

```json
{
  "url": "https://example.com/protected-page",
  "final_url": "https://example.com/protected-page",
  "status": 200,
  "content": "<!doctype html>…",
  "credits": 8,
  "proxy_source": "custom"
}
```

`status` is the target site's status. `--meta` prints the credits, target status and request id to stderr. Adding `"js_render": true` (CLI `--js`) gives the same result.

## What happens behind the scenes

1. **The first attempt.** Your request goes out on a plain fetch, or on JavaScript rendering if you set `js_render: true`, through your `proxy`.
2. **Detection.** The response is a challenge that names its vendor, for example Cloudflare's `cf-mitigated` header or its "Just a moment…" interstitial.
3. **Escalation.** Spicrawl abandons that attempt and retries on full-browser rendering, through the same proxy exit. Full-browser rendering runs with a real display and waits out the interstitial before it captures the page.
4. **The result.** The page that cleared the challenge comes back with HTTP 200. If every step was challenged, you get `ERR::UPSTREAM::CHALLENGE` and pay nothing.

Spicrawl does not solve CAPTCHAs. It handles the managed challenge automatically by rendering the page the way a real browser would. A page that has to climb took about 18 to 20 seconds in our testing, so give your HTTP client a generous timeout (the examples above use 180 s).

Read these on the response:

```http
HTTP/1.1 200 OK
X-Target-Status: 200
X-Credits-Charged: 8
X-Proxy-Source: custom
X-Warning: ESCALATED: ...
```

* `X-Warning: ESCALATED: ...` explains the climb. There is one header per abandoned step, and it is sent on success too, so you can see why the price rose. The text after the code gives the reason.
* `X-Credits-Charged` is the price of the tier that served the page.
* `X-Proxy-Source: custom` confirms the request went through your proxy.

All headers are listed in [Response headers](https://docs.spicrawl.com/response-headers.md). Full detail on the climb and its limits is under [Automatic escalation](https://docs.spicrawl.com/guides/anti-bot.md#automatic-escalation-coming-soon).

## Cost

You pay only for the tier that served the page. A failed request costs 0 credits.

| Outcome                                                                   | Credits |
| ------------------------------------------------------------------------- | ------- |
| Plain fetch passed (no challenge)                                         | 1       |
| `js_render: true` passed (no challenge)                                   | 3       |
| A challenge moved the request to full-browser rendering, which cleared it | 8       |
| Every step was challenged (`ERR::UPSTREAM::CHALLENGE`)                    | 0       |

Your own proxy adds no surcharge. If you set `max_cost`, keep it at 8 or more, or the climb to full-browser rendering is left off. See [Credits](https://docs.spicrawl.com/credits.md).

## Troubleshooting

| Symptom                                                     | What to do                                                                                                                                                                                                                                      |
| ----------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `ERR::UPSTREAM::CHALLENGE` (HTTP 502, 0 credits, retryable) | Retry with your own residential proxy in `proxy`. On the default shared pool a challenged page can still return this error. Retry once or twice more: the verdict is per exit and per moment, so a retry often passes (`--retry 3` on the CLI). |
| Still challenged on a residential proxy                     | The site is showing a challenge that automation does not always clear. Open the request with `spicrawl logs --status blocked`, then `spicrawl logs get <request-id>`, to see the vendor and the signal.                                         |
| `ERR::UPSTREAM::TIMEOUT` (HTTP 504, 0 credits, retryable)   | Retry. Also check that your own HTTP client timeout is longer than the climb, which took about 18 to 20 seconds in our testing.                                                                                                                 |
| No climb happened                                           | Rendering can only use an `http://` proxy, and the climb is left off when `max_cost` is below 8, when the method is not `GET`, or when the request sets `actions`.                                                                              |
| HTTP 200 with `X-Target-Status: 403` or `429` and 0 credits | A bare 403 or 429 with no vendor marker is the site's own answer, not a challenge, and does not escalate. For a 429, slow down.                                                                                                                 |
| `ERR::ENGINE::UNAVAILABLE` (HTTP 503)                       | Rendering capacity is busy. Retry after `Retry-After`.                                                                                                                                                                                          |

**Residential or datacenter proxy.** Anti-bot vendors keep reputation lists of datacenter ranges and challenge them on sight, so a datacenter exit is the likeliest to fail. A residential exit passes far more often. Managed premium exits (`premium_proxy`) (coming soon) are coming soon; until then, bring your own `proxy`. See [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md).

## FAQ

**Can I scrape Cloudflare-protected sites?**

Yes, when you send your own proxy in `proxy`. Spicrawl detects the Cloudflare challenge, escalates to full-browser rendering on the same exit and returns the real page, billed at 8 credits. Success depends on the site and the exit, so some sites with stricter rules can still block.

**Do I need to run my own browser?**

No. You do not run a browser, pick one or set a rendering flag. A plain request is enough, and Spicrawl climbs to full-browser rendering by itself when the page is challenged.

**Does it handle Cloudflare Turnstile?**

It handles the "Just a moment…" managed challenge page (Turnstile) automatically: a request for such a page returned the real page in about 18 to 20 seconds. It does this by rendering the page and waiting out the interstitial, not by solving CAPTCHAs.

**Do I need a proxy?**

Today, yes. Cloudflare-protected pages are supported when you bring your own proxy, and residential works best. Managed premium exits (`premium_proxy`) (coming soon) are coming soon. On the default shared pool a challenged page can return `ERR::UPSTREAM::CHALLENGE`, and the fix is to retry with your own residential proxy.

**What does it cost to scrape a Cloudflare-protected page?**

You pay only for the tier that served the page: 8 credits for full-browser rendering. A request that fails costs 0 credits.

## Related

* [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md) for the escalation ladder and how to read blocks in your logs
* [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md) for `proxy`, `proxy_verify` and `proxy_country`
* [JavaScript rendering](https://docs.spicrawl.com/guides/javascript-rendering.md)
* [Credits](https://docs.spicrawl.com/credits.md)
* [Errors](https://docs.spicrawl.com/errors.md)
