How to scrape Cloudflare-protected pages with Spicrawl
Bypass Cloudflare's managed challenge with one API call. Spicrawl escalates to full-browser rendering on your proxy and bills only the tier that served.
To scrape a Cloudflare-protected page, send POST /v1/scrape with your own proxy in proxy. When the site answers with a Cloudflare "Just a moment…" managed challenge (Turnstile), Spicrawl detects it, moves the request to full-browser rendering on the same proxy exit, and returns the real page. You set no browser and no rendering flag. On a live Cloudflare-protected page this took about 18 to 20 seconds and cost 8 credits.
This works with a plain request and with js_render: true. Residential proxies work best. Some sites with stricter rules can still block, so treat success as likely, not guaranteed.
Why Cloudflare blocks scrapers
Cloudflare scores each request on its IP reputation, its TLS fingerprint and its browser behaviour. A low score gets an interstitial ("Just a moment…", a managed challenge, or "Access denied") instead of the page. A plain HTTP client cannot run the checks the interstitial needs, and datacenter IP ranges are challenged on sight by most vendors. A challenged request returns no content, which is why scrapers see a 403 or a challenge page where a browser sees the site.
Scrape a Cloudflare-protected page
Pass your proxy in proxy as an http:// URL and request the page as usual. Replace http://user:pass@your-proxy:8000 with your own proxy.
curl --max-time 180 https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/protected-page", "proxy": "http://user:pass@your-proxy:8000", "response_format": "json"}'The JSON response carries the real page, not the challenge:
{
"url": "https://example.com/protected-page",
"final_url": "https://example.com/protected-page",
"status": 200,
"content": "<!doctype html>…",
"credits": 8,
"proxy_source": "custom"
}status is the target site's status. --meta prints the credits, target status and request id to stderr. Adding "js_render": true (CLI --js) gives the same result.
What happens behind the scenes
- The first attempt. Your request goes out on a plain fetch, or on JavaScript rendering if you set
js_render: true, through yourproxy. - Detection. The response is a challenge that names its vendor, for example Cloudflare's
cf-mitigatedheader or its "Just a moment…" interstitial. - Escalation. Spicrawl abandons that attempt and retries on full-browser rendering, through the same proxy exit. Full-browser rendering runs with a real display and waits out the interstitial before it captures the page.
- The result. The page that cleared the challenge comes back with HTTP 200. If every step was challenged, you get
ERR::UPSTREAM::CHALLENGEand pay nothing.
Spicrawl does not solve CAPTCHAs. It handles the managed challenge automatically by rendering the page the way a real browser would. A page that has to climb took about 18 to 20 seconds in our testing, so give your HTTP client a generous timeout (the examples above use 180 s).
Read these on the response:
HTTP/1.1 200 OK
X-Target-Status: 200
X-Credits-Charged: 8
X-Proxy-Source: custom
X-Warning: ESCALATED: ...X-Warning: ESCALATED: ...explains the climb. There is one header per abandoned step, and it is sent on success too, so you can see why the price rose. The text after the code gives the reason.X-Credits-Chargedis the price of the tier that served the page.X-Proxy-Source: customconfirms the request went through your proxy.
All headers are listed in Response headers. Full detail on the climb and its limits is under Automatic escalation.
Cost
You pay only for the tier that served the page. A failed request costs 0 credits.
| Outcome | Credits |
|---|---|
| Plain fetch passed (no challenge) | 1 |
js_render: true passed (no challenge) | 3 |
| A challenge moved the request to full-browser rendering, which cleared it | 8 |
Every step was challenged (ERR::UPSTREAM::CHALLENGE) | 0 |
Your own proxy adds no surcharge. If you set max_cost, keep it at 8 or more, or the climb to full-browser rendering is left off. See Credits.
Troubleshooting
| Symptom | What to do |
|---|---|
ERR::UPSTREAM::CHALLENGE (HTTP 502, 0 credits, retryable) | Retry with your own residential proxy in proxy. On the default shared pool a challenged page can still return this error. Retry once or twice more: the verdict is per exit and per moment, so a retry often passes (--retry 3 on the CLI). |
| Still challenged on a residential proxy | The site is showing a challenge that automation does not always clear. Open the request with spicrawl logs --status blocked, then spicrawl logs get <request-id>, to see the vendor and the signal. |
ERR::UPSTREAM::TIMEOUT (HTTP 504, 0 credits, retryable) | Retry. Also check that your own HTTP client timeout is longer than the climb, which took about 18 to 20 seconds in our testing. |
| No climb happened | Rendering can only use an http:// proxy, and the climb is left off when max_cost is below 8, when the method is not GET, or when the request sets actions. |
HTTP 200 with X-Target-Status: 403 or 429 and 0 credits | A bare 403 or 429 with no vendor marker is the site's own answer, not a challenge, and does not escalate. For a 429, slow down. |
ERR::ENGINE::UNAVAILABLE (HTTP 503) | Rendering capacity is busy. Retry after Retry-After. |
Residential or datacenter proxy. Anti-bot vendors keep reputation lists of datacenter ranges and challenge them on sight, so a datacenter exit is the likeliest to fail. A residential exit passes far more often. Managed premium exits (premium_proxy) Coming soon are coming soon; until then, bring your own proxy. See Proxies and geo.
FAQ
Yes, when you send your own proxy in proxy. Spicrawl detects the Cloudflare challenge, escalates to full-browser rendering on the same exit and returns the real page, billed at 8 credits. Success depends on the site and the exit, so some sites with stricter rules can still block.
No. You do not run a browser, pick one or set a rendering flag. A plain request is enough, and Spicrawl climbs to full-browser rendering by itself when the page is challenged.
It handles the "Just a moment…" managed challenge page (Turnstile) automatically: a request for such a page returned the real page in about 18 to 20 seconds. It does this by rendering the page and waiting out the interstitial, not by solving CAPTCHAs.
Today, yes. Cloudflare-protected pages are supported when you bring your own proxy, and residential works best. Managed premium exits (premium_proxy) Coming soon are coming soon. On the default shared pool a challenged page can return ERR::UPSTREAM::CHALLENGE, and the fix is to retry with your own residential proxy.
You pay only for the tier that served the page: 8 credits for full-browser rendering. A request that fails costs 0 credits.
Related
- Anti-bot for the escalation ladder and how to read blocks in your logs
- Proxies and geo for
proxy,proxy_verifyandproxy_country - JavaScript rendering
- Credits
- Errors
Anti-bot
How Cloudflare-protected pages are handled through your own proxy, what to do on ERR::UPSTREAM::CHALLENGE, the escalation ladder with the credit cost of each rung, the automatic climb to stronger rendering, and how to read blocks in your request logs.
Proxies and geo
Route requests through your own proxy with proxy and proxy_verify, which also lets Cloudflare-protected pages clear automatically, and read how a request was routed. Spicrawl's managed proxy pool is coming soon.