spicrawlspicrawlDocs

How to scrape Cloudflare-protected pages with Spicrawl

Bypass Cloudflare's managed challenge with one API call. Spicrawl escalates to full-browser rendering on your proxy and bills only the tier that served.

To scrape a Cloudflare-protected page, send POST /v1/scrape with your own proxy in proxy. When the site answers with a Cloudflare "Just a moment…" managed challenge (Turnstile), Spicrawl detects it, moves the request to full-browser rendering on the same proxy exit, and returns the real page. You set no browser and no rendering flag. On a live Cloudflare-protected page this took about 18 to 20 seconds and cost 8 credits.

This works with a plain request and with js_render: true. Residential proxies work best. Some sites with stricter rules can still block, so treat success as likely, not guaranteed.

Why Cloudflare blocks scrapers

Cloudflare scores each request on its IP reputation, its TLS fingerprint and its browser behaviour. A low score gets an interstitial ("Just a moment…", a managed challenge, or "Access denied") instead of the page. A plain HTTP client cannot run the checks the interstitial needs, and datacenter IP ranges are challenged on sight by most vendors. A challenged request returns no content, which is why scrapers see a 403 or a challenge page where a browser sees the site.

Scrape a Cloudflare-protected page

Pass your proxy in proxy as an http:// URL and request the page as usual. Replace http://user:pass@your-proxy:8000 with your own proxy.

curl --max-time 180 https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/protected-page", "proxy": "http://user:pass@your-proxy:8000", "response_format": "json"}'

The JSON response carries the real page, not the challenge:

{
  "url": "https://example.com/protected-page",
  "final_url": "https://example.com/protected-page",
  "status": 200,
  "content": "<!doctype html>…",
  "credits": 8,
  "proxy_source": "custom"
}

status is the target site's status. --meta prints the credits, target status and request id to stderr. Adding "js_render": true (CLI --js) gives the same result.

What happens behind the scenes

  1. The first attempt. Your request goes out on a plain fetch, or on JavaScript rendering if you set js_render: true, through your proxy.
  2. Detection. The response is a challenge that names its vendor, for example Cloudflare's cf-mitigated header or its "Just a moment…" interstitial.
  3. Escalation. Spicrawl abandons that attempt and retries on full-browser rendering, through the same proxy exit. Full-browser rendering runs with a real display and waits out the interstitial before it captures the page.
  4. The result. The page that cleared the challenge comes back with HTTP 200. If every step was challenged, you get ERR::UPSTREAM::CHALLENGE and pay nothing.

Spicrawl does not solve CAPTCHAs. It handles the managed challenge automatically by rendering the page the way a real browser would. A page that has to climb took about 18 to 20 seconds in our testing, so give your HTTP client a generous timeout (the examples above use 180 s).

Read these on the response:

HTTP/1.1 200 OK
X-Target-Status: 200
X-Credits-Charged: 8
X-Proxy-Source: custom
X-Warning: ESCALATED: ...
  • X-Warning: ESCALATED: ... explains the climb. There is one header per abandoned step, and it is sent on success too, so you can see why the price rose. The text after the code gives the reason.
  • X-Credits-Charged is the price of the tier that served the page.
  • X-Proxy-Source: custom confirms the request went through your proxy.

All headers are listed in Response headers. Full detail on the climb and its limits is under Automatic escalation.

Cost

You pay only for the tier that served the page. A failed request costs 0 credits.

OutcomeCredits
Plain fetch passed (no challenge)1
js_render: true passed (no challenge)3
A challenge moved the request to full-browser rendering, which cleared it8
Every step was challenged (ERR::UPSTREAM::CHALLENGE)0

Your own proxy adds no surcharge. If you set max_cost, keep it at 8 or more, or the climb to full-browser rendering is left off. See Credits.

Troubleshooting

SymptomWhat to do
ERR::UPSTREAM::CHALLENGE (HTTP 502, 0 credits, retryable)Retry with your own residential proxy in proxy. On the default shared pool a challenged page can still return this error. Retry once or twice more: the verdict is per exit and per moment, so a retry often passes (--retry 3 on the CLI).
Still challenged on a residential proxyThe site is showing a challenge that automation does not always clear. Open the request with spicrawl logs --status blocked, then spicrawl logs get <request-id>, to see the vendor and the signal.
ERR::UPSTREAM::TIMEOUT (HTTP 504, 0 credits, retryable)Retry. Also check that your own HTTP client timeout is longer than the climb, which took about 18 to 20 seconds in our testing.
No climb happenedRendering can only use an http:// proxy, and the climb is left off when max_cost is below 8, when the method is not GET, or when the request sets actions.
HTTP 200 with X-Target-Status: 403 or 429 and 0 creditsA bare 403 or 429 with no vendor marker is the site's own answer, not a challenge, and does not escalate. For a 429, slow down.
ERR::ENGINE::UNAVAILABLE (HTTP 503)Rendering capacity is busy. Retry after Retry-After.

Residential or datacenter proxy. Anti-bot vendors keep reputation lists of datacenter ranges and challenge them on sight, so a datacenter exit is the likeliest to fail. A residential exit passes far more often. Managed premium exits (premium_proxy) Coming soon are coming soon; until then, bring your own proxy. See Proxies and geo.

FAQ

On this page