spicrawlspicrawlDocs

Scrape bot-protected sites (Cloudflare, PerimeterX, more)

Bypass Cloudflare and PerimeterX blocks when scraping with one API call. Automatic challenge handling on your own proxy, billed only on success.

Yes: send POST /v1/scrape with your own proxy in proxy. When a Cloudflare, PerimeterX or similar challenge answers instead of the page, Spicrawl moves to full-browser rendering on the same exit, waits it out and returns the real page at 8 credits (0 if it never clears).

How do I scrape a protected page?

Request the page as usual with your proxy in proxy; Spicrawl escalates only if the plain fetch meets a challenge (how it works).

curl --max-time 180 https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/community/thread/42", "proxy": "http://user:pass@your-proxy:8000", "response_format": "markdown"}' \
  -D -

The headers of a real Cloudflare-protected request, trimmed:

HTTP/1.1 200 OK
X-Target-Status: 200
X-Credits-Charged: 8
X-Proxy-Source: custom
X-Warning: ESCALATED: fetch (1 credits): blocked by challenge page (cloudflare: cf-mitigated response header)

X-Credits-Charged is the tier that served the page; X-Warning: ESCALATED says why it moved up and is sent on success too. Keep your client timeout at 180 s.

Which bot-protection vendors does it handle, and what does it cost?

VendorWhat we measuredCredits
Cloudflare (managed challenge, Turnstile page)Cleared 3 of 3, 15 to 36 s (9 Oct 2026)8
PerimeterX (HUMAN)Cleared once; a repeat request was challenged8 on success, 0 if challenged
AkamaiBlock pages detected; some cleared8 if cleared, 0 if blocked
Kasada, ImpervaBlock pages detected0
DataDomeTwo listing pages stayed challenged0

Results depend on the site and the exit. You pay for the tier that served the page: 1 credit for a plain fetch, 3 for js_render: true, 8 for full-browser rendering, a PDF or a screenshot. Keep max_cost at 8 or more, or the climb is off; the beta's 1,000 monthly credits are 125 pages at 8 (Credits).

Which business scenarios does this fit?

Three we tested:

  • Forum archive behind Cloudflare: request threads as markdown with the default cache. One came back as 11.7 KB at 8 credits.
  • Crypto dashboard or token pair page behind Cloudflare: send cache: false, since results are cached for 48 hours; poll every few minutes. A pair page came back as 17.6 KB at 8 credits (Caching).
  • Finance research behind PerimeterX: treat a challenged answer as a retry. An unprotected quote page billed 1 credit, so one code path covers a mixed watchlist.

How do I save a PDF or screenshot of a protected page?

Add js_render: true with response_format: "pdf", or send screenshot: true. Both wait out the challenge first and add nothing to the 8 credits. A Cloudflare-protected forum thread came back as a 4-page, 941 KB PDF.

curl --max-time 180 https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/community/thread/42", "js_render": true, "proxy": "http://user:pass@your-proxy:8000", "response_format": "pdf"}' \
  -o thread.pdf
head -c 5 thread.pdf   # %PDF-

A screenshot comes back in the JSON envelope, base64-encoded under screenshots[].data (Screenshots and PDF):

curl --max-time 180 https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/community/thread/42", "js_render": true, "proxy": "http://user:pass@your-proxy:8000", "screenshot": true, "screenshot_fullpage": true}' \
  | jq -r '.screenshots[0].data' | base64 -d > thread.png

How do I run it across a watchlist?

Retry retryable failures, keep requests in flight at or under your Concurrency-Limit header (Rate limits), and use /v1/scrape: /v1/batch never climbs to full-browser rendering.

import os, time, requests
from concurrent.futures import ThreadPoolExecutor

URLS = ["https://example.com/pair/1", "https://example.com/pair/2"]

def fetch(url, tries=3):
    for _ in range(tries):
        r = requests.post(
            "https://api.spicrawl.com/v1/scrape",
            headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
            json={"url": url, "proxy": os.environ["MY_PROXY_URL"], "response_format": "markdown", "cache": False},
            timeout=180,
        )
        if r.ok:
            return url, int(r.headers["X-Credits-Charged"]), r.text
        problem = r.json() if r.headers.get("content-type", "").startswith("application/problem") else {}
        if not problem.get("retryable"):
            break
        time.sleep(int(r.headers.get("Retry-After", 2)))
    return url, 0, None

with ThreadPoolExecutor(max_workers=4) as pool:  # stay under Concurrency-Limit
    for url, credits, text in pool.map(fetch, URLS):
        print(url, credits, "ok" if text else "failed")

If the site still blocks you (all error codes):

SymptomWhat to do
ERR::UPSTREAM::CHALLENGE (502, 0 credits, retryable)Retry once or twice with a residential proxy. A DataDome page may keep holding; no setting forces it through.
ERR::PROXY::UNREACHABLE or ERR::PROXY::AUTH_FAILED (502)Fix the proxy URL or credentials; proxy_verify: true fails before any rendering cost.
Challenge returned, no climbThe climb is off when max_cost is below 8, the method is not GET, the request sets actions, session_id or mode=auto, or proxy is not an http:// URL.

Frequently asked questions

Start scraping in minutes

1,000 free credits every month. One API key, one request.

Disclaimer

For educational purposes

The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and robots.txt, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.

On this page