Scrape bot-protected sites (Cloudflare, PerimeterX, more)
Bypass Cloudflare and PerimeterX blocks when scraping with one API call. Automatic challenge handling on your own proxy, billed only on success.
Yes: send POST /v1/scrape with your own proxy in proxy. When a Cloudflare, PerimeterX or similar challenge answers instead of the page, Spicrawl moves to full-browser rendering on the same exit, waits it out and returns the real page at 8 credits (0 if it never clears).
How do I scrape a protected page?
Request the page as usual with your proxy in proxy; Spicrawl escalates only if the plain fetch meets a challenge (how it works).
curl --max-time 180 https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/community/thread/42", "proxy": "http://user:pass@your-proxy:8000", "response_format": "markdown"}' \
-D -The headers of a real Cloudflare-protected request, trimmed:
HTTP/1.1 200 OK
X-Target-Status: 200
X-Credits-Charged: 8
X-Proxy-Source: custom
X-Warning: ESCALATED: fetch (1 credits): blocked by challenge page (cloudflare: cf-mitigated response header)X-Credits-Charged is the tier that served the page; X-Warning: ESCALATED says why it moved up and is sent on success too. Keep your client timeout at 180 s.
Which bot-protection vendors does it handle, and what does it cost?
| Vendor | What we measured | Credits |
|---|---|---|
| Cloudflare (managed challenge, Turnstile page) | Cleared 3 of 3, 15 to 36 s (9 Oct 2026) | 8 |
| PerimeterX (HUMAN) | Cleared once; a repeat request was challenged | 8 on success, 0 if challenged |
| Akamai | Block pages detected; some cleared | 8 if cleared, 0 if blocked |
| Kasada, Imperva | Block pages detected | 0 |
| DataDome | Two listing pages stayed challenged | 0 |
Results depend on the site and the exit. You pay for the tier that served the page: 1 credit for a plain fetch, 3 for js_render: true, 8 for full-browser rendering, a PDF or a screenshot. Keep max_cost at 8 or more, or the climb is off; the beta's 1,000 monthly credits are 125 pages at 8 (Credits).
Which business scenarios does this fit?
Three we tested:
- Forum archive behind Cloudflare: request threads as
markdownwith the default cache. One came back as 11.7 KB at 8 credits. - Crypto dashboard or token pair page behind Cloudflare: send
cache: false, since results are cached for 48 hours; poll every few minutes. A pair page came back as 17.6 KB at 8 credits (Caching). - Finance research behind PerimeterX: treat a challenged answer as a retry. An unprotected quote page billed 1 credit, so one code path covers a mixed watchlist.
How do I save a PDF or screenshot of a protected page?
Add js_render: true with response_format: "pdf", or send screenshot: true. Both wait out the challenge first and add nothing to the 8 credits. A Cloudflare-protected forum thread came back as a 4-page, 941 KB PDF.
curl --max-time 180 https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/community/thread/42", "js_render": true, "proxy": "http://user:pass@your-proxy:8000", "response_format": "pdf"}' \
-o thread.pdf
head -c 5 thread.pdf # %PDF-A screenshot comes back in the JSON envelope, base64-encoded under screenshots[].data (Screenshots and PDF):
curl --max-time 180 https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/community/thread/42", "js_render": true, "proxy": "http://user:pass@your-proxy:8000", "screenshot": true, "screenshot_fullpage": true}' \
| jq -r '.screenshots[0].data' | base64 -d > thread.pngHow do I run it across a watchlist?
Retry retryable failures, keep requests in flight at or under your Concurrency-Limit header (Rate limits), and use /v1/scrape: /v1/batch never climbs to full-browser rendering.
import os, time, requests
from concurrent.futures import ThreadPoolExecutor
URLS = ["https://example.com/pair/1", "https://example.com/pair/2"]
def fetch(url, tries=3):
for _ in range(tries):
r = requests.post(
"https://api.spicrawl.com/v1/scrape",
headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
json={"url": url, "proxy": os.environ["MY_PROXY_URL"], "response_format": "markdown", "cache": False},
timeout=180,
)
if r.ok:
return url, int(r.headers["X-Credits-Charged"]), r.text
problem = r.json() if r.headers.get("content-type", "").startswith("application/problem") else {}
if not problem.get("retryable"):
break
time.sleep(int(r.headers.get("Retry-After", 2)))
return url, 0, None
with ThreadPoolExecutor(max_workers=4) as pool: # stay under Concurrency-Limit
for url, credits, text in pool.map(fetch, URLS):
print(url, credits, "ok" if text else "failed")If the site still blocks you (all error codes):
| Symptom | What to do |
|---|---|
ERR::UPSTREAM::CHALLENGE (502, 0 credits, retryable) | Retry once or twice with a residential proxy. A DataDome page may keep holding; no setting forces it through. |
ERR::PROXY::UNREACHABLE or ERR::PROXY::AUTH_FAILED (502) | Fix the proxy URL or credentials; proxy_verify: true fails before any rendering cost. |
| Challenge returned, no climb | The climb is off when max_cost is below 8, the method is not GET, the request sets actions, session_id or mode=auto, or proxy is not an http:// URL. |
Frequently asked questions
Yes, with your own proxy in proxy. Spicrawl detects the "Just a moment…" challenge, moves to full-browser rendering on the same exit and returns the real page at 8 credits; it does not solve CAPTCHAs. On 9 October 2026 a Cloudflare-protected page cleared 3 of 3 times in 15 to 36 s, but stricter sites can still block. See the Cloudflare guide.
PerimeterX: yes in our tests, a finance research page cleared at 8 credits, but the verdict is per exit and per moment, so retry. Akamai, Kasada and Imperva block pages return ERR::UPSTREAM::CHALLENGE at 0 credits, never a billed "Access Denied" page. DataDome may still block.
For protected sites, yes: bring your own in proxy, residential if you can, because a datacenter exit is the likeliest to be challenged. Managed premium exits (premium_proxy) are not available yet.
You pay only for the tier that served the page: 1 credit for a plain fetch, 3 for JavaScript rendering, 8 for full-browser rendering. Failed and challenged attempts cost 0 credits; a retry of a success bills again.
You get ERR::UPSTREAM::CHALLENGE (HTTP 502, retryable) at 0 credits. Retry once or twice and use a residential proxy; see Anti-bot.
That is for you to decide per site: Spicrawl returns the page a visitor would see and grants no right to the content. Check the site's terms and robots.txt, pace your requests and keep to public pages.
Related
- How to scrape Cloudflare-protected pages
- Anti-bot
- Proxies and geo
- Screenshots and PDF
- Caching, Credits, Rate limits, Errors
- Job postings for job boards behind these challenges
Start scraping in minutes
1,000 free credits every month. One API key, one request.
Disclaimer
For educational purposes
The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and robots.txt, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.
Job postings
Job scraping API cookbook: scrape job boards, infinite-scroll listings and public job APIs with Spicrawl. Tested requests, credit costs and limits.
AI knowledge base
Scrape websites and docs sites to clean markdown for RAG and AI agents: batch ingestion, chunking, the MCP server, llms.txt and scheduled refreshes.