# Scrape bot-protected sites (Cloudflare, PerimeterX, more)

> Bypass Cloudflare and PerimeterX blocks when scraping with one API call. Automatic challenge handling on your own proxy, billed only on success.

Source: https://docs.spicrawl.com/use-cases/protected-sites

Yes: send `POST /v1/scrape` with your own proxy in `proxy`. When a Cloudflare, PerimeterX or similar challenge answers instead of the page, Spicrawl moves to full-browser rendering on the same exit, waits it out and returns the real page at 8 credits (0 if it never clears).

 

## How do I scrape a protected page?

Request the page as usual with your proxy in `proxy`; Spicrawl escalates only if the plain fetch meets a challenge ([how it works](https://docs.spicrawl.com/guides/anti-bot.md#automatic-escalation-coming-soon)).

```bash title="curl"
curl --max-time 180 https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/community/thread/42", "proxy": "http://user:pass@your-proxy:8000", "response_format": "markdown"}' \
  -D -
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={"url": "https://example.com/community/thread/42", "proxy": os.environ["MY_PROXY_URL"], "response_format": "markdown"},
    timeout=180,
)
r.raise_for_status()
print(r.headers["X-Credits-Charged"], r.headers.get("X-Warning"))
print(r.text[:500])
```

```typescript title="TypeScript"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url: "https://example.com/community/thread/42", proxy: process.env.MY_PROXY_URL, response_format: "markdown" }),
  signal: AbortSignal.timeout(180_000),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
console.log(r.headers.get("X-Credits-Charged"), r.headers.get("X-Warning"));
console.log((await r.text()).slice(0, 500));
```

```bash title="CLI"
spicrawl scrape https://example.com/community/thread/42 --proxy "$MY_PROXY_URL" --format markdown --retry 3 --meta
```

The headers of a real Cloudflare-protected request, trimmed:

```http
HTTP/1.1 200 OK
X-Target-Status: 200
X-Credits-Charged: 8
X-Proxy-Source: custom
X-Warning: ESCALATED: fetch (1 credits): blocked by challenge page (cloudflare: cf-mitigated response header)
```

`X-Credits-Charged` is the tier that served the page; `X-Warning: ESCALATED` says why it moved up and is sent on success too. Keep your client timeout at 180 s.

## Which bot-protection vendors does it handle, and what does it cost?

| Vendor                                         | What we measured                              | Credits                       |
| ---------------------------------------------- | --------------------------------------------- | ----------------------------- |
| Cloudflare (managed challenge, Turnstile page) | Cleared 3 of 3, 15 to 36 s (9 Oct 2026)       | 8                             |
| PerimeterX (HUMAN)                             | Cleared once; a repeat request was challenged | 8 on success, 0 if challenged |
| Akamai                                         | Block pages detected; some cleared            | 8 if cleared, 0 if blocked    |
| Kasada, Imperva                                | Block pages detected                          | 0                             |
| DataDome                                       | Two listing pages stayed challenged           | 0                             |

Results depend on the site and the exit. You pay for the tier that served the page: 1 credit for a plain fetch, 3 for `js_render: true`, 8 for full-browser rendering, a PDF or a screenshot. Keep `max_cost` at 8 or more, or the climb is off; the beta's 1,000 monthly credits are 125 pages at 8 ([Credits](https://docs.spicrawl.com/credits.md)).

## Which business scenarios does this fit?

Three we tested:

* **Forum archive behind Cloudflare:** request threads as `markdown` with the default cache. One came back as 11.7 KB at 8 credits.
* **Crypto dashboard or token pair page behind Cloudflare:** send `cache: false`, since results are cached for 48 hours; poll every few minutes. A pair page came back as 17.6 KB at 8 credits ([Caching](https://docs.spicrawl.com/guides/caching.md)).
* **Finance research behind PerimeterX:** treat a challenged answer as a retry. An unprotected quote page billed 1 credit, so one code path covers a mixed watchlist.

### How do I save a PDF or screenshot of a protected page?

Add `js_render: true` with `response_format: "pdf"`, or send `screenshot: true`. Both wait out the challenge first and add nothing to the 8 credits. A Cloudflare-protected forum thread came back as a 4-page, 941 KB PDF.

```bash title="curl"
curl --max-time 180 https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/community/thread/42", "js_render": true, "proxy": "http://user:pass@your-proxy:8000", "response_format": "pdf"}' \
  -o thread.pdf
head -c 5 thread.pdf   # %PDF-
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={"url": "https://example.com/community/thread/42", "js_render": True,
          "proxy": os.environ["MY_PROXY_URL"], "response_format": "pdf"},
    timeout=180,
)
r.raise_for_status()
assert r.content.startswith(b"%PDF-")
open("thread.pdf", "wb").write(r.content)
print(r.headers["X-Credits-Charged"])
```

```bash title="CLI"
spicrawl scrape https://example.com/community/thread/42 --render --proxy "$MY_PROXY_URL" --format pdf -o thread.pdf
```

A screenshot comes back in the JSON envelope, base64-encoded under `screenshots[].data` ([Screenshots and PDF](https://docs.spicrawl.com/guides/screenshots-and-pdf.md)):

```bash title="curl"
curl --max-time 180 https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/community/thread/42", "js_render": true, "proxy": "http://user:pass@your-proxy:8000", "screenshot": true, "screenshot_fullpage": true}' \
  | jq -r '.screenshots[0].data' | base64 -d > thread.png
```

```bash title="CLI"
spicrawl scrape https://example.com/community/thread/42 --render --proxy "$MY_PROXY_URL" \
  --screenshot --screenshot-full-page -o thread.html
# saves thread.html plus thread.final.png next to it
```

## How do I run it across a watchlist?

Retry retryable failures, keep requests in flight at or under your `Concurrency-Limit` header ([Rate limits](https://docs.spicrawl.com/rate-limits.md)), and use `/v1/scrape`: `/v1/batch` never climbs to full-browser rendering.

```python title="Python"
import os, time, requests
from concurrent.futures import ThreadPoolExecutor

URLS = ["https://example.com/pair/1", "https://example.com/pair/2"]

def fetch(url, tries=3):
    for _ in range(tries):
        r = requests.post(
            "https://api.spicrawl.com/v1/scrape",
            headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
            json={"url": url, "proxy": os.environ["MY_PROXY_URL"], "response_format": "markdown", "cache": False},
            timeout=180,
        )
        if r.ok:
            return url, int(r.headers["X-Credits-Charged"]), r.text
        problem = r.json() if r.headers.get("content-type", "").startswith("application/problem") else {}
        if not problem.get("retryable"):
            break
        time.sleep(int(r.headers.get("Retry-After", 2)))
    return url, 0, None

with ThreadPoolExecutor(max_workers=4) as pool:  # stay under Concurrency-Limit
    for url, credits, text in pool.map(fetch, URLS):
        print(url, credits, "ok" if text else "failed")
```

```bash title="CLI"
# one URL per line; 4 at a time; keep the failures
spicrawl scrape - --proxy "$MY_PROXY_URL" --format markdown --no-cache --concurrency 4 --retry 3 < urls.txt \
  | jq -c 'select(.error)'
```

If the site still blocks you ([all error codes](https://docs.spicrawl.com/errors.md#UPSTREAM_CHALLENGE)):

| Symptom                                                      | What to do                                                                                                                                                         |
| ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `ERR::UPSTREAM::CHALLENGE` (502, 0 credits, retryable)       | Retry once or twice with a residential proxy. A DataDome page may keep holding; no setting forces it through.                                                      |
| `ERR::PROXY::UNREACHABLE` or `ERR::PROXY::AUTH_FAILED` (502) | Fix the proxy URL or credentials; `proxy_verify: true` fails before any rendering cost.                                                                            |
| Challenge returned, no climb                                 | The climb is off when `max_cost` is below 8, the method is not `GET`, the request sets `actions`, `session_id` or `mode=auto`, or `proxy` is not an `http://` URL. |

## Frequently asked questions

**Can I scrape Cloudflare-protected sites, including Turnstile pages?**

Yes, with your own proxy in `proxy`. Spicrawl detects the "Just a moment…" challenge, moves to full-browser rendering on the same exit and returns the real page at 8 credits; it does not solve CAPTCHAs. On 9 October 2026 a Cloudflare-protected page cleared 3 of 3 times in 15 to 36 s, but stricter sites can still block. See [the Cloudflare guide](https://docs.spicrawl.com/guides/cloudflare.md).

**Can I scrape PerimeterX, Akamai, Kasada, Imperva or DataDome sites?**

PerimeterX: yes in our tests, a finance research page cleared at 8 credits, but the verdict is per exit and per moment, so retry. Akamai, Kasada and Imperva block pages return `ERR::UPSTREAM::CHALLENGE` at 0 credits, never a billed "Access Denied" page. DataDome may still block.

**Do I need a proxy?**

For protected sites, yes: bring your own in `proxy`, residential if you can, because a datacenter exit is the likeliest to be challenged. Managed premium exits (`premium_proxy`) are not available yet.

**What does it cost, and what if it fails?**

You pay only for the tier that served the page: 1 credit for a plain fetch, 3 for JavaScript rendering, 8 for full-browser rendering. Failed and challenged attempts cost 0 credits; a retry of a success bills again.

**What if the site still blocks me?**

You get `ERR::UPSTREAM::CHALLENGE` (HTTP 502, retryable) at 0 credits. Retry once or twice and use a residential proxy; see [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md#if-you-get-a-challenge).

**Is it OK to scrape a protected site?**

That is for you to decide per site: Spicrawl returns the page a visitor would see and grants no right to the content. Check the site's terms and `robots.txt`, pace your requests and keep to public pages.

## Related

* [How to scrape Cloudflare-protected pages](https://docs.spicrawl.com/guides/cloudflare.md)
* [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md)
* [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md)
* [Screenshots and PDF](https://docs.spicrawl.com/guides/screenshots-and-pdf.md)
* [Caching](https://docs.spicrawl.com/guides/caching.md), [Credits](https://docs.spicrawl.com/credits.md), [Rate limits](https://docs.spicrawl.com/rate-limits.md), [Errors](https://docs.spicrawl.com/errors.md)
* [Job postings](https://docs.spicrawl.com/use-cases/job-postings.md) for job boards behind these challenges

 

## Disclaimer

> **For educational purposes:** The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and `robots.txt`, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.
