# Turn websites into LLM-ready data for RAG and AI agents

> Scrape websites and docs sites to clean markdown for RAG and AI agents: batch ingestion, chunking, the MCP server, llms.txt and scheduled refreshes.

Source: https://docs.spicrawl.com/use-cases/ai-knowledge-base

Send a URL with `response_format: "markdown"` to get clean main-content markdown for embeddings and LLM context. Batch up to 10,000 URLs for a knowledge base, or connect the hosted MCP server for live agent access; Spicrawl does not crawl, chunk or embed, so the URL list and vector store are yours.

 

## How do I turn a web page into LLM-ready markdown?

Send the URL with `response_format: "markdown"`; `include_tags` and `exclude_tags` pick the content container before conversion. See [Markdown for LLMs](https://docs.spicrawl.com/guides/markdown.md) for every option.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/docs/install",
    "response_format": "markdown",
    "include_tags": ["article"],
    "max_cost": 1
  }'
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={
        "url": "https://example.com/docs/install",
        "response_format": "markdown",
        "include_tags": ["article"],
        "max_cost": 1,
    },
    timeout=120,
)
r.raise_for_status()
if r.headers["X-Target-Status"] == "200":   # the site's status, not Spicrawl's
    markdown = r.text
```

```bash title="CLI"
spicrawl scrape https://example.com/docs/install --format markdown --include article --max-cost 1
```

> **Warning:** A `200` from Spicrawl is not a `200` from the site: a page that answered 403 still returns HTTP 200 at 0 credits, so read `X-Target-Status` before you index. A 404 or 410 means the page is gone.

### How do I handle JavaScript-heavy documentation sites?

A near-empty body with `X-Target-Status: 200` means JavaScript builds the page. Add `js_render: true` and a `wait_for` selector that exists only once the article is built (3 credits). `X-Warning: RENDER_DEGRADED` means `wait_for` never matched, so fix the selector.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/docs/api/reference",
    "js_render": true,
    "wait_for": "article h1",
    "response_format": "markdown",
    "include_tags": ["article"],
    "max_cost": 3
  }'
```

```bash title="CLI"
spicrawl scrape https://example.com/docs/api/reference --render --wait-for 'article h1' \
  --format markdown --include article --max-cost 3 --meta
```

* **Probe first.** Fetch one page per section plain (1 credit) and render only the sections that come back empty. In a batch, set `js_render` on the job or on individual `items`.
* **Interaction.** Content behind "Load more", a tab or infinite scroll needs `actions` (8 credits), which batch refuses: use `/v1/scrape`.
* **Bot walls.** `ERR::UPSTREAM::CHALLENGE` (retryable, 0 credits) means a challenge; route through your own residential `proxy`, which is likely to work but not guaranteed. See [Cloudflare-protected pages](https://docs.spicrawl.com/guides/cloudflare.md).

## How do I scrape a whole documentation site for RAG?

Spicrawl has no crawl endpoint, so you collect the URLs, submit one batch job and page the results into your chunker.

**Step 1: Collect the URLs**

Take them from a sitemap, an `llms.txt` file (a markdown list of a site's pages for LLMs) or a page's `links` (`links: true`, no extra credits, JSON envelope). Fetch sitemaps and `llms.txt` with `response_format: "html"`, which returns the file as served.

```python title="urls.py"
import os, re, requests

API = "https://api.spicrawl.com"
H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}

def fetch_html(url: str) -> str:
    # "html" returns XML and llms.txt as served; markdown refuses non-HTML documents
    r = requests.post(f"{API}/v1/scrape", headers=H, timeout=120,
                      json={"url": url, "response_format": "html", "max_cost": 1})
    r.raise_for_status()
    return r.text

# A sitemap index lists other sitemaps: fetch each in turn.
locs = re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", fetch_html("https://example.com/sitemap.xml"))
# llms.txt instead: re.findall(r"\]\((https?://[^)\s]+)\)", fetch_html("https://example.com/llms.txt"))
urls = sorted({u for u in locs if u.startswith("https://example.com/docs/")})
```

**Step 2: Submit one batch job**

Set `main_content_only: true` and the content selector once for the job, and add a `credit_budget` so a wrong URL list cannot overspend.

```python title="Python"
import time

job = requests.post(f"{API}/v1/batch", headers=H, timeout=60, json={
    "name": "docs-example-2026-10-09",
    "urls": urls,
    "response_format": "markdown",
    "main_content_only": True,
    "include_tags": ["article"],
    "exclude_tags": [".cookie-banner", "nav.toc"],
    "credit_budget": 200,
}).json()

while job["status"] not in ("completed", "failed", "cancelled"):
    time.sleep(10)
    job = requests.get(f"{API}/v1/batch/{job['id']}", headers=H, timeout=30).json()
```

```bash title="CLI"
spicrawl batch submit urls.txt --format markdown --main-content --include article \
  --exclude .cookie-banner --credit-budget 200 --name docs-example-2026-10-09 \
  --wait --max-wait 30m
```

Submitting returns `202` with `estimated_credits` (a hold, not a charge). Do not resubmit if `warnings` mentions unqueued items: a second submission is a second job and bills again.

**Step 3: Read the results into your chunker**

`GET /v1/batch/{id}/results` returns one JSON object per line and pages with the `X-Next-Cursor` header. `result.content` is a JSON string literal, so decode it once, then split on headings, keep fenced code whole and tag each chunk with its heading path.

````python title="chunk.py"
import re

def chunk_markdown(md: str, max_chars: int = 2000) -> list[dict]:
    """Split on H1-H3, keep fenced code whole, tag each chunk with its heading path."""
    chunks, path, buf, in_fence = [], [], [], False

    def flush():
        text = "\n".join(buf).strip()
        if text:
            chunks.append({"heading": " > ".join(path), "text": text})
        buf.clear()

    for line in md.splitlines():
        if line.startswith("```"):
            in_fence = not in_fence
        m = None if in_fence else re.match(r"(#{1,3})\s+(.+)", line)
        if m:
            flush()
            path[:] = path[: len(m[1]) - 1] + [m[2].strip()]
        buf.append(line)
        # Character budget, no overlap. Swap in a tokenizer if you need exact token limits.
        if not in_fence and not line.strip() and sum(map(len, buf)) > max_chars:
            flush()
    flush()
    return chunks
````

```python title="ingest.py"
import json  # API, H, requests and chunk_markdown as defined above

def results(job_id: str):
    cursor = None
    while True:
        r = requests.get(f"{API}/v1/batch/{job_id}/results", headers=H, timeout=120,
                         params={"cursor": cursor} if cursor else {})
        r.raise_for_status()
        for line in r.text.splitlines():
            yield json.loads(line)
        cursor = r.headers.get("X-Next-Cursor")
        if not cursor:
            return

for item in results(job["id"]):
    if item["status"] != "succeeded":
        print("not succeeded:", item["url"], item["status"], (item.get("error") or {}).get("code"))
        continue
    if item["http_status"] in (404, 410):          # the page is gone: drop it from the index
        delete_from_index(item["url"])
        continue
    res = item["result"]
    if item["http_status"] != 200 or res.get("truncated") or res.get("omitted") or "content" not in res:
        continue                                    # refused by the site, or the body did not fit inline
    for c in chunk_markdown(json.loads(res["content"])):
        upsert(url=item["url"], heading=c["heading"], text=c["text"], fetched_at=item["finished_at"])
```

`upsert` and `delete_from_index` stand for your vector store's calls.

* **Chunk size.** Embed `heading + "\n\n" + text`, aim for roughly 200 to 500 tokens (2,000 characters is about 500), and add overlap only if retrieval tests show answers split across chunks.
* **Provenance.** Store `url`, `finished_at` and a hash of the page markdown with every chunk, so answers cite a source and a refresh sees what changed.
* **Retries.** Items with `error.retryable: true` re-run with `POST /v1/batch/{id}/retry`; succeeded items never re-run.
* **Oversize.** `result.truncated: true` means the item hit the 512 KiB cap and `content` is not valid JSON; `result.omitted: true` means the page budget ran out, so lower `limit`.
* **PDFs.** Batch fails a non-HTML item with `ERR::EXTRACTION::FAILED` at 0 credits; fetch it with `POST /v1/scrape`, which parses a PDF to text.
* **Discovery.** Submit with `open: true`, add URLs with `POST /v1/batch/{id}/items`, then call `/close`; an open job never completes on its own.

### What does it cost to build a knowledge base?

Markdown, tags, `links` and PDF parsing add no credits.

| Job                                              | Credits                 |
| ------------------------------------------------ | ----------------------- |
| 400-page docs site, every page server-rendered   | 400 (400 x 1)           |
| The same site, every page JavaScript-rendered    | 1,200 (400 x 3)         |
| 300 server-rendered pages and 100 rendered pages | 600 (300 x 1 + 100 x 3) |
| One changelog page with scroll `actions`         | 8                       |

Failed fetches cost 0. Against the beta's 1,000-credit monthly allowance, weekly refreshes of a 400-page plain-fetch site use 1,600 to 2,000 credits a month, so refresh only the sections that change. See [Credits](https://docs.spicrawl.com/credits.md).

## How do I give an AI agent access to the web with MCP?

Add the hosted MCP server and the agent calls `spicrawl_scrape`, `spicrawl_batch_*` and `spicrawl_docs_*` as native tools. For Claude Code:

```bash
claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
  --header "Authorization: Bearer $SPICRAWL_API_KEY"
```

`spicrawl init` configures Claude Code, Cursor, VS Code and Codex in one step; other clients are under [Build with agents](https://docs.spicrawl.com/agents/overview.md) and every tool is on the [MCP server](https://docs.spicrawl.com/agents/mcp.md) page.

* **Status and cost.** MCP tools do not pass response headers, so ask for `format: "json"` and read `status` and `credits` from the envelope.
* **Least privilege.** `spicrawl_scrape` is not read-only (`method` and `actions` can submit forms), so give agents only the tools they need.
* **Live or indexed.** Pre-index slow-changing pages that are asked about often; read live when the answer must reflect the page now.
* **Docs for agents.** Point `CLAUDE.md` or `AGENTS.md` at `https://docs.spicrawl.com/llms.txt` (also `llms-full.txt`, and `.md` on every page). See [Docs for agents](https://docs.spicrawl.com/agents/llms-txt.md).

## How do I keep a knowledge base fresh?

Re-scrape on a schedule you control, skip the cache, and re-embed only pages whose content changed.

* **Single scrapes.** `/v1/scrape` serves a copy younger than `cache_ttl` (48 hours by default) as `Cache-State: hit`. Send `"cache": false` (CLI `--no-cache`) for a live fetch; a hit bills like a fetch, so this costs nothing extra.
* **Batch.** Jobs never read the cache, and `cache` on `POST /v1/batch` is refused with a `400`.
* **Scheduling.** Spicrawl has no scheduler: run the re-scrape from cron or CI, and set `webhook_endpoint_id` for a callback when it finishes.

```bash title="refresh.sh"
#!/usr/bin/env bash
# crontab: 0 3 * * 1 /opt/kb/refresh.sh   (SPICRAWL_API_KEY must be set in the cron environment)
set -euo pipefail
job=$(spicrawl batch submit urls.txt --format markdown --main-content --include article \
        --credit-budget 500 --name "docs-$(date +%F)" --wait --max-wait 60m)
[ "$(jq -r .status <<<"$job")" = completed ] || { echo "$job" >&2; exit 1; }
spicrawl batch results "$(jq -r .id <<<"$job")" --all -o results.jsonl
```

Hash each page's markdown and re-index only a URL whose hash changed; a 404 or 410 means delete its vectors.

```python
import hashlib

digest = hashlib.sha256(markdown.encode()).hexdigest()
if digest != seen.get(item["url"]):        # seen: url -> last digest, kept in your own store
    reindex(item["url"], markdown)
    seen[item["url"]] = digest
```

Every refresh bills every URL, so keep one `urls.txt` per cadence (changelog daily, reference weekly, stable guides monthly); `credit_budget` caps a run.

## Frequently asked questions

**What is LLM-ready data?**

Web content reduced to the main article as markdown, with headings, tables and code blocks intact and navigation, scripts and styles removed. Set `response_format: "markdown"`; it costs the same as the fetch.

**Can Spicrawl crawl an entire website for me?**

No, Spicrawl fetches only the URLs you list. Take them from a sitemap, an `llms.txt` file or a page's `links`, then submit up to 10,000 as one batch job; see [the documentation-site recipe](#how-do-i-scrape-a-whole-documentation-site-for-rag).

**How much does it cost to scrape a documentation site for RAG?**

A page costs 1 credit for a plain fetch or 3 with JavaScript rendering, so a 400-page site is 400 credits, or 1,200 fully rendered; markdown adds nothing and failures cost 0. See [the cost table](#what-does-it-cost-to-build-a-knowledge-base).

**Does Spicrawl work on JavaScript-rendered documentation sites?**

Yes, send `js_render: true` with a `wait_for` selector (3 credits); content behind a click or scroll uses `actions` (8 credits). See [the JavaScript recipe](#how-do-i-handle-javascript-heavy-documentation-sites).

**How do I keep my RAG index up to date?**

Re-scrape from cron or CI, re-embed only pages whose content hash changed, and delete pages that return 404 or 410. See [the refresh recipe](#how-do-i-keep-a-knowledge-base-fresh).

**How do I give an AI agent web access with MCP?**

Connect `https://mcp.spicrawl.com/mcp` with an `Authorization: Bearer` header, or run `spicrawl init`; the agent gets 25 `spicrawl_*` tools. See [the MCP section](#how-do-i-give-an-ai-agent-access-to-the-web-with-mcp).

## Related

* [Markdown for LLMs](https://docs.spicrawl.com/guides/markdown.md), [Batch jobs](https://docs.spicrawl.com/guides/batch.md) and [Caching](https://docs.spicrawl.com/guides/caching.md)
* [JavaScript rendering](https://docs.spicrawl.com/guides/javascript-rendering.md), [Browser actions](https://docs.spicrawl.com/guides/browser-actions.md) and [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md)
* [Agent quickstart](https://docs.spicrawl.com/agents/quickstart.md), [Agent frameworks](https://docs.spicrawl.com/agents/frameworks.md) and [Best practices for agents](https://docs.spicrawl.com/agents/best-practices.md)
* [Credits](https://docs.spicrawl.com/credits.md) and [Rate limits](https://docs.spicrawl.com/rate-limits.md)

 

## Disclaimer

> **For educational purposes:** The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and `robots.txt`, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.
