spicrawlspicrawlDocs

Turn websites into LLM-ready data for RAG and AI agents

Scrape websites and docs sites to clean markdown for RAG and AI agents: batch ingestion, chunking, the MCP server, llms.txt and scheduled refreshes.

Send a URL with response_format: "markdown" to get clean main-content markdown for embeddings and LLM context. Batch up to 10,000 URLs for a knowledge base, or connect the hosted MCP server for live agent access; Spicrawl does not crawl, chunk or embed, so the URL list and vector store are yours.

How do I turn a web page into LLM-ready markdown?

Send the URL with response_format: "markdown"; include_tags and exclude_tags pick the content container before conversion. See Markdown for LLMs for every option.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/docs/install",
    "response_format": "markdown",
    "include_tags": ["article"],
    "max_cost": 1
  }'

A 200 from Spicrawl is not a 200 from the site: a page that answered 403 still returns HTTP 200 at 0 credits, so read X-Target-Status before you index. A 404 or 410 means the page is gone.

How do I handle JavaScript-heavy documentation sites?

A near-empty body with X-Target-Status: 200 means JavaScript builds the page. Add js_render: true and a wait_for selector that exists only once the article is built (3 credits). X-Warning: RENDER_DEGRADED means wait_for never matched, so fix the selector.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/docs/api/reference",
    "js_render": true,
    "wait_for": "article h1",
    "response_format": "markdown",
    "include_tags": ["article"],
    "max_cost": 3
  }'
  • Probe first. Fetch one page per section plain (1 credit) and render only the sections that come back empty. In a batch, set js_render on the job or on individual items.
  • Interaction. Content behind "Load more", a tab or infinite scroll needs actions (8 credits), which batch refuses: use /v1/scrape.
  • Bot walls. ERR::UPSTREAM::CHALLENGE (retryable, 0 credits) means a challenge; route through your own residential proxy, which is likely to work but not guaranteed. See Cloudflare-protected pages.

How do I scrape a whole documentation site for RAG?

Spicrawl has no crawl endpoint, so you collect the URLs, submit one batch job and page the results into your chunker.

Collect the URLs

Take them from a sitemap, an llms.txt file (a markdown list of a site's pages for LLMs) or a page's links (links: true, no extra credits, JSON envelope). Fetch sitemaps and llms.txt with response_format: "html", which returns the file as served.

urls.py
import os, re, requests

API = "https://api.spicrawl.com"
H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}

def fetch_html(url: str) -> str:
    # "html" returns XML and llms.txt as served; markdown refuses non-HTML documents
    r = requests.post(f"{API}/v1/scrape", headers=H, timeout=120,
                      json={"url": url, "response_format": "html", "max_cost": 1})
    r.raise_for_status()
    return r.text

# A sitemap index lists other sitemaps: fetch each in turn.
locs = re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", fetch_html("https://example.com/sitemap.xml"))
# llms.txt instead: re.findall(r"\]\((https?://[^)\s]+)\)", fetch_html("https://example.com/llms.txt"))
urls = sorted({u for u in locs if u.startswith("https://example.com/docs/")})

Submit one batch job

Set main_content_only: true and the content selector once for the job, and add a credit_budget so a wrong URL list cannot overspend.

import time

job = requests.post(f"{API}/v1/batch", headers=H, timeout=60, json={
    "name": "docs-example-2026-10-09",
    "urls": urls,
    "response_format": "markdown",
    "main_content_only": True,
    "include_tags": ["article"],
    "exclude_tags": [".cookie-banner", "nav.toc"],
    "credit_budget": 200,
}).json()

while job["status"] not in ("completed", "failed", "cancelled"):
    time.sleep(10)
    job = requests.get(f"{API}/v1/batch/{job['id']}", headers=H, timeout=30).json()

Submitting returns 202 with estimated_credits (a hold, not a charge). Do not resubmit if warnings mentions unqueued items: a second submission is a second job and bills again.

Read the results into your chunker

GET /v1/batch/{id}/results returns one JSON object per line and pages with the X-Next-Cursor header. result.content is a JSON string literal, so decode it once, then split on headings, keep fenced code whole and tag each chunk with its heading path.

chunk.py
import re

def chunk_markdown(md: str, max_chars: int = 2000) -> list[dict]:
    """Split on H1-H3, keep fenced code whole, tag each chunk with its heading path."""
    chunks, path, buf, in_fence = [], [], [], False

    def flush():
        text = "\n".join(buf).strip()
        if text:
            chunks.append({"heading": " > ".join(path), "text": text})
        buf.clear()

    for line in md.splitlines():
        if line.startswith("```"):
            in_fence = not in_fence
        m = None if in_fence else re.match(r"(#{1,3})\s+(.+)", line)
        if m:
            flush()
            path[:] = path[: len(m[1]) - 1] + [m[2].strip()]
        buf.append(line)
        # Character budget, no overlap. Swap in a tokenizer if you need exact token limits.
        if not in_fence and not line.strip() and sum(map(len, buf)) > max_chars:
            flush()
    flush()
    return chunks
ingest.py
import json  # API, H, requests and chunk_markdown as defined above

def results(job_id: str):
    cursor = None
    while True:
        r = requests.get(f"{API}/v1/batch/{job_id}/results", headers=H, timeout=120,
                         params={"cursor": cursor} if cursor else {})
        r.raise_for_status()
        for line in r.text.splitlines():
            yield json.loads(line)
        cursor = r.headers.get("X-Next-Cursor")
        if not cursor:
            return

for item in results(job["id"]):
    if item["status"] != "succeeded":
        print("not succeeded:", item["url"], item["status"], (item.get("error") or {}).get("code"))
        continue
    if item["http_status"] in (404, 410):          # the page is gone: drop it from the index
        delete_from_index(item["url"])
        continue
    res = item["result"]
    if item["http_status"] != 200 or res.get("truncated") or res.get("omitted") or "content" not in res:
        continue                                    # refused by the site, or the body did not fit inline
    for c in chunk_markdown(json.loads(res["content"])):
        upsert(url=item["url"], heading=c["heading"], text=c["text"], fetched_at=item["finished_at"])

upsert and delete_from_index stand for your vector store's calls.

  • Chunk size. Embed heading + "\n\n" + text, aim for roughly 200 to 500 tokens (2,000 characters is about 500), and add overlap only if retrieval tests show answers split across chunks.
  • Provenance. Store url, finished_at and a hash of the page markdown with every chunk, so answers cite a source and a refresh sees what changed.
  • Retries. Items with error.retryable: true re-run with POST /v1/batch/{id}/retry; succeeded items never re-run.
  • Oversize. result.truncated: true means the item hit the 512 KiB cap and content is not valid JSON; result.omitted: true means the page budget ran out, so lower limit.
  • PDFs. Batch fails a non-HTML item with ERR::EXTRACTION::FAILED at 0 credits; fetch it with POST /v1/scrape, which parses a PDF to text.
  • Discovery. Submit with open: true, add URLs with POST /v1/batch/{id}/items, then call /close; an open job never completes on its own.

What does it cost to build a knowledge base?

Markdown, tags, links and PDF parsing add no credits.

JobCredits
400-page docs site, every page server-rendered400 (400 x 1)
The same site, every page JavaScript-rendered1,200 (400 x 3)
300 server-rendered pages and 100 rendered pages600 (300 x 1 + 100 x 3)
One changelog page with scroll actions8

Failed fetches cost 0. Against the beta's 1,000-credit monthly allowance, weekly refreshes of a 400-page plain-fetch site use 1,600 to 2,000 credits a month, so refresh only the sections that change. See Credits.

How do I give an AI agent access to the web with MCP?

Add the hosted MCP server and the agent calls spicrawl_scrape, spicrawl_batch_* and spicrawl_docs_* as native tools. For Claude Code:

claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
  --header "Authorization: Bearer $SPICRAWL_API_KEY"

spicrawl init configures Claude Code, Cursor, VS Code and Codex in one step; other clients are under Build with agents and every tool is on the MCP server page.

  • Status and cost. MCP tools do not pass response headers, so ask for format: "json" and read status and credits from the envelope.
  • Least privilege. spicrawl_scrape is not read-only (method and actions can submit forms), so give agents only the tools they need.
  • Live or indexed. Pre-index slow-changing pages that are asked about often; read live when the answer must reflect the page now.
  • Docs for agents. Point CLAUDE.md or AGENTS.md at https://docs.spicrawl.com/llms.txt (also llms-full.txt, and .md on every page). See Docs for agents.

How do I keep a knowledge base fresh?

Re-scrape on a schedule you control, skip the cache, and re-embed only pages whose content changed.

  • Single scrapes. /v1/scrape serves a copy younger than cache_ttl (48 hours by default) as Cache-State: hit. Send "cache": false (CLI --no-cache) for a live fetch; a hit bills like a fetch, so this costs nothing extra.
  • Batch. Jobs never read the cache, and cache on POST /v1/batch is refused with a 400.
  • Scheduling. Spicrawl has no scheduler: run the re-scrape from cron or CI, and set webhook_endpoint_id for a callback when it finishes.
refresh.sh
#!/usr/bin/env bash
# crontab: 0 3 * * 1 /opt/kb/refresh.sh   (SPICRAWL_API_KEY must be set in the cron environment)
set -euo pipefail
job=$(spicrawl batch submit urls.txt --format markdown --main-content --include article \
        --credit-budget 500 --name "docs-$(date +%F)" --wait --max-wait 60m)
[ "$(jq -r .status <<<"$job")" = completed ] || { echo "$job" >&2; exit 1; }
spicrawl batch results "$(jq -r .id <<<"$job")" --all -o results.jsonl

Hash each page's markdown and re-index only a URL whose hash changed; a 404 or 410 means delete its vectors.

import hashlib

digest = hashlib.sha256(markdown.encode()).hexdigest()
if digest != seen.get(item["url"]):        # seen: url -> last digest, kept in your own store
    reindex(item["url"], markdown)
    seen[item["url"]] = digest

Every refresh bills every URL, so keep one urls.txt per cadence (changelog daily, reference weekly, stable guides monthly); credit_budget caps a run.

Frequently asked questions

Start scraping in minutes

1,000 free credits every month. One API key, one request.

Disclaimer

For educational purposes

The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and robots.txt, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.

On this page