Turn websites into LLM-ready data for RAG and AI agents
Scrape websites and docs sites to clean markdown for RAG and AI agents: batch ingestion, chunking, the MCP server, llms.txt and scheduled refreshes.
Send a URL with response_format: "markdown" to get clean main-content markdown for embeddings and LLM context. Batch up to 10,000 URLs for a knowledge base, or connect the hosted MCP server for live agent access; Spicrawl does not crawl, chunk or embed, so the URL list and vector store are yours.
How do I turn a web page into LLM-ready markdown?
Send the URL with response_format: "markdown"; include_tags and exclude_tags pick the content container before conversion. See Markdown for LLMs for every option.
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/docs/install",
"response_format": "markdown",
"include_tags": ["article"],
"max_cost": 1
}'A 200 from Spicrawl is not a 200 from the site: a page that answered 403 still returns HTTP 200 at 0 credits, so read X-Target-Status before you index. A 404 or 410 means the page is gone.
How do I handle JavaScript-heavy documentation sites?
A near-empty body with X-Target-Status: 200 means JavaScript builds the page. Add js_render: true and a wait_for selector that exists only once the article is built (3 credits). X-Warning: RENDER_DEGRADED means wait_for never matched, so fix the selector.
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/docs/api/reference",
"js_render": true,
"wait_for": "article h1",
"response_format": "markdown",
"include_tags": ["article"],
"max_cost": 3
}'- Probe first. Fetch one page per section plain (1 credit) and render only the sections that come back empty. In a batch, set
js_renderon the job or on individualitems. - Interaction. Content behind "Load more", a tab or infinite scroll needs
actions(8 credits), which batch refuses: use/v1/scrape. - Bot walls.
ERR::UPSTREAM::CHALLENGE(retryable, 0 credits) means a challenge; route through your own residentialproxy, which is likely to work but not guaranteed. See Cloudflare-protected pages.
How do I scrape a whole documentation site for RAG?
Spicrawl has no crawl endpoint, so you collect the URLs, submit one batch job and page the results into your chunker.
Collect the URLs
Take them from a sitemap, an llms.txt file (a markdown list of a site's pages for LLMs) or a page's links (links: true, no extra credits, JSON envelope). Fetch sitemaps and llms.txt with response_format: "html", which returns the file as served.
import os, re, requests
API = "https://api.spicrawl.com"
H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
def fetch_html(url: str) -> str:
# "html" returns XML and llms.txt as served; markdown refuses non-HTML documents
r = requests.post(f"{API}/v1/scrape", headers=H, timeout=120,
json={"url": url, "response_format": "html", "max_cost": 1})
r.raise_for_status()
return r.text
# A sitemap index lists other sitemaps: fetch each in turn.
locs = re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", fetch_html("https://example.com/sitemap.xml"))
# llms.txt instead: re.findall(r"\]\((https?://[^)\s]+)\)", fetch_html("https://example.com/llms.txt"))
urls = sorted({u for u in locs if u.startswith("https://example.com/docs/")})Submit one batch job
Set main_content_only: true and the content selector once for the job, and add a credit_budget so a wrong URL list cannot overspend.
import time
job = requests.post(f"{API}/v1/batch", headers=H, timeout=60, json={
"name": "docs-example-2026-10-09",
"urls": urls,
"response_format": "markdown",
"main_content_only": True,
"include_tags": ["article"],
"exclude_tags": [".cookie-banner", "nav.toc"],
"credit_budget": 200,
}).json()
while job["status"] not in ("completed", "failed", "cancelled"):
time.sleep(10)
job = requests.get(f"{API}/v1/batch/{job['id']}", headers=H, timeout=30).json()Submitting returns 202 with estimated_credits (a hold, not a charge). Do not resubmit if warnings mentions unqueued items: a second submission is a second job and bills again.
Read the results into your chunker
GET /v1/batch/{id}/results returns one JSON object per line and pages with the X-Next-Cursor header. result.content is a JSON string literal, so decode it once, then split on headings, keep fenced code whole and tag each chunk with its heading path.
import re
def chunk_markdown(md: str, max_chars: int = 2000) -> list[dict]:
"""Split on H1-H3, keep fenced code whole, tag each chunk with its heading path."""
chunks, path, buf, in_fence = [], [], [], False
def flush():
text = "\n".join(buf).strip()
if text:
chunks.append({"heading": " > ".join(path), "text": text})
buf.clear()
for line in md.splitlines():
if line.startswith("```"):
in_fence = not in_fence
m = None if in_fence else re.match(r"(#{1,3})\s+(.+)", line)
if m:
flush()
path[:] = path[: len(m[1]) - 1] + [m[2].strip()]
buf.append(line)
# Character budget, no overlap. Swap in a tokenizer if you need exact token limits.
if not in_fence and not line.strip() and sum(map(len, buf)) > max_chars:
flush()
flush()
return chunksimport json # API, H, requests and chunk_markdown as defined above
def results(job_id: str):
cursor = None
while True:
r = requests.get(f"{API}/v1/batch/{job_id}/results", headers=H, timeout=120,
params={"cursor": cursor} if cursor else {})
r.raise_for_status()
for line in r.text.splitlines():
yield json.loads(line)
cursor = r.headers.get("X-Next-Cursor")
if not cursor:
return
for item in results(job["id"]):
if item["status"] != "succeeded":
print("not succeeded:", item["url"], item["status"], (item.get("error") or {}).get("code"))
continue
if item["http_status"] in (404, 410): # the page is gone: drop it from the index
delete_from_index(item["url"])
continue
res = item["result"]
if item["http_status"] != 200 or res.get("truncated") or res.get("omitted") or "content" not in res:
continue # refused by the site, or the body did not fit inline
for c in chunk_markdown(json.loads(res["content"])):
upsert(url=item["url"], heading=c["heading"], text=c["text"], fetched_at=item["finished_at"])upsert and delete_from_index stand for your vector store's calls.
- Chunk size. Embed
heading + "\n\n" + text, aim for roughly 200 to 500 tokens (2,000 characters is about 500), and add overlap only if retrieval tests show answers split across chunks. - Provenance. Store
url,finished_atand a hash of the page markdown with every chunk, so answers cite a source and a refresh sees what changed. - Retries. Items with
error.retryable: truere-run withPOST /v1/batch/{id}/retry; succeeded items never re-run. - Oversize.
result.truncated: truemeans the item hit the 512 KiB cap andcontentis not valid JSON;result.omitted: truemeans the page budget ran out, so lowerlimit. - PDFs. Batch fails a non-HTML item with
ERR::EXTRACTION::FAILEDat 0 credits; fetch it withPOST /v1/scrape, which parses a PDF to text. - Discovery. Submit with
open: true, add URLs withPOST /v1/batch/{id}/items, then call/close; an open job never completes on its own.
What does it cost to build a knowledge base?
Markdown, tags, links and PDF parsing add no credits.
| Job | Credits |
|---|---|
| 400-page docs site, every page server-rendered | 400 (400 x 1) |
| The same site, every page JavaScript-rendered | 1,200 (400 x 3) |
| 300 server-rendered pages and 100 rendered pages | 600 (300 x 1 + 100 x 3) |
One changelog page with scroll actions | 8 |
Failed fetches cost 0. Against the beta's 1,000-credit monthly allowance, weekly refreshes of a 400-page plain-fetch site use 1,600 to 2,000 credits a month, so refresh only the sections that change. See Credits.
How do I give an AI agent access to the web with MCP?
Add the hosted MCP server and the agent calls spicrawl_scrape, spicrawl_batch_* and spicrawl_docs_* as native tools. For Claude Code:
claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
--header "Authorization: Bearer $SPICRAWL_API_KEY"spicrawl init configures Claude Code, Cursor, VS Code and Codex in one step; other clients are under Build with agents and every tool is on the MCP server page.
- Status and cost. MCP tools do not pass response headers, so ask for
format: "json"and readstatusandcreditsfrom the envelope. - Least privilege.
spicrawl_scrapeis not read-only (methodandactionscan submit forms), so give agents only the tools they need. - Live or indexed. Pre-index slow-changing pages that are asked about often; read live when the answer must reflect the page now.
- Docs for agents. Point
CLAUDE.mdorAGENTS.mdathttps://docs.spicrawl.com/llms.txt(alsollms-full.txt, and.mdon every page). See Docs for agents.
How do I keep a knowledge base fresh?
Re-scrape on a schedule you control, skip the cache, and re-embed only pages whose content changed.
- Single scrapes.
/v1/scrapeserves a copy younger thancache_ttl(48 hours by default) asCache-State: hit. Send"cache": false(CLI--no-cache) for a live fetch; a hit bills like a fetch, so this costs nothing extra. - Batch. Jobs never read the cache, and
cacheonPOST /v1/batchis refused with a400. - Scheduling. Spicrawl has no scheduler: run the re-scrape from cron or CI, and set
webhook_endpoint_idfor a callback when it finishes.
#!/usr/bin/env bash
# crontab: 0 3 * * 1 /opt/kb/refresh.sh (SPICRAWL_API_KEY must be set in the cron environment)
set -euo pipefail
job=$(spicrawl batch submit urls.txt --format markdown --main-content --include article \
--credit-budget 500 --name "docs-$(date +%F)" --wait --max-wait 60m)
[ "$(jq -r .status <<<"$job")" = completed ] || { echo "$job" >&2; exit 1; }
spicrawl batch results "$(jq -r .id <<<"$job")" --all -o results.jsonlHash each page's markdown and re-index only a URL whose hash changed; a 404 or 410 means delete its vectors.
import hashlib
digest = hashlib.sha256(markdown.encode()).hexdigest()
if digest != seen.get(item["url"]): # seen: url -> last digest, kept in your own store
reindex(item["url"], markdown)
seen[item["url"]] = digestEvery refresh bills every URL, so keep one urls.txt per cadence (changelog daily, reference weekly, stable guides monthly); credit_budget caps a run.
Frequently asked questions
Web content reduced to the main article as markdown, with headings, tables and code blocks intact and navigation, scripts and styles removed. Set response_format: "markdown"; it costs the same as the fetch.
No, Spicrawl fetches only the URLs you list. Take them from a sitemap, an llms.txt file or a page's links, then submit up to 10,000 as one batch job; see the documentation-site recipe.
A page costs 1 credit for a plain fetch or 3 with JavaScript rendering, so a 400-page site is 400 credits, or 1,200 fully rendered; markdown adds nothing and failures cost 0. See the cost table.
Yes, send js_render: true with a wait_for selector (3 credits); content behind a click or scroll uses actions (8 credits). See the JavaScript recipe.
Re-scrape from cron or CI, re-embed only pages whose content hash changed, and delete pages that return 404 or 410. See the refresh recipe.
Connect https://mcp.spicrawl.com/mcp with an Authorization: Bearer header, or run spicrawl init; the agent gets 25 spicrawl_* tools. See the MCP section.
Related
- Markdown for LLMs, Batch jobs and Caching
- JavaScript rendering, Browser actions and Anti-bot
- Agent quickstart, Agent frameworks and Best practices for agents
- Credits and Rate limits
Start scraping in minutes
1,000 free credits every month. One API key, one request.
Disclaimer
For educational purposes
The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and robots.txt, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.
Protected sites
Bypass Cloudflare and PerimeterX blocks when scraping with one API call. Automatic challenge handling on your own proxy, billed only on success.
Lead generation
Scrape company directories, startup lists and listing sites into structured leads (name, website, location, size), in batches, exported to CSV or a CRM.