# Spicrawl > Web data API for agents: fetch, render and extract any URL as markdown or structured JSON. One HTTPS call (`POST /v1/scrape`) turns a URL into markdown, HTML, text, JSON or a screenshot; Spicrawl picks the engine, proxy and retries. Authenticate with `Authorization: Bearer $SPICRAWL_API_KEY`. Each link below is the Markdown version of a docs page; fetch only the pages you need. The OpenAPI spec is the source of truth for request and response fields. Every page is also available as HTML at the same URL without `.md`, and all pages concatenated are at https://docs.spicrawl.com/llms-full.txt. ## Get started - [Spicrawl](https://docs.spicrawl.com/index.md): Spicrawl is a web data API for agents: fetch, render and extract any URL as markdown, HTML or structured JSON with one request. - [Quickstart](https://docs.spicrawl.com/quickstart.md): Get a Spicrawl API key, scrape your first page as markdown with curl, Python, TypeScript or the CLI, and read the response headers that matter. - [Authentication](https://docs.spicrawl.com/authentication.md): Authenticate Spicrawl API calls with a Bearer API key, choose live or test keys, and grant the scopes each route needs. - [Credits](https://docs.spicrawl.com/credits.md): What each Spicrawl request costs by engine, what is free, and how to cap spend with max_cost and batch budgets. - [Errors](https://docs.spicrawl.com/errors.md): Every Spicrawl error code with its HTTP status, whether to retry, what it means and which parameter to change. - [Rate limits](https://docs.spicrawl.com/rate-limits.md): How Spicrawl limits request rate, concurrency, live sessions and monthly credits, the headers that report each limit, and how to back off. - [Response headers](https://docs.spicrawl.com/response-headers.md): Every header Spicrawl returns on /v1/scrape: request id, credits, engine, target status, proxy, cache state, limits and X-Warning codes. ## Build with agents - [Build with agents](https://docs.spicrawl.com/agents/overview.md): The four ways an AI agent can use Spicrawl (MCP server, agent skill, CLI, raw HTTP), when to pick each, and how agents read these docs. - [Agent quickstart](https://docs.spicrawl.com/agents/quickstart.md): Connect your coding agent to Spicrawl in one command with spicrawl init, or by pasting one self-contained setup prompt into the agent. - [MCP server](https://docs.spicrawl.com/agents/mcp.md): Connect any MCP client to the hosted Spicrawl MCP server at https://mcp.spicrawl.com/mcp and use its 25 spicrawl_* tools to scrape, batch, manage sessions, inspect usage and search the docs. - [Agent skill](https://docs.spicrawl.com/agents/skill.md): Install the Spicrawl agent skill (SKILL.md from https://app.spicrawl.com/skill.md) so a coding agent writes correct calls to the Spicrawl HTTP API. - [Docs for agents](https://docs.spicrawl.com/agents/llms-txt.md): How an AI agent should read the Spicrawl docs: /llms.txt, /llms-full.txt, a .md version of every page, the search API, the OpenAPI spec, the Spicrawl MCP server's docs tools, and spicrawl docs in a terminal. - [Best practices for agents](https://docs.spicrawl.com/agents/best-practices.md): Rules for agents built on Spicrawl: an error-handling loop, a cheapest-first escalation ladder, cost and token control, result verification, and a reference fetch_page tool in Python and TypeScript. ## Coding agents - [Claude Code](https://docs.spicrawl.com/agents/claude-code.md): Set up Spicrawl in Claude Code: add the hosted MCP server, install the agent skill, and add CLAUDE.md rules so the agent fetches web pages cheaply and handles errors correctly. - [Cursor](https://docs.spicrawl.com/agents/cursor.md): Set up Spicrawl in Cursor: add the hosted MCP server to .cursor/mcp.json, install the agent skill, and add a project rule so the agent fetches web pages cheaply and handles errors correctly. - [Codex](https://docs.spicrawl.com/agents/codex.md): Set up Spicrawl in the Codex CLI: add the hosted MCP server to config.toml, install the agent skill, and add AGENTS.md rules so the agent fetches web pages cheaply and handles errors correctly. - [GitHub Copilot and VS Code](https://docs.spicrawl.com/agents/github-copilot.md): Set up Spicrawl in VS Code agent mode, the Copilot CLI and the Copilot cloud agent: add the hosted MCP server, install the agent skill, and add repository instructions. - [OpenCode](https://docs.spicrawl.com/agents/opencode.md): Set up Spicrawl in OpenCode: add the hosted MCP server to opencode.json with OAuth turned off, install the agent skill, and add AGENTS.md rules so the agent fetches web pages cheaply and handles errors correctly. - [Gemini CLI](https://docs.spicrawl.com/agents/gemini-cli.md): Set up Spicrawl in Google's Gemini CLI: add the hosted MCP server to settings.json, install the agent skill, and add GEMINI.md rules so the agent fetches web pages cheaply and handles errors correctly. - [Windsurf (Devin Desktop)](https://docs.spicrawl.com/agents/windsurf.md): Set up Spicrawl in Windsurf, now named Devin Desktop: add the hosted MCP server, install the agent skill, and add a rule so the agent fetches web pages cheaply and handles errors correctly. - [Cline](https://docs.spicrawl.com/agents/cline.md): Set up Spicrawl in Cline, the VS Code extension: add the hosted MCP server to cline_mcp_settings.json, install the agent skill, and add a rule so the agent fetches web pages cheaply and handles errors correctly. - [Kilo Code](https://docs.spicrawl.com/agents/kilo-code.md): Set up Spicrawl in Kilo Code: add the hosted MCP server to kilo.jsonc, install the agent skill, and add AGENTS.md rules so the agent fetches web pages cheaply and handles errors correctly. - [Zed](https://docs.spicrawl.com/agents/zed.md): Set up Spicrawl in Zed: add the hosted MCP server to context_servers in settings.json, install the agent skill, and add AGENTS.md instructions so the agent fetches web pages cheaply and handles errors correctly. - [JetBrains AI Assistant and Junie](https://docs.spicrawl.com/agents/jetbrains.md): Set up Spicrawl in JetBrains IDEs: add the hosted MCP server to AI Assistant and Junie, install the agent skill, and add project instructions so the agent fetches web pages cheaply and handles errors correctly. - [Goose](https://docs.spicrawl.com/agents/goose.md): Set up Spicrawl in Goose: add the hosted MCP server as a Streamable HTTP extension, install the agent skill, and add AGENTS.md rules so the agent fetches web pages cheaply and handles errors correctly. - [Kiro](https://docs.spicrawl.com/agents/kiro.md): Set up Spicrawl in Kiro: add the hosted MCP server to .kiro/settings/mcp.json, install the agent skill, and add a steering file so the agent fetches web pages cheaply and handles errors correctly. - [Factory Droid](https://docs.spicrawl.com/agents/factory.md): Set up Spicrawl in Factory's Droid CLI: add the hosted MCP server with an API-key header, install the agent skill, and add AGENTS.md rules so the agent fetches web pages cheaply and handles errors correctly. - [More MCP clients](https://docs.spicrawl.com/agents/other-clients.md): Connect Crush, Amp, Kimi Code CLI, Warp, Qwen Code, Continue, Augment, Amazon Q Developer and Google Antigravity to the hosted Spicrawl MCP server at https://mcp.spicrawl.com/mcp. ## Cloud agents and app builders - [Devin](https://docs.spicrawl.com/agents/devin.md): Set up Spicrawl in Devin: add the hosted MCP server as a custom MCP with an Auth Header, store the key in Devin Secrets, and give Devin a skill so it fetches web pages cheaply and handles errors correctly. - [OpenHands](https://docs.spicrawl.com/agents/openhands.md): Set up Spicrawl in OpenHands (web UI, CLI or SDK): add the hosted MCP server with your API key, install the agent skill, and add AGENTS.md rules so the agent fetches web pages cheaply and handles errors correctly. - [AI app builders](https://docs.spicrawl.com/agents/app-builders.md): Connect Spicrawl's hosted MCP server to Replit Agent, v0, Lovable and Bolt so the builder can read web pages, and call the Spicrawl API from the app it builds. ## Agent frameworks - [Agent frameworks](https://docs.spicrawl.com/agents/frameworks.md): Give an agent built with the OpenAI Agents SDK, Vercel AI SDK, LangChain, CrewAI, Google ADK, Mastra or smolagents the 25 spicrawl_* tools by connecting to the hosted Spicrawl MCP server. ## Guides - [Get clean markdown for an LLM](https://docs.spicrawl.com/guides/markdown.md): Turn any page or PDF into main-content markdown with response_format=markdown, and cut tokens with include_tags, exclude_tags and main_content_only. - [Render JavaScript pages](https://docs.spicrawl.com/guides/javascript-rendering.md): Run a page in a browser with js_render, pick an engine, wait for content with wait and wait_for, and let mode=auto escalate only when it has to. - [Get past anti-bot challenges](https://docs.spicrawl.com/guides/anti-bot.md): What ERR::UPSTREAM::CHALLENGE means, the escalation ladder from a browser-fingerprinted fetch to rendering to your own proxy with the credit cost of each rung, the automatic climb to the stealth browser, and how to read blocks in your request logs. - [Use your own proxy](https://docs.spicrawl.com/guides/proxies-and-geo.md): Route requests through your own proxy with proxy and proxy_verify, and read how a request was routed. Spicrawl's managed proxy pool is coming soon. - [Extract structured data](https://docs.spicrawl.com/guides/structured-data.md): Get JSON out of a page three ways: autoparse for the page's own metadata, a selector map for exact fields, or a JSON Schema with selectors for typed, validated output. - [Extract data with a model (coming soon)](https://docs.spicrawl.com/guides/ai-extraction.md): Coming soon: describe the data in plain language or as a JSON Schema with ai_extract and get it back under data, for 4 credits on top of the engine price. - [Capture the page's own API calls](https://docs.spicrawl.com/guides/network-capture.md): Record the XHR and fetch responses a page makes while it renders with network_capture, and read their JSON bodies from the network array. - [Capture screenshots and PDFs](https://docs.spicrawl.com/guides/screenshots-and-pdf.md): Capture a page as a PNG, JPEG or WebP screenshot with screenshot=true, or print it to PDF with response_format=pdf on the chromium engine. - [Click, type and scroll with browser actions](https://docs.spicrawl.com/guides/browser-actions.md): Drive a rendered page with the actions array: wait_for, wait_for_navigation, click, fill, select, scroll, evaluate and screenshot steps, up to 50 per request, before the page is captured. - [Stay logged in with sessions](https://docs.spicrawl.com/guides/sessions-and-logins.md): Create a session with POST /v1/sessions, pass its id as session_id on /v1/scrape to reuse cookies, storage and a pinned exit IP, and release it when you are done. - [Scrape thousands of URLs with a batch job](https://docs.spicrawl.com/guides/batch.md): Submit up to 10,000 URLs per call to POST /v1/batch, poll the job until it is terminal, and page the JSONL results before they expire after 72 hours. - [Drive a cloud browser over CDP (coming soon)](https://docs.spicrawl.com/guides/cdp-browser.md): Coming soon: connect Puppeteer or Playwright to a browser hosted by Spicrawl. - [Cache results and skip repeat charges](https://docs.spicrawl.com/guides/caching.md): How the /v1/scrape result cache works: on by default with a 48-hour window, cache hits cost 0 credits, and cache=false or cache_ttl=0 forces a fresh fetch. ## CLI - [Spicrawl CLI](https://docs.spicrawl.com/cli/overview.md): The spicrawl command-line client scrapes, renders and extracts web pages, and prints JSON when piped so scripts and AI agents can use it directly. - [Install the CLI](https://docs.spicrawl.com/cli/install.md): Install the spicrawl binary with npm or npx, then verify, upgrade or remove it. The curl installer and Homebrew are coming soon. - [CLI authentication and config](https://docs.spicrawl.com/cli/authentication.md): Save an API key with spicrawl login, inspect it with spicrawl auth status, edit the config file, and pass the key through the environment in CI. - [spicrawl scrape](https://docs.spicrawl.com/cli/scrape.md): Fetch one URL, or many from stdin, as markdown, html, text, PDF or extracted JSON with spicrawl scrape; every flag maps to a POST /v1/scrape field. - [spicrawl batch](https://docs.spicrawl.com/cli/batch.md): Submit up to 10,000 URLs as one asynchronous batch job, wait for it, and download the results as JSON Lines with spicrawl batch. - [Sessions and cloud browsers](https://docs.spicrawl.com/cli/sessions-and-browser.md): Keep cookies and engine across scrapes with spicrawl sessions. Cloud browsers (spicrawl browser url) are coming soon. - [Logs, usage and status](https://docs.spicrawl.com/cli/logs-and-usage.md): Inspect recent requests with spicrawl logs, read billed usage with spicrawl usage, and check the API end to end with spicrawl status. - [Set up AI agents with the CLI](https://docs.spicrawl.com/cli/agent-setup.md): Connect Claude Code, Cursor, VS Code and Codex to Spicrawl with spicrawl init, mcp install and skill install, and give agents docs and request schemas offline with spicrawl docs and spicrawl schema. - [Scripting with the CLI](https://docs.spicrawl.com/cli/scripting.md): Recipes for using spicrawl in shell scripts, agent tool calls and CI: jq pipelines, URL files, exit-code branching, retries and per-URL files. - [CLI exit codes](https://docs.spicrawl.com/cli/exit-codes.md): The stable exit codes of the spicrawl CLI, which API error codes map to each, and the JSON error written to stderr on failure. ## API reference - [API reference](https://docs.spicrawl.com/api-reference/introduction.md): Base URL, authentication, content types, strict JSON, the GET and POST forms of /v1/scrape, pagination and retries for the Spicrawl REST API. ## API reference: Scrape - [Scrape a URL (query-string form)](https://docs.spicrawl.com/api-reference/scrape/scrape-get.md): Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data. - [Scrape a URL](https://docs.spicrawl.com/api-reference/scrape/scrape-post.md): Fetch one URL synchronously and return it as html, markdown, text, a PDF, or a JSON envelope with extracted data. ## API reference: Batch - [List batch jobs](https://docs.spicrawl.com/api-reference/batch/batch-list.md): Lists this API key's project's jobs, newest first (`submitted_at` desc, then id desc). - [Submit a batch job](https://docs.spicrawl.com/api-reference/batch/batch-create.md): Queues many URLs as one asynchronous job and returns `202` with the job object immediately. - [Get a batch job and its progress](https://docs.spicrawl.com/api-reference/batch/batch-get.md): Returns the job with live progress counters. Cheap to call: progress is read from sharded counters, not by counting items. - [Page through a job's finished items (JSONL)](https://docs.spicrawl.com/api-reference/batch/batch-results.md): Streams one page of FINISHED items (status succeeded, failed, cancelled or skipped) in `seq` order as JSON Lines: one `BatchResultLine` object per line, each terminated by `\n`, no array, no trailing pagination object. - [Get one item's payload](https://docs.spicrawl.com/api-reference/batch/batch-task-content.md): Returns a single finished item's payload verbatim (not wrapped), capped at 512 KiB. - [Cancel a batch job](https://docs.spicrawl.com/api-reference/batch/batch-cancel.md): Requests cancellation of a `queued`, `running`, `paused` or `cancelling` job; takes no body. - [Re-run a job's failed items](https://docs.spicrawl.com/api-reference/batch/batch-retry.md): Resets every item with status `failed` back to queued and re-dispatches it; takes no body. - [Append items to an open job](https://docs.spicrawl.com/api-reference/batch/batch-append.md): Adds items to a job created with `open: true`. Same body shape as batchCreate: `urls` or `items`, at most 10,000 per call; a job holds at most 100,000 items across all appends (exceeding it is a 400). - [Stop an open job accepting items](https://docs.spicrawl.com/api-reference/batch/batch-close.md): Sets `open` to false so the job completes once its queued work drains; takes no body. ## API reference: Sessions - [List the key's project's sessions, newest first](https://docs.spicrawl.com/api-reference/sessions/sessions-list.md): Lists sessions in the API key's project (not the whole organization), ordered by `created_at` then `id`, descending. - [Create a persisted browser session](https://docs.spicrawl.com/api-reference/sessions/sessions-create.md): Creates a session: a sealed cookie jar and web-storage snapshot, pinned to one engine and one exit intent, that you reuse by passing its `id` as `session_id` on `/v1/scrape`. - [Get one session's metadata](https://docs.spicrawl.com/api-reference/sessions/sessions-get.md): Returns the session's metadata and a non-secret context summary. - [Delete a session and its context](https://docs.spicrawl.com/api-reference/sessions/sessions-delete.md): Removes the session row and purges its context permanently. Unlike release, the id stops resolving afterwards. - [Dump the session's cookies and storage](https://docs.spicrawl.com/api-reference/sessions/sessions-get-context.md): The only endpoint that returns credentials. Works only while the session is `active`. - [End a session and purge its context](https://docs.spicrawl.com/api-reference/sessions/sessions-release.md): Marks the session `released` and deletes its stored cookies and storage in the same transaction. ## API reference: Browser - [Open a Chrome DevTools Protocol WebSocket to a cloud browser (coming soon)](https://docs.spicrawl.com/api-reference/browser/browser-connect.md): A WebSocket upgrade, not a JSON endpoint. Connect with `puppeteer.connect({browserWSEndpoint})` or `chromium.connectOverCDP(url)`. - [Mint a single-use connect URL for /v1/browser (coming soon)](https://docs.spicrawl.com/api-reference/browser/browser-token.md): Returns a `/v1/browser` URL that carries a short-lived token instead of the API key, for CDP clients that accept only a URL. ## API reference: Requests - [List recent requests in the key's project, or in every project of its organization](https://docs.spicrawl.com/api-reference/requests/requests-list.md): Per-request records from the ephemeral request log, newest first, for the API key's project (every key in that project, not only this key). - [Get one request record](https://docs.spicrawl.com/api-reference/requests/requests-get.md): Returns one record by request id (the `X-Request-Id` / `request_id` you received). ## API reference: Usage - [Billed usage over a window, broken down by one dimension](https://docs.spicrawl.com/api-reference/usage/usage-get.md): Reads the durable billing rollup (`usage_daily`) for the key's organization. - [Current-period usage against the plan allowance](https://docs.spicrawl.com/api-reference/usage/usage-summary.md): Usage for the current period, with credit allowance arithmetic in micro-credits (1 credit = 1000000), plus the monthly allowance the API enforces. - [Report drift between billed usage and the request log for one day](https://docs.spicrawl.com/api-reference/usage/usage-reconciliation.md): Compares billed usage for one UTC day with usage re-derived from the Postgres request log. ## API reference: Fleet - [List the deployment's workers, grouped by kind](https://docs.spicrawl.com/api-reference/fleet/fleet-list-workers.md): Reports deployment shape: for each worker kind (`browser`, `egress`, `extract`, `batch`, always in that order and always all four), how many workers are publishing, how many are ready, and each one's load. ## API reference: Health - [Liveness](https://docs.spicrawl.com/api-reference/health/health-live.md): Outside `/v1` and unauthenticated. Answers 200 whenever the process can serve HTTP, and checks no dependency, so it stays 200 during a database or worker outage. - [Readiness](https://docs.spicrawl.com/api-reference/health/health-ready.md): Outside `/v1` and unauthenticated. Checks Postgres, Redis (when configured) and each configured data-plane worker (`fetch`, `extract`, `render`) with a 2-second budget, and reports whether the process is draining. ## Machine-readable resources - [OpenAPI spec](https://docs.spicrawl.com/openapi.yaml): OpenAPI 3.1 description of every endpoint, field, error code and credit price (YAML). - [Agent skill](https://docs.spicrawl.com/skill.md): Compact skill file for coding agents: auth, the scrape parameters, batch, sessions, credits and error handling. - [Full docs text](https://docs.spicrawl.com/llms-full.txt): Every page above as Markdown in one file. - [CLI](https://docs.spicrawl.com/cli/overview.md): The spicrawl command-line client scrapes, renders and extracts web pages, and prints JSON when piped so scripts and AI agents can use it directly. - [MCP server](https://docs.spicrawl.com/agents/mcp.md): Connect any MCP client to the hosted Spicrawl MCP server at https://mcp.spicrawl.com/mcp and use its 25 spicrawl_* tools to scrape, batch, manage sessions, inspect usage and search the docs.