# Spicrawl

> Spicrawl is a web data API for agents: fetch, render and extract any URL as markdown, HTML or structured JSON with one request.

Source: https://docs.spicrawl.com

Spicrawl turns any URL into data your code or AI agent can use: clean markdown, HTML, text, a PDF, or JSON extracted with selectors (extraction with a model (coming soon)). You send one request shape to `https://api.spicrawl.com/v1/scrape`, say what you need with flags (`js_render`, `proxy`, `extract`), and Spicrawl picks the engine and retries for you. You pay credits only for successful requests; failures and cache hits cost 0.

```bash
curl -sS "https://api.spicrawl.com/v1/scrape" \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/products/42", "response_format": "markdown"}'
```

The response body is the page as markdown. The headers say what the site answered (`X-Target-Status`), which engine ran (`X-Engine`) and what it cost (`X-Credits-Charged`).

- [Quickstart](https://docs.spicrawl.com/quickstart.md): Get a key and scrape your first page in curl, Python, TypeScript or the CLI.

- [Build with agents](https://docs.spicrawl.com/agents/overview.md): Connect Claude Code, Cursor, Codex, GitHub Copilot, OpenCode, Devin and other agents, or your own, through MCP, the skill or the CLI.

- [CLI](https://docs.spicrawl.com/cli/overview.md): The `spicrawl` binary: scrape, batch, sessions and logs from a shell or script.

- [Guides](https://docs.spicrawl.com/guides/markdown.md): Render JavaScript, get past anti-bot checks, extract structured data, run batches.

- [API reference](https://docs.spicrawl.com/api-reference/introduction.md): Every endpoint, parameter and response, generated from the OpenAPI spec.

- [Errors](https://docs.spicrawl.com/errors.md): Every error code, whether to retry, and which parameter to change.

## For AI agents

Agents can use Spicrawl, and read these docs, without a human in the loop:

| Entry point               | Where                                                                      | Use it to                                                                                                            |
| ------------------------- | -------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| `llms.txt`                | [`https://docs.spicrawl.com/llms.txt`](https://docs.spicrawl.com/llms.txt) | Load an index of these docs. Append `.md` to any page URL for its Markdown. See [llms.txt](https://docs.spicrawl.com/agents/llms-txt.md).        |
| MCP server                | `https://mcp.spicrawl.com/mcp` (streamable HTTP)                           | Give an MCP client `spicrawl_*` tools. See [MCP](https://docs.spicrawl.com/agents/mcp.md).                                                       |
| Agent skill               | [`https://app.spicrawl.com/skill.md`](https://app.spicrawl.com/skill.md)   | Teach a coding agent the HTTP API so it writes correct calls. See [Agent skill](https://docs.spicrawl.com/agents/skill.md).                      |
| CLI                       | `spicrawl init`                                                            | Detect your agent tooling and install the MCP server and skill in one step. See [CLI agent setup](https://docs.spicrawl.com/cli/agent-setup.md). |
| JavaScript/TypeScript SDK | `npm install @spicrawl/sdk`                                                | Call the API from Node.js 18+ with typed methods instead of raw HTTP.                                                |

Rules an agent should hold on to:

* **Two statuses.** The HTTP status is Spicrawl's. The site's status is `X-Target-Status`. A site that refused you is still HTTP `200`, at 0 credits.
* **Switch on `code`, retry on `retryable`.** Errors are `application/problem+json`. Honour `Retry-After`, and read `diagnostics.hint`: it names the parameter to change. See [Errors](https://docs.spicrawl.com/errors.md).
* **Failures cost 0.** `X-Credits-Charged` is what was billed.
* **Not idempotent.** Retrying a `POST /v1/scrape` that succeeded performs and bills a second scrape.
* **Unknown fields are rejected.** A misspelled parameter is a `400`, never ignored.
