# Get clean markdown for an LLM

> Turn any page or PDF into main-content markdown with response_format=markdown, and cut tokens with include_tags, exclude_tags and main_content_only.

Source: https://docs.spicrawl.com/guides/markdown

Use this when you are feeding a page to a model: a RAG index, an agent's context window, a summariser. Raw HTML spends most of its tokens on markup, navigation and scripts. `response_format=markdown` returns the page's main article as markdown, with navigation, footers and asides stripped by default, and it parses PDFs to text as well. Markdown conversion costs nothing extra: you pay the engine price only.

## Minimal request

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/blog/pricing-update", "response_format": "markdown"}'
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={"url": "https://example.com/blog/pricing-update", "response_format": "markdown"},
    timeout=120,
)
r.raise_for_status()
markdown = r.text
print(r.headers["X-Credits-Charged"], r.headers["X-Target-Status"])
```

```typescript title="TypeScript"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url: "https://example.com/blog/pricing-update", response_format: "markdown" }),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
const markdown = await r.text();
```

```bash title="CLI"
spicrawl scrape https://example.com/blog/pricing-update --format markdown
```

## What comes back

The body is the markdown itself (`Content-Type: text/markdown`). Metadata is in the headers:

```http
HTTP/1.1 200 OK
Content-Type: text/markdown; charset=utf-8
X-Engine: fetch
X-Target-Status: 200
X-Credits-Charged: 1
Cache-State: miss
```

```markdown
# Pricing update for 2026

From 1 March, the Starter plan includes 50,000 requests per month.

| Plan | Price |
|---|---|
| Starter | $29 |
| Scale | $199 |
```

The HTTP status is the platform's. A page that answered 403 still comes back as HTTP 200 at 0 credits, so read `X-Target-Status` before you trust the content.

If you also set `links`, `extract`, `autoparse`, `ai_extract`, `network_capture` or `screenshot`, the response becomes the JSON envelope instead, with the markdown under `content` and an `X-Warning: FORMAT_COERCED` header. Set `response_format=json` yourself when you want the envelope from the start.

## Options that matter

| Field               | Default                           | Effect                                                                                                                                                              |
| ------------------- | --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `response_format`   | `html`                            | `markdown` or `text` for LLM input. `text` drops all markdown syntax.                                                                                               |
| `main_content_only` | unset (isolation on for markdown) | `true` keeps only the main article. `false` keeps the whole document, including nav and footer. Leaving it unset keeps the markdown default, which is isolation on. |
| `include_tags`      | none                              | CSS selectors. Only matching subtrees are kept. Applied first.                                                                                                      |
| `exclude_tags`      | none                              | CSS selectors. Matching elements are removed. Applied after `include_tags`, before `main_content_only`.                                                             |
| `links`             | `false`                           | Adds the page's absolute, de-duplicated `<a href>` targets under `links`. Forces the JSON envelope. No extra credits.                                               |
| `parse_pdf`         | `true`                            | A PDF target is parsed to text. `false` refuses it with `ERR::EXTRACT::FAILED`.                                                                                     |
| `js_render`         | `false`                           | Needed when the article is built by JavaScript. See [JavaScript rendering](https://docs.spicrawl.com/guides/javascript-rendering.md).                                                           |

A project can store defaults for `response_format`, `main_content_only`, `links` and `parse_pdf`. A value in the request always wins.

### Scope the page to cut tokens

`include_tags` and `exclude_tags` run before markdown conversion, so anything you drop never reaches your model:

```json
{
  "url": "https://example.com/docs/install",
  "response_format": "markdown",
  "include_tags": ["article", ".docs-content"],
  "exclude_tags": [".cookie-banner", "aside.related", "pre.changelog"]
}
```

CLI: `spicrawl scrape https://example.com/docs/install --format markdown --include article --include .docs-content --exclude .cookie-banner`.

Token-saving checklist:

* Keep the default main-content isolation unless you need navigation text.
* Use `include_tags` for the one container you need. It is the largest single saving on documentation and news pages.
* Use `text` instead of `markdown` when your model does not need headings, links or tables.
* Leave `links` off unless you are crawling. Links arrive in a separate array; they do not inflate `content`.
* Leave `cache` on (the default). A cache hit costs 0 credits and returns the same markdown.

## PDFs

When the target returns a PDF and you asked for `markdown` or `text`, the PDF is parsed to plain text and returned in the body at the normal engine price. Headings and tables are not reconstructed: a PDF has no HTML structure, so markdown and text return the same plain text.

* `response_format=html` returns the raw PDF bytes instead.
* `extract`, `autoparse` and `links` on a PDF fail with `ERR::EXTRACT::FAILED`, because a PDF has no DOM.
* An encrypted or scanned (image-only) PDF yields no text and fails with `ERR::EXTRACT::FAILED`; the `detail` says which.

## Failure modes

| Code                              | HTTP | What to do                                                                                                                                                                            |
| --------------------------------- | ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `ERR::EXTRACT::FAILED`            | 502  | The target is not HTML (an image or binary), is a PDF with `parse_pdf=false`, or is an unreadable PDF. Request `response_format=html` for the raw bytes, or remove `parse_pdf=false`. |
| `ERR::UPSTREAM::CHALLENGE`        | 502  | The site served a bot challenge. See [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md).                                                                                                                    |
| `ERR::UPSTREAM::TIMEOUT`          | 504  | Retry. If the page is JavaScript-built, add `js_render=true`.                                                                                                                         |
| `ERR::REQUEST::INVALID_PARAMETER` | 400  | A field has the wrong type or an unknown value. The `detail` names it.                                                                                                                |

A markdown body that is nearly empty with `X-Target-Status: 200` usually means the content is rendered by JavaScript. Retry with `js_render=true`.

## Cost

Markdown, text, `include_tags`, `exclude_tags`, `links` and PDF parsing add 0 credits. You pay the engine and proxy tier: 1 credit on the default fetch engine, 3 with `js_render`. Failures, cache hits and non-billable target statuses (anything except 200, 404, 410 or your `allowed_status_codes`) cost 0. `X-Credits-Charged` is the amount billed. See [Credits](https://docs.spicrawl.com/credits.md).

## Related

* [JavaScript rendering](https://docs.spicrawl.com/guides/javascript-rendering.md) when the markdown comes back empty
* [Structured data](https://docs.spicrawl.com/guides/structured-data.md) when you want fields, not prose
* [AI extraction](https://docs.spicrawl.com/guides/ai-extraction.md) (coming soon) to have a model read the markdown for you
* [Caching](https://docs.spicrawl.com/guides/caching.md)
