spicrawlspicrawlDocs

Get clean markdown for an LLM

Turn any page or PDF into main-content markdown with response_format=markdown, and cut tokens with include_tags, exclude_tags and main_content_only.

Use this when you are feeding a page to a model: a RAG index, an agent's context window, a summariser. Raw HTML spends most of its tokens on markup, navigation and scripts. response_format=markdown returns the page's main article as markdown, with navigation, footers and asides stripped by default, and it parses PDFs to text as well. Markdown conversion costs nothing extra: you pay the engine price only.

Minimal request

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/blog/pricing-update", "response_format": "markdown"}'

What comes back

The body is the markdown itself (Content-Type: text/markdown). Metadata is in the headers:

HTTP/1.1 200 OK
Content-Type: text/markdown; charset=utf-8
X-Engine: fetch
X-Target-Status: 200
X-Credits-Charged: 1
Cache-State: miss
# Pricing update for 2026

From 1 March, the Starter plan includes 50,000 requests per month.

| Plan | Price |
|---|---|
| Starter | $29 |
| Scale | $199 |

The HTTP status is the platform's. A page that answered 403 still comes back as HTTP 200 at 0 credits, so read X-Target-Status before you trust the content.

If you also set links, extract, autoparse, ai_extract, network_capture or screenshot, the response becomes the JSON envelope instead, with the markdown under content and an X-Warning: FORMAT_COERCED header. Set response_format=json yourself when you want the envelope from the start.

Options that matter

FieldDefaultEffect
response_formathtmlmarkdown or text for LLM input. text drops all markdown syntax.
main_content_onlyunset (isolation on for markdown)true keeps only the main article. false keeps the whole document, including nav and footer. Leaving it unset keeps the markdown default, which is isolation on.
include_tagsnoneCSS selectors. Only matching subtrees are kept. Applied first.
exclude_tagsnoneCSS selectors. Matching elements are removed. Applied after include_tags, before main_content_only.
linksfalseAdds the page's absolute, de-duplicated <a href> targets under links. Forces the JSON envelope. No extra credits.
parse_pdftrueA PDF target is parsed to text. false refuses it with ERR::EXTRACT::FAILED.
js_renderfalseNeeded when the article is built by JavaScript. See JavaScript rendering.

A project can store defaults for response_format, main_content_only, links and parse_pdf. A value in the request always wins.

Scope the page to cut tokens

include_tags and exclude_tags run before markdown conversion, so anything you drop never reaches your model:

{
  "url": "https://example.com/docs/install",
  "response_format": "markdown",
  "include_tags": ["article", ".docs-content"],
  "exclude_tags": [".cookie-banner", "aside.related", "pre.changelog"]
}

CLI: spicrawl scrape https://example.com/docs/install --format markdown --include article --include .docs-content --exclude .cookie-banner.

Token-saving checklist:

  • Keep the default main-content isolation unless you need navigation text.
  • Use include_tags for the one container you need. It is the largest single saving on documentation and news pages.
  • Use text instead of markdown when your model does not need headings, links or tables.
  • Leave links off unless you are crawling. Links arrive in a separate array; they do not inflate content.
  • Leave cache on (the default). A cache hit costs 0 credits and returns the same markdown.

PDFs

When the target returns a PDF and you asked for markdown or text, the PDF is parsed to plain text and returned in the body at the normal engine price. Headings and tables are not reconstructed: a PDF has no HTML structure, so markdown and text return the same plain text.

  • response_format=html returns the raw PDF bytes instead.
  • extract, autoparse and links on a PDF fail with ERR::EXTRACT::FAILED, because a PDF has no DOM.
  • An encrypted or scanned (image-only) PDF yields no text and fails with ERR::EXTRACT::FAILED; the detail says which.

Failure modes

CodeHTTPWhat to do
ERR::EXTRACT::FAILED502The target is not HTML (an image or binary), is a PDF with parse_pdf=false, or is an unreadable PDF. Request response_format=html for the raw bytes, or remove parse_pdf=false.
ERR::UPSTREAM::CHALLENGE502The site served a bot challenge. See Anti-bot.
ERR::UPSTREAM::TIMEOUT504Retry. If the page is JavaScript-built, add js_render=true.
ERR::REQUEST::INVALID_PARAMETER400A field has the wrong type or an unknown value. The detail names it.

A markdown body that is nearly empty with X-Target-Status: 200 usually means the content is rendered by JavaScript. Retry with js_render=true.

Cost

Markdown, text, include_tags, exclude_tags, links and PDF parsing add 0 credits. You pay the engine and proxy tier: 1 credit on the default fetch engine, 3 with js_render. Failures, cache hits and non-billable target statuses (anything except 200, 404, 410 or your allowed_status_codes) cost 0. X-Credits-Charged is the amount billed. See Credits.

On this page