Get clean markdown for an LLM
Turn any page or PDF into main-content markdown with response_format=markdown, and cut tokens with include_tags, exclude_tags and main_content_only.
Use this when you are feeding a page to a model: a RAG index, an agent's context window, a summariser. Raw HTML spends most of its tokens on markup, navigation and scripts. response_format=markdown returns the page's main article as markdown, with navigation, footers and asides stripped by default, and it parses PDFs to text as well. Markdown conversion costs nothing extra: you pay the engine price only.
Minimal request
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/blog/pricing-update", "response_format": "markdown"}'What comes back
The body is the markdown itself (Content-Type: text/markdown). Metadata is in the headers:
HTTP/1.1 200 OK
Content-Type: text/markdown; charset=utf-8
X-Engine: fetch
X-Target-Status: 200
X-Credits-Charged: 1
Cache-State: miss# Pricing update for 2026
From 1 March, the Starter plan includes 50,000 requests per month.
| Plan | Price |
|---|---|
| Starter | $29 |
| Scale | $199 |The HTTP status is the platform's. A page that answered 403 still comes back as HTTP 200 at 0 credits, so read X-Target-Status before you trust the content.
If you also set links, extract, autoparse, ai_extract, network_capture or screenshot, the response becomes the JSON envelope instead, with the markdown under content and an X-Warning: FORMAT_COERCED header. Set response_format=json yourself when you want the envelope from the start.
Options that matter
| Field | Default | Effect |
|---|---|---|
response_format | html | markdown or text for LLM input. text drops all markdown syntax. |
main_content_only | unset (isolation on for markdown) | true keeps only the main article. false keeps the whole document, including nav and footer. Leaving it unset keeps the markdown default, which is isolation on. |
include_tags | none | CSS selectors. Only matching subtrees are kept. Applied first. |
exclude_tags | none | CSS selectors. Matching elements are removed. Applied after include_tags, before main_content_only. |
links | false | Adds the page's absolute, de-duplicated <a href> targets under links. Forces the JSON envelope. No extra credits. |
parse_pdf | true | A PDF target is parsed to text. false refuses it with ERR::EXTRACT::FAILED. |
js_render | false | Needed when the article is built by JavaScript. See JavaScript rendering. |
A project can store defaults for response_format, main_content_only, links and parse_pdf. A value in the request always wins.
Scope the page to cut tokens
include_tags and exclude_tags run before markdown conversion, so anything you drop never reaches your model:
{
"url": "https://example.com/docs/install",
"response_format": "markdown",
"include_tags": ["article", ".docs-content"],
"exclude_tags": [".cookie-banner", "aside.related", "pre.changelog"]
}CLI: spicrawl scrape https://example.com/docs/install --format markdown --include article --include .docs-content --exclude .cookie-banner.
Token-saving checklist:
- Keep the default main-content isolation unless you need navigation text.
- Use
include_tagsfor the one container you need. It is the largest single saving on documentation and news pages. - Use
textinstead ofmarkdownwhen your model does not need headings, links or tables. - Leave
linksoff unless you are crawling. Links arrive in a separate array; they do not inflatecontent. - Leave
cacheon (the default). A cache hit costs 0 credits and returns the same markdown.
PDFs
When the target returns a PDF and you asked for markdown or text, the PDF is parsed to plain text and returned in the body at the normal engine price. Headings and tables are not reconstructed: a PDF has no HTML structure, so markdown and text return the same plain text.
response_format=htmlreturns the raw PDF bytes instead.extract,autoparseandlinkson a PDF fail withERR::EXTRACT::FAILED, because a PDF has no DOM.- An encrypted or scanned (image-only) PDF yields no text and fails with
ERR::EXTRACT::FAILED; thedetailsays which.
Failure modes
| Code | HTTP | What to do |
|---|---|---|
ERR::EXTRACT::FAILED | 502 | The target is not HTML (an image or binary), is a PDF with parse_pdf=false, or is an unreadable PDF. Request response_format=html for the raw bytes, or remove parse_pdf=false. |
ERR::UPSTREAM::CHALLENGE | 502 | The site served a bot challenge. See Anti-bot. |
ERR::UPSTREAM::TIMEOUT | 504 | Retry. If the page is JavaScript-built, add js_render=true. |
ERR::REQUEST::INVALID_PARAMETER | 400 | A field has the wrong type or an unknown value. The detail names it. |
A markdown body that is nearly empty with X-Target-Status: 200 usually means the content is rendered by JavaScript. Retry with js_render=true.
Cost
Markdown, text, include_tags, exclude_tags, links and PDF parsing add 0 credits. You pay the engine and proxy tier: 1 credit on the default fetch engine, 3 with js_render. Failures, cache hits and non-billable target statuses (anything except 200, 404, 410 or your allowed_status_codes) cost 0. X-Credits-Charged is the amount billed. See Credits.
Related
- JavaScript rendering when the markdown comes back empty
- Structured data when you want fields, not prose
- AI extraction Coming soon to have a model read the markdown for you
- Caching
Agent frameworks
Give an agent built with the OpenAI Agents SDK, Vercel AI SDK, LangChain, CrewAI, Google ADK, Mastra or smolagents the 25 spicrawl_* tools by connecting to the hosted Spicrawl MCP server.
JavaScript rendering
Run a page in a browser with js_render, pick an engine, wait for content with wait and wait_for, and let mode=auto escalate only when it has to.