spicrawlspicrawlDocs

Cursor

Set up Spicrawl in Cursor: add the hosted MCP server to .cursor/mcp.json, install the agent skill, and add a project rule so the agent fetches web pages cheaply and handles errors correctly.

Three steps give Cursor's agent Spicrawl's tools, the knowledge to write Spicrawl code, and rules for using both well. Set your key first, in the environment Cursor starts from:

export SPICRAWL_API_KEY=spicrawl_live_...   # add to your shell profile, then restart Cursor

Or do all three with spicrawl init --client cursor.

1. Add the MCP server

Create or edit .cursor/mcp.json in the project (or ~/.cursor/mcp.json for every project). Cursor replaces ${env:SPICRAWL_API_KEY} with the variable's value, so the file is safe to commit:

.cursor/mcp.json
{
  "mcpServers": {
    "spicrawl": {
      "url": "https://mcp.spicrawl.com/mcp",
      "headers": { "Authorization": "Bearer ${env:SPICRAWL_API_KEY}" }
    }
  }
}

Open Cursor Settings → MCP and check that spicrawl is enabled and lists 25 spicrawl_* tools. See MCP server for each tool.

On macOS, apps started from the Dock do not see variables exported in your shell profile. Start Cursor from a terminal (cursor .) or set the variable where your login session picks it up.

2. Install the skill

Cursor loads skills from .agents/skills/, .cursor/skills/, ~/.agents/skills/ and ~/.cursor/skills/:

spicrawl skill install --client cursor

This writes .cursor/skills/spicrawl/SKILL.md. Without the CLI:

mkdir -p .cursor/skills/spicrawl
curl -fsSL https://app.spicrawl.com/skill.md -o .cursor/skills/spicrawl/SKILL.md

See Agent skill.

3. Add a project rule

Save this as .cursor/rules/spicrawl.mdc. With alwaysApply: false and a description, the agent includes the rule when a task involves web data. Set alwaysApply: true to include it in every chat.

.cursor/rules/spicrawl.mdc
---
description: Fetching, scraping or extracting data from web pages with Spicrawl
alwaysApply: false
---

# Web data (Spicrawl)

Use Spicrawl to read web pages: the spicrawl_* MCP tools in chat, or
POST https://api.spicrawl.com/v1/scrape with `Authorization: Bearer $SPICRAWL_API_KEY` in code.
Docs: https://docs.spicrawl.com/llms.txt (append .md to any page URL for Markdown).

- Ask for markdown: `response_format: "markdown"` (API default is html); on spicrawl_scrape, `format: "markdown"`.
- Check the site's status, not only the HTTP status: `X-Target-Status` header, or `status`
  in the JSON envelope (spicrawl_scrape with `format: "json"`). 200 is the page; 404/410
  mean it does not exist; 403/429/503 mean the site refused, so escalate.
- On an error, switch on `code`. Retry only when `retryable` is true, after
  `retry_after_seconds`. Read `diagnostics.hint` and change what it names first.
  Never retry ERR::REQUEST::*, ERR::AUTH::* or ERR::LIMIT::QUOTA_EXCEEDED.
- Escalate one step at a time and stop at the first that works:
  plain fetch (1 credit) -> `js_render: true` (3) -> add the user's own `proxy`
  if they have one. Empty or skeleton content means render; ERR::UPSTREAM::CHALLENGE
  or a 403 target status means retry once, then the user's own proxy.
- Set `max_cost` on every request to the price of the step you intend.
- Keep the cache on (default). Set `cache: false` only for prices, stock or other live data.
- Trim tokens with `main_content_only` (on by default for markdown), `include_tags`, `exclude_tags`.
- For more than 20 URLs, use a batch job (spicrawl_batch_submit / POST /v1/batch).
  Batch items return raw HTML and apply only render, proxy and block_resources settings.
- Never print or commit SPICRAWL_API_KEY. Log `X-Request-Id` for failures.

Cursor also reads AGENTS.md at the project root. If you already keep agent instructions there, paste the same section into it instead of creating a rule.

A first task to try

Open the agent (Agent mode) and ask:

Use Spicrawl to read https://example.com/pricing as markdown. List each plan with its monthly price.
If the page comes back empty, retry with rendering.

The agent should call spicrawl_scrape with url and format: "markdown", and add render: true only if the first result is empty. For a coding task:

Add a fetchPage(url) function to this project that calls the Spicrawl API for markdown,
following the Spicrawl rule's error-handling steps. Read the key from process.env.SPICRAWL_API_KEY.

Compare the result with the reference implementation in Best practices.

Cloud agents

Cursor's cloud (background) agents run in a remote VM, so your shell's SPICRAWL_API_KEY and ${env:...} references are not available to them. They get MCP servers from two places:

  • Dashboard. Team admins add shared servers under Dashboard → Plugins & MCPs. You add personal servers from the MCP dropdown at cursor.com/agents.
  • Repository. .cursor/mcp.json in the workspace.

Use the HTTP server, which Cursor recommends for cloud agents. Cursor keeps an HTTP server's configuration out of the agent's VM and encrypts its headers at rest, and it cannot be read back after you save it. Enter the values in the dashboard form:

FieldValue
URLhttps://mcp.spicrawl.com/mcp
HeaderAuthorization: Bearer spicrawl_live_... (your key)

Do not commit the key to .cursor/mcp.json. Spicrawl has no OAuth, so the server needs this header. The skill from step 2 and the rule from step 3 are committed to the repository, so cloud agents read them from the checkout.

Troubleshooting

SymptomFix
The server shows an error or 401 in Cursor Settings → MCPSPICRAWL_API_KEY was not set when Cursor started, so the header was empty. Export it and restart Cursor from that shell.
The server is listed but has no toolsToggle it off and on in Cursor Settings → MCP, then open a new chat.
The agent writes js_render on spicrawl_scrape and gets "unknown argument"The tool's argument is render. The API field is js_render.
Usage tools return ERR::AUTH::INSUFFICIENT_SCOPEGrant the read scope to the key in the dashboard.

On this page