spicrawlspicrawlDocs

JetBrains AI Assistant and Junie

Set up Spicrawl in JetBrains IDEs: add the hosted MCP server to AI Assistant and Junie, install the agent skill, and add project instructions so the agent fetches web pages cheaply and handles errors correctly.

JetBrains IDEs have two agents that connect to MCP servers: AI Assistant (its own settings page and its own rules) and Junie (its own mcp.json, skills and guidelines). Both talk to Streamable HTTP servers. The Spicrawl CLI has no JetBrains target (--client accepts claude, cursor, vscode, codex and all), so you do each step by hand. Have your key ready:

export SPICRAWL_API_KEY=spicrawl_live_...   # for your own code and the skill's examples

1. Add the MCP server

Both agents use the same server entry: the hosted endpoint plus your key as a bearer header. Spicrawl has no OAuth, so the header is required.

Open Settings | Tools | AI Assistant | Model Context Protocol (MCP) and click Add. JetBrains documents only a url for HTTP servers, not request headers, so connect through the mcp-remote bridge (needs Node.js) on the STDIO transport. Paste this JSON:

{
  "mcpServers": {
    "spicrawl": {
      "command": "npx",
      "args": [
        "-y", "mcp-remote", "https://mcp.spicrawl.com/mcp",
        "--header", "Authorization:${SPICRAWL_AUTH}"
      ],
      "env": { "SPICRAWL_AUTH": "Bearer spicrawl_live_..." }
    }
  }
}

The header value goes through the SPICRAWL_AUTH variable because some clients split arguments on spaces. Click OK, then Apply. The server appears in the list with its connection status and tools; it should show 25 spicrawl_* tools. Set the level to Global to use it in every project, or Project to scope it. The key is stored in the IDE's MCP settings, not in your repository.

If your AI Assistant version accepts headers on the HTTP transport, the direct form is {"mcpServers": {"spicrawl": {"url": "https://mcp.spicrawl.com/mcp", "headers": {"Authorization": "Bearer spicrawl_live_..."}}}}. If the server then fails with 401, use the bridge above.

See MCP server for each tool.

2. Install the skill

Junie loads skills from .junie/skills/ in the project and ~/.junie/skills/ for every project, and also from .agents/skills/. Each skill is a folder with a SKILL.md:

mkdir -p .agents/skills/spicrawl
curl -fsSL https://app.spicrawl.com/skill.md -o .agents/skills/spicrawl/SKILL.md

spicrawl skill install --client codex writes the same file. Junie selects the skill when a task matches its description; run /skills to see what it discovered, or /spicrawl to apply it directly. See Agent skill.

AI Assistant does not read this folder; it relies on the rules in step 3.

3. Add instructions

Junie reads AGENTS.md at the project root (with .junie/playbook.md and .junie/rules/*.md if you have them), or .junie/AGENTS.md, and the older .junie/guidelines.md. Use AGENTS.md, which other agents read too. ~/.junie/AGENTS.md holds global guidelines; project guidelines win on conflict.

AGENTS.md
## Web data (Spicrawl)

Use Spicrawl to read web pages: the spicrawl_* MCP tools in chat, or
POST https://api.spicrawl.com/v1/scrape with `Authorization: Bearer $SPICRAWL_API_KEY` in code.
Docs: https://docs.spicrawl.com/llms.txt (append .md to any page URL for Markdown).

- Ask for markdown: `response_format: "markdown"` (API default is html); on spicrawl_scrape, `format: "markdown"`.
- Check the site's status, not only the HTTP status: `X-Target-Status` header, or `status`
  in the JSON envelope (spicrawl_scrape with `format: "json"`). 200 is the page; 404/410
  mean it does not exist; 403/429/503 mean the site refused, so escalate.
- On an error, switch on `code`. Retry only when `retryable` is true, after
  `retry_after_seconds`. Read `diagnostics.hint` and change what it names first.
  Never retry ERR::REQUEST::*, ERR::AUTH::* or ERR::LIMIT::QUOTA_EXCEEDED.
- Escalate one step at a time and stop at the first that works:
  plain fetch (1 credit) -> `js_render: true` (3) -> add the user's own `proxy`
  if they have one. Empty or skeleton content means render; ERR::UPSTREAM::CHALLENGE
  or a 403 target status means retry once, then the user's own proxy.
- Set `max_cost` on every request to the price of the step you intend.
- Keep the cache on (default). Set `cache: false` only for prices, stock or other live data.
- Trim tokens with `main_content_only` (on by default for markdown), `include_tags`, `exclude_tags`.
- For more than 20 URLs, use a batch job (spicrawl_batch_submit / POST /v1/batch).
  Batch items return raw HTML and apply only render, proxy and block_resources settings.
- Never print or commit SPICRAWL_API_KEY. Log `X-Request-Id` for failures.

A first task to try

Open the agent chat and ask:

Use Spicrawl to read https://example.com/pricing as markdown. List each plan with its monthly price.
If the page comes back empty, retry with rendering.

The agent should call spicrawl_scrape with url and format: "markdown", and add render: true only if the first result is empty. For a coding task:

Add a fetchPage(url) function to this project that calls the Spicrawl API for markdown,
following the Web data instructions. Read the key from process.env.SPICRAWL_API_KEY.

Compare the result with the reference implementation in Best practices.

Troubleshooting

SymptomFix
The server shows as failed, or the endpoint answers 401 Unauthorized: send your Spicrawl API keyThe Authorization header is missing, wrong, or has no Bearer prefix. Fix the JSON and apply again.
The server connects but shows no toolsOpen a new chat, and check the server's tools on the MCP settings page (or /mcp in Junie CLI).
A rule in .aiassistant/rules/ is ignoredIts type is Off or Manually. Set it to Always or By model decision.
The agent writes js_render on spicrawl_scrape and gets "unknown argument"The tool's argument is render. The API field is js_render.
Usage tools return ERR::AUTH::INSUFFICIENT_SCOPEGrant the read scope to the key in the dashboard.

On this page