spicrawlspicrawlDocs

Goose

Set up Spicrawl in Goose: add the hosted MCP server as a Streamable HTTP extension, install the agent skill, and add AGENTS.md rules so the agent fetches web pages cheaply and handles errors correctly.

Three steps give Goose Spicrawl's tools, the knowledge to write Spicrawl code, and rules for using both well. Goose calls MCP servers extensions. Spicrawl connects as a remote extension of type Streamable HTTP.

1. Add the MCP server

Run goose configure and answer the prompts:

Choose the extension type

Select Add Extension, then Remote Extension (Streamable HTTP).

Enter the endpoint

Give the extension a name such as spicrawl, and https://mcp.spicrawl.com/mcp as the endpoint URI. Set the timeout in seconds (the docs' examples use 300).

Add the auth header

Answer yes to adding custom headers. Use the header name Authorization and the value Bearer spicrawl_live_... with your real key.

The extension is saved to ~/.config/goose/config.yaml. You can also write it by hand:

~/.config/goose/config.yaml
extensions:
  spicrawl:
    name: Spicrawl
    type: streamable_http
    uri: https://mcp.spicrawl.com/mcp
    headers:
      Authorization: "Bearer spicrawl_live_..."
    enabled: true
    timeout: 300

Goose stores the header value as written, so the key sits in config.yaml in plain text. Do not commit or share that file. The Goose docs do not describe $VAR substitution for headers, so this page does not rely on it.

Spicrawl authenticates with the static header only and does not support OAuth. In Goose Desktop, Extensions in the sidebar then Add custom extension opens a similar form. See MCP server for the 25 spicrawl_* tools.

2. Install the skill

Goose supports Agent Skills. It loads them from .agents/skills/ in the project and ~/.agents/skills/ for every project (it still reads .goose/skills/ and .claude/skills/ for backward compatibility):

spicrawl skill install --client codex

This writes .agents/skills/spicrawl/SKILL.md. The --client value names the CLI's own target, not the agent that reads it. Without the CLI:

mkdir -p .agents/skills/spicrawl
curl -fsSL https://app.spicrawl.com/skill.md -o .agents/skills/spicrawl/SKILL.md

See Agent skill.

3. Add rules to AGENTS.md

By default Goose reads AGENTS.md, then .goosehints, from the project root. Paste this section into AGENTS.md. To apply it to every session instead, put it in ~/.config/goose/.goosehints.

AGENTS.md
## Web data (Spicrawl)

Use Spicrawl to read web pages: the spicrawl_* MCP tools in chat, or
POST https://api.spicrawl.com/v1/scrape with `Authorization: Bearer $SPICRAWL_API_KEY` in code.
Docs: https://docs.spicrawl.com/llms.txt (append .md to any page URL for Markdown).

- Ask for markdown: `response_format: "markdown"` (API default is html); on spicrawl_scrape, `format: "markdown"`.
- Check the site's status, not only the HTTP status: `X-Target-Status` header, or `status`
  in the JSON envelope (spicrawl_scrape with `format: "json"`). 200 is the page; 404/410
  mean it does not exist; 403/429/503 mean the site refused, so escalate.
- On an error, switch on `code`. Retry only when `retryable` is true, after
  `retry_after_seconds`. Read `diagnostics.hint` and change what it names first.
  Never retry ERR::REQUEST::*, ERR::AUTH::* or ERR::LIMIT::QUOTA_EXCEEDED.
- Escalate one step at a time and stop at the first that works:
  plain fetch (1 credit) -> `js_render: true` (3) -> add the user's own `proxy`
  if they have one. Empty or skeleton content means render; ERR::UPSTREAM::CHALLENGE
  or a 403 target status means retry once, then the user's own proxy.
- Set `max_cost` on every request to the price of the step you intend.
- Keep the cache on (default). Set `cache: false` only for prices, stock or other live data.
- Trim tokens with `main_content_only` (on by default for markdown), `include_tags`, `exclude_tags`.
- For more than 20 URLs, use a batch job (spicrawl_batch_submit / POST /v1/batch).
  Batch items return raw HTML and apply only render, proxy and block_resources settings.
- Never print or commit SPICRAWL_API_KEY. Log `X-Request-Id` for failures.

A first task to try

Start a Goose session in the project (goose session) and ask:

Use Spicrawl to read https://example.com/pricing as markdown. List each plan with its monthly price.
If the page comes back empty, retry with rendering.

The agent should call spicrawl_scrape with url and format: "markdown", and add render: true only if the first result is empty. For a coding task:

Add a fetchPage(url) function to this project that calls the Spicrawl API for markdown,
following the Web data rules in AGENTS.md. Read the key from process.env.SPICRAWL_API_KEY.

Compare the result with the reference implementation in Best practices.

Troubleshooting

SymptomFix
The extension fails to load or returns 401The Authorization header is missing, misspelled or holds a placeholder key. Check the value is Bearer followed by the key, with one space. Confirm the key with curl -i https://mcp.spicrawl.com/mcp -H "Authorization: Bearer $SPICRAWL_API_KEY".
The extension times out on a slow pageRaise timeout (seconds) on the extension.
The agent writes js_render on spicrawl_scrape and gets "unknown argument"The tool's argument is render. The API field is js_render.
Usage tools return ERR::AUTH::INSUFFICIENT_SCOPEGrant the read scope to the key in the dashboard.

On this page