Gemini CLI
Set up Spicrawl in Google's Gemini CLI: add the hosted MCP server to settings.json, install the agent skill, and add GEMINI.md rules so the agent fetches web pages cheaply and handles errors correctly.
Three steps give Gemini CLI Spicrawl's tools, the knowledge to write Spicrawl code, and rules for using both well. Set your key first:
export SPICRAWL_API_KEY=spicrawl_live_... # add to your shell profile
export SPICRAWL_MCP_BEARER="$SPICRAWL_API_KEY"The second variable exists because of how Gemini CLI expands variables in MCP headers (see the callout in step 1). The spicrawl init command does not support Gemini CLI, so set it up with the steps below.
1. Add the MCP server
Gemini CLI reads MCP servers from ~/.gemini/settings.json (every project) or .gemini/settings.json (one project). It connects to a Streamable HTTP server through the httpUrl key and expands $VAR and ${VAR} in headers values when it connects, so the file holds no key:
{
"mcpServers": {
"spicrawl": {
"httpUrl": "https://mcp.spicrawl.com/mcp",
"headers": { "Authorization": "Bearer $SPICRAWL_MCP_BEARER" }
}
}
}The same entry from the command line. Single quotes keep your shell from expanding the variable, so the file stores the reference. --scope user writes ~/.gemini/settings.json; the default, project, writes .gemini/settings.json:
gemini mcp add --scope user --transport http \
--header 'Authorization: Bearer $SPICRAWL_MCP_BEARER' \
spicrawl https://mcp.spicrawl.com/mcpStart gemini and run /mcp. spicrawl should be connected and list 25 spicrawl_* tools. gemini mcp list shows the same status outside a session. See MCP server for each tool.
Do not name the variable SPICRAWL_API_KEY in the header
Before expanding headers, Gemini CLI removes from the environment every variable whose name contains KEY, TOKEN, SECRET, PASSWORD, AUTH, CREDENTIAL, CERT or PRIVATE. $SPICRAWL_API_KEY would expand to an empty string and the server would answer 401. Export the key under a name without those words, such as SPICRAWL_MCP_BEARER.
Spicrawl's server does not use OAuth. Leave the oauth block out of the entry. If the server shows an authentication error, check that SPICRAWL_MCP_BEARER is set in the shell that started gemini, and do not run /mcp auth.
Gemini CLI also supports extensions, which bundle MCP servers, skills and commands. No Spicrawl extension is published, so use the settings entry above.
Qwen Code
Qwen Code, a fork of Gemini CLI, uses the same mcpServers format with httpUrl and headers in ~/.qwen/settings.json or .qwen/settings.json, and qwen mcp add --transport http --header "..." spicrawl https://mcp.spicrawl.com/mcp. Put the literal Authorization: Bearer spicrawl_live_... value in the user-level file, not in a committed project file.
2. Install the skill
Gemini CLI loads skills from .gemini/skills/ and .agents/skills/ in the project, and from ~/.gemini/skills/ and ~/.agents/skills/ for every project. Where a skill name exists in both, .agents/skills/ wins:
spicrawl skill install --client codexThis writes .agents/skills/spicrawl/SKILL.md, a location Gemini CLI reads (the codex value is the CLI's only option that writes there). Without the CLI:
mkdir -p .agents/skills/spicrawl
curl -fsSL https://app.spicrawl.com/skill.md -o .agents/skills/spicrawl/SKILL.mdRun /skills reload in a running session, then /skills list to confirm spicrawl appears. Gemini CLI shows the skill's name and description to the model at the start of a session and asks you to approve when it activates the skill. gemini skills install installs from a Git repository or local directory, not from a URL, so use the curl command above. See Agent skill.
3. Add rules to GEMINI.md
Gemini CLI concatenates GEMINI.md files (~/.gemini/GEMINI.md for all projects, and the ones in the project and its parent directories) and sends them with every prompt. Paste this section into GEMINI.md at the project root, then run /memory reload:
# Web data (Spicrawl)
Use Spicrawl to read web pages: the spicrawl_* MCP tools in chat, or
POST https://api.spicrawl.com/v1/scrape with `Authorization: Bearer $SPICRAWL_API_KEY` in code.
Docs: https://docs.spicrawl.com/llms.txt (append .md to any page URL for Markdown).
- Ask for markdown: `response_format: "markdown"` (API default is html); on spicrawl_scrape, `format: "markdown"`.
- Check the site's status, not only the HTTP status: `X-Target-Status` header, or `status`
in the JSON envelope (spicrawl_scrape with `format: "json"`). 200 is the page; 404/410
mean it does not exist; 403/429/503 mean the site refused, so escalate.
- On an error, switch on `code`. Retry only when `retryable` is true, after
`retry_after_seconds`. Read `diagnostics.hint` and change what it names first.
Never retry ERR::REQUEST::*, ERR::AUTH::* or ERR::LIMIT::QUOTA_EXCEEDED.
- Escalate one step at a time and stop at the first that works:
plain fetch (1 credit) -> `js_render: true` (3) -> add the user's own `proxy`
if they have one. Empty or skeleton content means render; ERR::UPSTREAM::CHALLENGE
or a 403 target status means retry once, then the user's own proxy.
- Set `max_cost` on every request to the price of the step you intend.
- Keep the cache on (default). Set `cache: false` only for prices, stock or other live data.
- Trim tokens with `main_content_only` (on by default for markdown), `include_tags`, `exclude_tags`.
- For more than 20 URLs, use a batch job (spicrawl_batch_submit / POST /v1/batch).
Batch items return raw HTML and apply only render, proxy and block_resources settings.
- Never print or commit SPICRAWL_API_KEY. Log `X-Request-Id` for failures.To have Gemini CLI read an existing AGENTS.md instead, list both names in settings.json:
{
"context": { "fileName": ["AGENTS.md", "GEMINI.md"] }
}A first task to try
Start gemini in the project and ask:
Use Spicrawl to read https://example.com/pricing as markdown. List each plan with its monthly price.
If the page comes back empty, retry with rendering.The agent should call spicrawl_scrape with url and format: "markdown", and add render: true only if the first result is empty. Gemini CLI asks you to confirm each tool call unless the server has "trust": true. For a coding task:
Add a fetchPage(url) function to this project that calls the Spicrawl API for markdown,
following the Web data rules in GEMINI.md. Read the key from process.env.SPICRAWL_API_KEY.Compare the result with the reference implementation in Best practices.
Troubleshooting
| Symptom | Fix |
|---|---|
/mcp shows spicrawl disconnected or a 401 | The header expanded to an empty value. Use a variable name without KEY, TOKEN, AUTH or the other blocked words (SPICRAWL_MCP_BEARER), and start gemini from a shell where it is set. |
The server shows no tools, or spicrawl is missing from /mcp | Check that the file is valid JSON, that the URL is under httpUrl (not url, which is SSE), and that a project-level .gemini/settings.json is not overriding your user entry. |
The agent writes js_render on spicrawl_scrape and gets "unknown argument" | The tool's argument is render. The API field is js_render. |
Usage tools return ERR::AUTH::INSUFFICIENT_SCOPE | Grant the read scope to the key in the dashboard. |
OpenCode
Set up Spicrawl in OpenCode: add the hosted MCP server to opencode.json with OAuth turned off, install the agent skill, and add AGENTS.md rules so the agent fetches web pages cheaply and handles errors correctly.
Windsurf
Set up Spicrawl in Windsurf, now named Devin Desktop: add the hosted MCP server, install the agent skill, and add a rule so the agent fetches web pages cheaply and handles errors correctly.