Codex
Set up Spicrawl in the Codex CLI: add the hosted MCP server to config.toml, install the agent skill, and add AGENTS.md rules so the agent fetches web pages cheaply and handles errors correctly.
Three steps give Codex Spicrawl's tools, the knowledge to write Spicrawl code, and rules for using both well. Set your key first:
export SPICRAWL_API_KEY=spicrawl_live_... # add to your shell profileOr do all three with spicrawl init --client codex.
1. Add the MCP server
codex mcp add spicrawl --url https://mcp.spicrawl.com/mcp --bearer-token-env-var SPICRAWL_API_KEYThis writes the server to ~/.codex/config.toml. Codex reads the key from SPICRAWL_API_KEY each time it connects and sends it as Authorization: Bearer, so the key never lands in the file. The equivalent config:
[mcp_servers.spicrawl]
url = "https://mcp.spicrawl.com/mcp"
bearer_token_env_var = "SPICRAWL_API_KEY"To scope the server to one project, put the same table in .codex/config.toml at the project root. Codex reads project config only in projects you have marked as trusted.
Check it with codex mcp list, or /mcp inside a session. The server shows 25 spicrawl_* tools. See MCP server for each tool.
2. Install the skill
Codex loads skills from .agents/skills/ in the working directory and each parent up to the repository root, and from $HOME/.agents/skills/ for every repository:
mkdir -p .agents/skills/spicrawl
curl -fsSL https://app.spicrawl.com/skill.md -o .agents/skills/spicrawl/SKILL.mdCodex picks the skill when a task matches its description (fetching, rendering or extracting web data). See Agent skill.
3. Add rules to AGENTS.md
Paste this section into AGENTS.md at the project root. Codex reads it before it starts work.
## Web data (Spicrawl)
Use Spicrawl to read web pages: the spicrawl_* MCP tools, the `spicrawl` CLI, or
POST https://api.spicrawl.com/v1/scrape with `Authorization: Bearer $SPICRAWL_API_KEY` in code.
Docs: https://docs.spicrawl.com/llms.txt (append .md to any page URL for Markdown).
- Ask for markdown: `response_format: "markdown"` (API default is html); on spicrawl_scrape, `format: "markdown"`.
- Check the site's status, not only the HTTP status: `X-Target-Status` header, or `status`
in the JSON envelope (spicrawl_scrape with `format: "json"`). 200 is the page; 404/410
mean it does not exist; 403/429/503 mean the site refused, so escalate.
- On an error, switch on `code`. Retry only when `retryable` is true, after
`retry_after_seconds`. Read `diagnostics.hint` and change what it names first.
Never retry ERR::REQUEST::*, ERR::AUTH::* or ERR::LIMIT::QUOTA_EXCEEDED.
- Escalate one step at a time and stop at the first that works:
plain fetch (1 credit) -> `js_render: true` (3) -> add the user's own `proxy`
if they have one. Empty or skeleton content means render; ERR::UPSTREAM::CHALLENGE
or a 403 target status means retry once, then the user's own proxy.
- Set `max_cost` on every request to the price of the step you intend.
- Keep the cache on (default). Set `cache: false` only for prices, stock or other live data.
- Trim tokens with `main_content_only` (on by default for markdown), `include_tags`, `exclude_tags`.
- For more than 20 URLs, use a batch job (spicrawl_batch_submit / POST /v1/batch).
Batch items return raw HTML and apply only render, proxy and block_resources settings.
- With the CLI, read stdout as JSON and branch on the exit code: 3 auth, 4 bad request,
5 limit, 6 site refused or failed, 7 proxy, 8 engine or extraction, 10 network.
- Never print or commit SPICRAWL_API_KEY. Log `X-Request-Id` for failures.Use the CLI instead of MCP
Codex works in a shell, so the spicrawl CLI is an alternative to the MCP server with no config at all. It prints JSON when stdout is not a terminal and exits with a stable code:
spicrawl scrape https://example.com/pricing --format markdown --max-cost 3 | jq -r .contentWhen the CLI gets the page as a document, its JSON output adds status (the site's status), credits_charged and request_id around content. See CLI overview.
A first task to try
Start codex in the project and ask:
Use Spicrawl to read https://example.com/pricing as markdown. List each plan with its monthly price.
If the page comes back empty, retry with rendering.For a coding task:
Add a fetch_page(url) function to this project that calls the Spicrawl API for markdown,
following the Web data rules in AGENTS.md. Read the key from SPICRAWL_API_KEY.Compare the result with the reference implementation in Best practices.
Codex cloud
Codex cloud tasks run in a hosted container, not on your machine. The ~/.codex/config.toml from step 1 belongs to the CLI, IDE extension and desktop app, and OpenAI's docs warn that local MCP servers and transports may not be available in the cloud. Do not count on the MCP server there. Use the CLI or plain HTTP instead, both of which need only the key.
Set the key as an environment variable
In the cloud environment's settings, add SPICRAWL_API_KEY as an environment variable, not a secret. Environment variables last for the whole task, setup script and agent phase. Secrets are only available to setup scripts and are removed before the agent phase starts, so the agent could not read the key from one.
Install the CLI in the setup script (optional)
The setup script runs with internet access. Add:
npm install -g @spicrawl/cliSkip this step to use plain HTTP with curl instead.
Allow api.spicrawl.com
Agent internet access is off by default. In the environment's settings, set it to On, choose the None allowlist preset and add the domain api.spicrawl.com. If you also restrict HTTP methods to GET, HEAD and OPTIONS, POST /v1/scrape is blocked, so allow POST too.
Add the rules to AGENTS.md
Commit the section from step 3 to the repository. Cloud tasks read AGENTS.md from the checkout.
Then the agent can run either of these:
spicrawl scrape https://example.com/pricing --format markdown --max-cost 3Enabling internet access lets the agent reach any domain you allow. OpenAI recommends allowing only the domains and methods you need, because web content can carry prompt injection.
Troubleshooting
| Symptom | Fix |
|---|---|
| The server fails to start or returns 401 | SPICRAWL_API_KEY is not set in the shell that started codex. Export it and start a new session. |
In a cloud task the agent says SPICRAWL_API_KEY is empty, or requests to api.spicrawl.com fail | The key was saved as a secret (removed before the agent phase) or agent internet access does not allow api.spicrawl.com. See Codex cloud. |
Tools from .codex/config.toml do not load | The project is not trusted. Trust it, or move the table to ~/.codex/config.toml. |
The agent writes js_render on spicrawl_scrape and gets "unknown argument" | The tool's argument is render. The API field is js_render. |
Usage tools return ERR::AUTH::INSUFFICIENT_SCOPE | Grant the read scope to the key in the dashboard. |
Cursor
Set up Spicrawl in Cursor: add the hosted MCP server to .cursor/mcp.json, install the agent skill, and add a project rule so the agent fetches web pages cheaply and handles errors correctly.
GitHub Copilot and VS Code
Set up Spicrawl in VS Code agent mode, the Copilot CLI and the Copilot cloud agent: add the hosted MCP server, install the agent skill, and add repository instructions.