# Codex

> Set up Spicrawl in the Codex CLI: add the hosted MCP server to config.toml, install the agent skill, and add AGENTS.md rules so the agent fetches web pages cheaply and handles errors correctly.

Source: https://docs.spicrawl.com/agents/codex

Three steps give Codex Spicrawl's tools, the knowledge to write Spicrawl code, and rules for using both well. Set your key first:

```bash
export SPICRAWL_API_KEY=spicrawl_live_...   # add to your shell profile
```

Or do all three with `spicrawl init --client codex`.

## 1. Add the MCP server

```bash
codex mcp add spicrawl --url https://mcp.spicrawl.com/mcp --bearer-token-env-var SPICRAWL_API_KEY
```

This writes the server to `~/.codex/config.toml`. Codex reads the key from `SPICRAWL_API_KEY` each time it connects and sends it as `Authorization: Bearer`, so the key never lands in the file. The equivalent config:

```toml title="~/.codex/config.toml"
[mcp_servers.spicrawl]
url = "https://mcp.spicrawl.com/mcp"
bearer_token_env_var = "SPICRAWL_API_KEY"
```

To scope the server to one project, put the same table in `.codex/config.toml` at the project root. Codex reads project config only in projects you have marked as trusted.

Check it with `codex mcp list`, or `/mcp` inside a session. The server shows 25 `spicrawl_*` tools. See [MCP server](https://docs.spicrawl.com/agents/mcp.md) for each tool.

## 2. Install the skill

Codex loads skills from `.agents/skills/` in the working directory and each parent up to the repository root, and from `$HOME/.agents/skills/` for every repository:

```bash
mkdir -p .agents/skills/spicrawl
curl -fsSL https://app.spicrawl.com/skill.md -o .agents/skills/spicrawl/SKILL.md
```

Codex picks the skill when a task matches its `description` (fetching, rendering or extracting web data). See [Agent skill](https://docs.spicrawl.com/agents/skill.md).

## 3. Add rules to AGENTS.md

Paste this section into `AGENTS.md` at the project root. Codex reads it before it starts work.

```markdown title="AGENTS.md"
## Web data (Spicrawl)

Use Spicrawl to read web pages: the spicrawl_* MCP tools, the `spicrawl` CLI, or
POST https://api.spicrawl.com/v1/scrape with `Authorization: Bearer $SPICRAWL_API_KEY` in code.
Docs: https://docs.spicrawl.com/llms.txt (append .md to any page URL for Markdown).

- Ask for markdown: `response_format: "markdown"` (API default is html); on spicrawl_scrape, `format: "markdown"`.
- Check the site's status, not only the HTTP status: `X-Target-Status` header, or `status`
  in the JSON envelope (spicrawl_scrape with `format: "json"`). 200 is the page; 404/410
  mean it does not exist; 403/429/503 mean the site refused, so escalate.
- On an error, switch on `code`. Retry only when `retryable` is true, after
  `retry_after_seconds`. Read `diagnostics.hint` and change what it names first.
  Never retry ERR::REQUEST::*, ERR::AUTH::* or ERR::LIMIT::QUOTA_EXCEEDED.
- Escalate one step at a time and stop at the first that works:
  plain fetch (1 credit) -> `js_render: true` (3) -> add the user's own `proxy`
  if they have one. Empty or skeleton content means render; ERR::UPSTREAM::CHALLENGE
  or a 403 target status means retry once, then the user's own proxy.
- Set `max_cost` on every request to the price of the step you intend.
- Keep the cache on (default). Set `cache: false` only for prices, stock or other live data.
- Trim tokens with `main_content_only` (on by default for markdown), `include_tags`, `exclude_tags`.
- For more than 20 URLs, use a batch job (spicrawl_batch_submit / POST /v1/batch).
  Batch items return raw HTML and apply only render, proxy and block_resources settings.
- With the CLI, read stdout as JSON and branch on the exit code: 3 auth, 4 bad request,
  5 limit, 6 site refused or failed, 7 proxy, 8 engine or extraction, 10 network.
- Never print or commit SPICRAWL_API_KEY. Log `X-Request-Id` for failures.
```

## Use the CLI instead of MCP

Codex works in a shell, so the `spicrawl` CLI is an alternative to the MCP server with no config at all. It prints JSON when stdout is not a terminal and exits with a stable code:

```bash
spicrawl scrape https://example.com/pricing --format markdown --max-cost 3 | jq -r .content
```

When the CLI gets the page as a document, its JSON output adds `status` (the site's status), `credits_charged` and `request_id` around `content`. See [CLI overview](https://docs.spicrawl.com/cli/overview.md).

## A first task to try

Start `codex` in the project and ask:

```text
Use Spicrawl to read https://example.com/pricing as markdown. List each plan with its monthly price.
If the page comes back empty, retry with rendering.
```

For a coding task:

```text
Add a fetch_page(url) function to this project that calls the Spicrawl API for markdown,
following the Web data rules in AGENTS.md. Read the key from SPICRAWL_API_KEY.
```

Compare the result with the reference implementation in [Best practices](https://docs.spicrawl.com/agents/best-practices.md).

## Codex cloud

Codex cloud tasks run in a hosted container, not on your machine. The `~/.codex/config.toml` from [step 1](#1-add-the-mcp-server) belongs to the CLI, IDE extension and desktop app, and OpenAI's docs warn that local MCP servers and transports may not be available in the cloud. Do not count on the MCP server there. Use the CLI or plain HTTP instead, both of which need only the key.

**Step 1: Set the key as an environment variable**

In the cloud environment's settings, add `SPICRAWL_API_KEY` as an **environment variable**, not a secret. Environment variables last for the whole task, setup script and agent phase. Secrets are only available to setup scripts and are removed before the agent phase starts, so the agent could not read the key from one.

**Step 2: Install the CLI in the setup script (optional)**

The setup script runs with internet access. Add:

```bash title="Setup script"
npm install -g @spicrawl/cli
```

Skip this step to use plain HTTP with `curl` instead.

**Step 3: Allow api.spicrawl.com**

Agent internet access is off by default. In the environment's settings, set it to **On**, choose the **None** allowlist preset and add the domain `api.spicrawl.com`. If you also restrict HTTP methods to `GET`, `HEAD` and `OPTIONS`, `POST /v1/scrape` is blocked, so allow `POST` too.

**Step 4: Add the rules to AGENTS.md**

Commit the section from [step 3](#3-add-rules-to-agentsmd) to the repository. Cloud tasks read `AGENTS.md` from the checkout.

Then the agent can run either of these:

```bash title="CLI"
spicrawl scrape https://example.com/pricing --format markdown --max-cost 3
```

```bash title="curl"
curl -sS https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com/pricing","response_format":"markdown","max_cost":3}'
```

Enabling internet access lets the agent reach any domain you allow. OpenAI recommends allowing only the domains and methods you need, because web content can carry prompt injection.

## Troubleshooting

| Symptom                                                                                            | Fix                                                                                                                                                         |
| -------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |
| The server fails to start or returns 401                                                           | `SPICRAWL_API_KEY` is not set in the shell that started `codex`. Export it and start a new session.                                                         |
| In a cloud task the agent says `SPICRAWL_API_KEY` is empty, or requests to `api.spicrawl.com` fail | The key was saved as a secret (removed before the agent phase) or agent internet access does not allow `api.spicrawl.com`. See [Codex cloud](#codex-cloud). |
| Tools from `.codex/config.toml` do not load                                                        | The project is not trusted. Trust it, or move the table to `~/.codex/config.toml`.                                                                          |
| The agent writes `js_render` on `spicrawl_scrape` and gets "unknown argument"                      | The tool's argument is `render`. The API field is `js_render`.                                                                                              |
| Usage tools return `ERR::AUTH::INSUFFICIENT_SCOPE`                                                 | Grant the `read` scope to the key in the dashboard.                                                                                                         |
