# Best practices for agents

> Rules for agents built on Spicrawl: an error-handling loop, a cheapest-first escalation ladder, cost and token control, result verification, and a reference fetch_page tool in Python and TypeScript.

Source: https://docs.spicrawl.com/agents/best-practices

An agent that uses Spicrawl well does four things: asks for markdown, checks the site's status as well as the HTTP status, acts on the error `code` instead of retrying blindly, and escalates one step at a time from the cheapest engine. This page gives the rules, then a complete `fetch_page(url)` tool that follows them.

The short version, for an agent's instructions:

```markdown
- Request `response_format: "markdown"`; set `max_cost` on every call.
- 200 from Spicrawl is not success until `X-Target-Status` is 200 (404/410: page does not exist).
- On error: switch on `code`; retry only if `retryable`, after `retry_after_seconds`; act on `diagnostics.hint`.
- Escalate: fetch -> js_render -> your own proxy (if you have one). Stop at the first that works.
- More than 20 URLs: batch. Logins: sessions.
```

## Handle errors by code

Every error is `application/problem+json`:

```json
{
  "type": "https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE",
  "title": "Target served a bot challenge",
  "status": 502,
  "code": "ERR::UPSTREAM::CHALLENGE",
  "retryable": true,
  "doc_url": "https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE",
  "request_id": "01J9Z7A1B2C3D4E5F6G7H8J9KA",
  "target_status": 403,
  "diagnostics": {
    "hint": "Every engine tier was served a bot challenge rather than the page. Nothing was charged. The verdict is per exit and per moment, so a retry often passes."
  }
}
```

The loop an agent runs on each failure:

**Step 1: Switch on code**

Never parse `title` or `detail`; they are for people. `code` is stable.

**Step 2: Act on diagnostics.hint**

When `diagnostics.hint` is present it names the parameter to change. Change it before retrying the same request.

**Step 3: Retry only when retryable is true**

Wait `retry_after_seconds` (or the `Retry-After` header) first; back off exponentially when neither is set. Cap retries at 2 or 3. Failed requests cost 0 credits, but a successful retry is billed as a new scrape because the API is not idempotent.

**Step 4: Otherwise, change the request or stop**

A non-retryable error will fail the same way again. Fix the request, escalate, or report to the user with the `request_id`.

| Code                                                                         | HTTP     | Retryable | What the agent does                                                                                                                 |
| ---------------------------------------------------------------------------- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `ERR::UPSTREAM::CHALLENGE`                                                   | 502      | yes       | The site served a bot wall. Retry once, then add `js_render: true`; if it persists, route through the user's own `proxy`.           |
| `ERR::UPSTREAM::TIMEOUT`                                                     | 504      | yes       | Retry once. On the fetch tier, try `js_render: true`; with `wait_for`, raise `wait_for_timeout` or fix the selector.                |
| `ERR::UPSTREAM::ERROR`, `CONNECTION_RESET`, `DNS_FAILED`, `TLS_FAILED`       | 502      | yes       | Retry once with backoff.                                                                                                            |
| `ERR::UPSTREAM::TOO_MANY_REDIRECTS`                                          | 502      | no        | Stop; report the URL.                                                                                                               |
| `ERR::ENGINE::RENDER_FAILED`                                                 | 502      | yes       | Retry once, then try `engine: "chromium"`.                                                                                          |
| `ERR::ENGINE::UNAVAILABLE`                                                   | 503      | yes       | Back off and retry, or drop the `engine` pin.                                                                                       |
| `ERR::LIMIT::RATE_LIMITED`, `CONCURRENCY_EXCEEDED`                           | 429      | yes       | Wait `retry_after_seconds`, then lower your parallelism.                                                                            |
| `ERR::LIMIT::MAX_COST_EXCEEDED`                                              | 400      | no        | The request would cost more than `max_cost`. Nothing ran. Raise `max_cost` deliberately or pick a cheaper option.                   |
| `ERR::LIMIT::QUOTA_EXCEEDED`                                                 | 402      | no        | Out of credits or over the monthly ceiling. Stop and tell the user.                                                                 |
| `ERR::REQUEST::INVALID_PARAMETER`, `MISSING_PARAMETER`, `INCOMPATIBLE_FLAGS` | 400      | no        | Fix the request; `detail` names the field. Unknown fields are always rejected.                                                      |
| `ERR::SECURITY::SSRF_BLOCKED`, `URL_BLOCKED`, `INVALID_URL`                  | 400      | no        | The URL cannot be scraped. Stop.                                                                                                    |
| `ERR::EXTRACT::FAILED`                                                       | 502      | no        | The target is not HTML (an image, a binary) or a PDF with `parse_pdf: false`. Ask for `response_format: "html"` or drop extraction. |
| `ERR::SESSION::BUSY`                                                         | 409      | yes       | Another request is using the session. Run requests on one session one at a time.                                                    |
| `ERR::SESSION::EXPIRED`, `RELEASED`                                          | 410      | no        | The session's cookies are gone. Create a new session and log in again.                                                              |
| `ERR::AUTH::*`                                                               | 401, 403 | no        | Stop. `INSUFFICIENT_SCOPE` names the missing scope.                                                                                 |

The full table is on [Errors](https://docs.spicrawl.com/errors.md).

## Escalate from the cheapest engine

Start with the plain fetch and move up only when the result tells you to. You pay only for the attempt that succeeds; failures cost 0.

| Step              | Request fields                               | Credits | Move up when                                                                                                                                    |
| ----------------- | -------------------------------------------- | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| 1. Fetch          | none                                         | 1       | The content is empty or a "Loading…" shell: go to 2. The site refused (`X-Target-Status` 403, 429, 503) or `ERR::UPSTREAM::CHALLENGE`: go to 3. |
| 2. Render         | `js_render: true`                            | 3       | Refused or challenged: go to 3.                                                                                                                 |
| 3. Your own proxy | `js_render: true, proxy: "<your proxy URL>"` | 3       | Stop and report; the page is not reachable today.                                                                                               |

Step 1 already passes sites that reject non-browser TLS fingerprints: `impersonate` is on by default and free, so do not send `impersonate: false`. Stealth mode (coming soon) (the Camoufox browser) and Spicrawl's [managed residential exits](https://docs.spicrawl.com/guides/proxies-and-geo.md#managed-proxy-pool-coming-soon) (coming soon) are not in this ladder yet; step 3 applies only when the user supplies a proxy. Where Camoufox runs, a step-3 request whose `max_cost` allows 25 moves to Camoufox by itself on a vendor challenge, billed only if Camoufox returns the page; `max_cost: 3` keeps it on step 3. See [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md#automatic-escalation-coming-soon).

If you do not want to write the ladder, send `mode: "auto"`: Spicrawl escalates fetch → obscura itself and bills only the rung that succeeded. `mode: "auto"` cannot be combined with `js_render` or `engine`, and its `max_cost` check uses the ladder's top rung (25).

> **Note:** Prices are the published table from `x-spicrawl-credits`. Pricing is currently switched off for every engine except Chromium, so most requests charge 0 today. `X-Request-Cost` always shows what the request would cost; `X-Credits-Charged` shows what was billed.

## Control cost

* **`max_cost` on every request.** A request whose dearest possible outcome costs more is refused before it runs with `ERR::LIMIT::MAX_COST_EXCEEDED` at 0 credits. Set it to the price of the step you intend: `1` for a fetch, `3` for a render.
* **Keep the cache on.** It is on by default; a hit costs 0 and shows `Cache-State: hit`. Set `cache: false` only for time-sensitive data (prices, stock). Requests with `session_id` or `actions` are never cached.
* **Cheapest engine first.** See the ladder above. Do not start with `js_render: true` on pages that do not need it.
* **Batches have two caps.** `max_cost` caps each item; `credit_budget` caps the whole job.
* **Read `X-Credits-Charged`** after each call and keep a running total in the agent's state.

## Choose the output

| You need                                       | Send                                                                           | Credits added | Notes                                                                                                                              |
| ---------------------------------------------- | ------------------------------------------------------------------------------ | ------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Text to read or summarise                      | `response_format: "markdown"`                                                  | 0             | Best for an LLM. The API default is `html`, so always set it.                                                                      |
| The page's own metadata (title, price, author) | `autoparse: true`                                                              | 0             | Returns JSON-LD, OpenGraph, microdata and embedded app state in `data`. No selectors, survives redesigns. Try it first.            |
| Specific fields, known layout                  | `extract: {"title": "h1", "price": ".price"}`                                  | 0             | Selector map, or a JSON Schema whose properties carry `selector`. Output in `data`; broken selectors are listed in `empty_fields`. |
| Specific fields, unknown layout                | `ai_extract: {"prompt": "product name, price and stock status"}` (coming soon) | 4             | Give `prompt` or `schema`, not both. Output in `data`.                                                                             |
| The data behind a JavaScript app               | `js_render: true, network_capture: {"resource_types": ["xhr", "fetch"]}`       | 0             | The page's own API responses, often cleaner than the DOM.                                                                          |

`extract`, `autoparse`, `links`, `ai_extract`, `network_capture` and `screenshot` all switch the response to the JSON envelope, with the page under `content` and the result under `data`.

## Stay inside a token budget

* `main_content_only` is on by default for markdown and strips nav, footer and aside. Leave it on.
* `include_tags: ["main article"]` keeps only matching subtrees; `exclude_tags: [".cookie-banner", "nav", ".related"]` removes elements. Both take CSS selectors and apply before `main_content_only`.
* For a list of links rather than the text, send `links: true` and read `links` from the envelope.
* If the envelope has `truncated: true`, the target body exceeded the size cap and was cut.

## Scale with batch

Above about 20 URLs, submit one `POST /v1/batch` job instead of a loop of scrapes. It accepts up to 10,000 URLs per call, runs them with server-side concurrency and retries, and you poll `GET /v1/batch/{id}` until `status` is `completed`, `failed` or `cancelled`.

> **Warning:** Batch items apply only `js_render`, `proxy` and `block_resources` today, and return the raw page (HTML), not markdown or extraction output. Convert or extract on your side, or use single scrapes when you need markdown per item.

Resubmitting a batch creates and bills a second job. See [Batch](https://docs.spicrawl.com/guides/batch.md).

## Use sessions for logins

For pages behind a login, create a session (`POST /v1/sessions`), log in once with `actions` and the session's `session_id`, then send the same `session_id` on every later scrape. The session keeps cookies, storage and one exit IP. A session runs one request at a time (`ERR::SESSION::BUSY` otherwise). Release it with `POST /v1/sessions/{id}/release` when done; releasing purges its cookies. See [Sessions and logins](https://docs.spicrawl.com/guides/sessions-and-logins.md).

## Verify results before using them

A `200` from Spicrawl means the scrape ran. Check the result before the agent trusts it:

| Check            | Where                                                             | Meaning                                                                                                                                                                                         |
| ---------------- | ----------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Site status      | `X-Target-Status`, or `status` in the envelope                    | 200 is the page. 404 and 410 are real answers (billed). Anything else was returned at 0 credits and means the site refused or failed.                                                           |
| Empty extraction | `empty_fields` in the envelope                                    | Selectors that matched nothing, usually a broken selector or an unrendered page.                                                                                                                |
| Warnings         | `X-Warning` headers (`CODE: message`), `warnings` in the envelope | `RENDER_DEGRADED` means `wait_for` never appeared and the page is as it stood. `FORMAT_COERCED` means your format was switched to JSON.                                                         |
| Content          | the body or `content`                                             | Empty or a "Loading" placeholder means render (step 2).                                                                                                                                         |
| Cost             | `X-Credits-Charged`                                               | What was billed.                                                                                                                                                                                |
| Balance          | `X-Credits-Remaining`                                             | Monthly allowance left after this request. Stop before it reaches the next step's `max_cost`, rather than discovering the ceiling as a `402`. Absent if your organization has no monthly limit. |

## Reference tool: fetch\_page(url)

Both implementations are a complete agent tool with no SDK. They request markdown, cap cost per step, check the site's status, retry only retryable errors, and escalate fetch → render → your own proxy (when `MY_PROXY_URL` is set). They return a small dict the agent can read, including the `code`, `hint` and `request_id` on failure.

```python title="fetch_page.py"
import os, time, requests

API = "https://api.spicrawl.com/v1/scrape"
# Cheapest first. The second value is max_cost: the published price of that step.
LADDER = [({}, 1), ({"js_render": True}, 3)]
if os.environ.get("MY_PROXY_URL"):  # your own proxy; Spicrawl's managed pool is coming soon
    LADDER.append(({"js_render": True, "proxy": os.environ["MY_PROXY_URL"]}, 3))
NEW_EXIT = 2  # first step that changes the exit IP

def fetch_page(url: str) -> dict:
    """Agent tool: return a web page as markdown, escalating only as far as needed."""
    headers = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
    step, retries = 0, 0
    while step < len(LADDER):
        flags, price = LADDER[step]
        body = {"url": url, "response_format": "markdown", "max_cost": price, **flags}
        r = requests.post(API, json=body, headers=headers, timeout=180)
        if r.ok:
            target = int(r.headers.get("X-Target-Status") or 0)
            if target in (404, 410):
                return {"ok": False, "error": f"page does not exist (HTTP {target})"}
            if target == 200 and (step > 0 or len(r.text.strip()) >= 200):
                return {"ok": True, "markdown": r.text, "engine": r.headers.get("X-Engine"),
                        "credits": int(r.headers.get("X-Credits-Charged") or 0)}
            # 200 but near-empty on the fetch tier: render. Refused (403/429/503): new exit.
            step = step + 1 if target == 200 else max(step + 1, NEW_EXIT)
            continue
        try:
            p = r.json()
        except ValueError:
            p = {"code": f"HTTP_{r.status_code}", "retryable": r.status_code >= 500}
        code = p.get("code", "")
        if code == "ERR::UPSTREAM::CHALLENGE":
            step = max(step + 1, NEW_EXIT)
        elif p.get("retryable") and retries < 2:
            retries += 1
            time.sleep(p.get("retry_after_seconds") or 2 ** retries)
        else:
            return {"ok": False, "code": code, "detail": p.get("detail"),
                    "hint": (p.get("diagnostics") or {}).get("hint"),
                    "request_id": p.get("request_id")}
    return {"ok": False, "code": "ESCALATION_EXHAUSTED",
            "detail": "every step of the ladder was refused by the site"}
```

```typescript title="fetchPage.ts"
const API = "https://api.spicrawl.com/v1/scrape";
// Cheapest first. The second value is max_cost: the published price of that step.
const LADDER: [Record<string, unknown>, number][] = [[{}, 1], [{ js_render: true }, 3]];
if (process.env.MY_PROXY_URL) // your own proxy; Spicrawl's managed pool is coming soon
  LADDER.push([{ js_render: true, proxy: process.env.MY_PROXY_URL }, 3]);
const NEW_EXIT = 2; // first step that changes the exit IP
const sleep = (s: number) => new Promise((r) => setTimeout(r, s * 1000));

/** Agent tool: return a web page as markdown, escalating only as far as needed. */
export async function fetchPage(url: string): Promise<Record<string, unknown>> {
  const headers = {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  };
  let step = 0, retries = 0;
  while (step < LADDER.length) {
    const [flags, price] = LADDER[step];
    const body = { url, response_format: "markdown", max_cost: price, ...flags };
    const r = await fetch(API, { method: "POST", headers, body: JSON.stringify(body),
                                 signal: AbortSignal.timeout(180_000) });
    if (r.ok) {
      const target = Number(r.headers.get("X-Target-Status") ?? 0);
      const text = await r.text();
      if (target === 404 || target === 410)
        return { ok: false, error: `page does not exist (HTTP ${target})` };
      if (target === 200 && (step > 0 || text.trim().length >= 200))
        return { ok: true, markdown: text, engine: r.headers.get("X-Engine"),
                 credits: Number(r.headers.get("X-Credits-Charged") ?? 0) };
      // 200 but near-empty on the fetch tier: render. Refused (403/429/503): new exit.
      step = target === 200 ? step + 1 : Math.max(step + 1, NEW_EXIT);
      continue;
    }
    const p = await r.json().catch(() => ({ code: `HTTP_${r.status}`, retryable: r.status >= 500 }));
    if (p.code === "ERR::UPSTREAM::CHALLENGE") {
      step = Math.max(step + 1, NEW_EXIT);
    } else if (p.retryable && retries < 2) {
      retries += 1;
      await sleep(p.retry_after_seconds ?? 2 ** retries);
    } else {
      return { ok: false, code: p.code, detail: p.detail,
               hint: p.diagnostics?.hint, request_id: p.request_id };
    }
  }
  return { ok: false, code: "ESCALATION_EXHAUSTED",
           detail: "every step of the ladder was refused by the site" };
}
```

What these tools do and why:

* **`max_cost` per step.** A step that would cost more than planned is refused before it runs, so the most one call can bill is the price of the step that succeeded (3 credits).
* **`ERR::LIMIT::MAX_COST_EXCEEDED` on a step** means the deployment prices that step above the table (for example, rendering served by Chromium at 8 credits). The tool returns it with the `detail`, which states the real price; raise that step's `max_cost` if you accept it.
* **Near-empty threshold.** 200 characters of markdown on the fetch tier is treated as an unrendered shell and moves to rendering. Tune it for your targets; a short page such as `https://example.com` renders the same on every step.
* **Target status.** A refusal (403, 429, 503) comes back as HTTP 200 at 0 credits, so the tool reads `X-Target-Status` rather than trusting `r.ok`. Refusals skip rendering and go straight to your own proxy when one is set, because rendering does not fix IP reputation.
* **Retries.** Only errors with `retryable: true`, at most twice per call, after `retry_after_seconds`. `ERR::UPSTREAM::CHALLENGE` escalates instead of repeating the same request.
* **Failure output.** Non-retryable errors return `code`, `detail`, `hint` and `request_id`, so the agent can explain the failure or fix its request. Pass `request_id` to `GET /v1/requests/{id}` for the full trace.

Register the function as a tool with a description such as: "Fetch a web page and return its main content as markdown. Use for reading any public URL. Returns ok=false with an error code when the page cannot be fetched."
