spicrawlspicrawlDocs

Best practices for agents

Rules for agents built on Spicrawl: an error-handling loop, a cheapest-first escalation ladder, cost and token control, result verification, and a reference fetch_page tool in Python and TypeScript.

An agent that uses Spicrawl well does four things: asks for markdown, checks the site's status as well as the HTTP status, acts on the error code instead of retrying blindly, and escalates one step at a time from the cheapest engine. This page gives the rules, then a complete fetch_page(url) tool that follows them.

The short version, for an agent's instructions:

- Request `response_format: "markdown"`; set `max_cost` on every call.
- 200 from Spicrawl is not success until `X-Target-Status` is 200 (404/410: page does not exist).
- On error: switch on `code`; retry only if `retryable`, after `retry_after_seconds`; act on `diagnostics.hint`.
- Escalate: fetch -> js_render -> your own proxy (if you have one). Stop at the first that works.
- More than 20 URLs: batch. Logins: sessions.

Handle errors by code

Every error is application/problem+json:

{
  "type": "https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE",
  "title": "Target served a bot challenge",
  "status": 502,
  "code": "ERR::UPSTREAM::CHALLENGE",
  "retryable": true,
  "doc_url": "https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE",
  "request_id": "01J9Z7A1B2C3D4E5F6G7H8J9KA",
  "target_status": 403,
  "diagnostics": {
    "hint": "Every engine tier was served a bot challenge rather than the page. Nothing was charged. The verdict is per exit and per moment, so a retry often passes."
  }
}

The loop an agent runs on each failure:

Switch on code

Never parse title or detail; they are for people. code is stable.

Act on diagnostics.hint

When diagnostics.hint is present it names the parameter to change. Change it before retrying the same request.

Retry only when retryable is true

Wait retry_after_seconds (or the Retry-After header) first; back off exponentially when neither is set. Cap retries at 2 or 3. Failed requests cost 0 credits, but a successful retry is billed as a new scrape because the API is not idempotent.

Otherwise, change the request or stop

A non-retryable error will fail the same way again. Fix the request, escalate, or report to the user with the request_id.

CodeHTTPRetryableWhat the agent does
ERR::UPSTREAM::CHALLENGE502yesThe site served a bot wall. Retry once, then add js_render: true; if it persists, route through the user's own proxy.
ERR::UPSTREAM::TIMEOUT504yesRetry once. On the fetch tier, try js_render: true; with wait_for, raise wait_for_timeout or fix the selector.
ERR::UPSTREAM::ERROR, CONNECTION_RESET, DNS_FAILED, TLS_FAILED502yesRetry once with backoff.
ERR::UPSTREAM::TOO_MANY_REDIRECTS502noStop; report the URL.
ERR::ENGINE::RENDER_FAILED502yesRetry once, then try engine: "chromium".
ERR::ENGINE::UNAVAILABLE503yesBack off and retry, or drop the engine pin.
ERR::LIMIT::RATE_LIMITED, CONCURRENCY_EXCEEDED429yesWait retry_after_seconds, then lower your parallelism.
ERR::LIMIT::MAX_COST_EXCEEDED400noThe request would cost more than max_cost. Nothing ran. Raise max_cost deliberately or pick a cheaper option.
ERR::LIMIT::QUOTA_EXCEEDED402noOut of credits or over the monthly ceiling. Stop and tell the user.
ERR::REQUEST::INVALID_PARAMETER, MISSING_PARAMETER, INCOMPATIBLE_FLAGS400noFix the request; detail names the field. Unknown fields are always rejected.
ERR::SECURITY::SSRF_BLOCKED, URL_BLOCKED, INVALID_URL400noThe URL cannot be scraped. Stop.
ERR::EXTRACT::FAILED502noThe target is not HTML (an image, a binary) or a PDF with parse_pdf: false. Ask for response_format: "html" or drop extraction.
ERR::SESSION::BUSY409yesAnother request is using the session. Run requests on one session one at a time.
ERR::SESSION::EXPIRED, RELEASED410noThe session's cookies are gone. Create a new session and log in again.
ERR::AUTH::*401, 403noStop. INSUFFICIENT_SCOPE names the missing scope.

The full table is on Errors.

Escalate from the cheapest engine

Start with the plain fetch and move up only when the result tells you to. You pay only for the attempt that succeeds; failures cost 0.

StepRequest fieldsCreditsMove up when
1. Fetchnone1The content is empty or a "Loading…" shell: go to 2. The site refused (X-Target-Status 403, 429, 503) or ERR::UPSTREAM::CHALLENGE: go to 3.
2. Renderjs_render: true3Refused or challenged: go to 3.
3. Your own proxyjs_render: true, proxy: "<your proxy URL>"3Stop and report; the page is not reachable today.

Step 1 already passes sites that reject non-browser TLS fingerprints: impersonate is on by default and free, so do not send impersonate: false. Stealth mode Coming soon (the Camoufox browser) and Spicrawl's managed residential exits Coming soon are not in this ladder yet; step 3 applies only when the user supplies a proxy. Where Camoufox runs, a step-3 request whose max_cost allows 25 moves to Camoufox by itself on a vendor challenge, billed only if Camoufox returns the page; max_cost: 3 keeps it on step 3. See Anti-bot.

If you do not want to write the ladder, send mode: "auto": Spicrawl escalates fetch → obscura itself and bills only the rung that succeeded. mode: "auto" cannot be combined with js_render or engine, and its max_cost check uses the ladder's top rung (25).

Prices are the published table from x-spicrawl-credits. Pricing is currently switched off for every engine except Chromium, so most requests charge 0 today. X-Request-Cost always shows what the request would cost; X-Credits-Charged shows what was billed.

Control cost

  • max_cost on every request. A request whose dearest possible outcome costs more is refused before it runs with ERR::LIMIT::MAX_COST_EXCEEDED at 0 credits. Set it to the price of the step you intend: 1 for a fetch, 3 for a render.
  • Keep the cache on. It is on by default; a hit costs 0 and shows Cache-State: hit. Set cache: false only for time-sensitive data (prices, stock). Requests with session_id or actions are never cached.
  • Cheapest engine first. See the ladder above. Do not start with js_render: true on pages that do not need it.
  • Batches have two caps. max_cost caps each item; credit_budget caps the whole job.
  • Read X-Credits-Charged after each call and keep a running total in the agent's state.

Choose the output

You needSendCredits addedNotes
Text to read or summariseresponse_format: "markdown"0Best for an LLM. The API default is html, so always set it.
The page's own metadata (title, price, author)autoparse: true0Returns JSON-LD, OpenGraph, microdata and embedded app state in data. No selectors, survives redesigns. Try it first.
Specific fields, known layoutextract: {"title": "h1", "price": ".price"}0Selector map, or a JSON Schema whose properties carry selector. Output in data; broken selectors are listed in empty_fields.
Specific fields, unknown layoutai_extract: {"prompt": "product name, price and stock status"} Coming soon4Give prompt or schema, not both. Output in data.
The data behind a JavaScript appjs_render: true, network_capture: {"resource_types": ["xhr", "fetch"]}0The page's own API responses, often cleaner than the DOM.

extract, autoparse, links, ai_extract, network_capture and screenshot all switch the response to the JSON envelope, with the page under content and the result under data.

Stay inside a token budget

  • main_content_only is on by default for markdown and strips nav, footer and aside. Leave it on.
  • include_tags: ["main article"] keeps only matching subtrees; exclude_tags: [".cookie-banner", "nav", ".related"] removes elements. Both take CSS selectors and apply before main_content_only.
  • For a list of links rather than the text, send links: true and read links from the envelope.
  • If the envelope has truncated: true, the target body exceeded the size cap and was cut.

Scale with batch

Above about 20 URLs, submit one POST /v1/batch job instead of a loop of scrapes. It accepts up to 10,000 URLs per call, runs them with server-side concurrency and retries, and you poll GET /v1/batch/{id} until status is completed, failed or cancelled.

Batch items apply only js_render, proxy and block_resources today, and return the raw page (HTML), not markdown or extraction output. Convert or extract on your side, or use single scrapes when you need markdown per item.

Resubmitting a batch creates and bills a second job. See Batch.

Use sessions for logins

For pages behind a login, create a session (POST /v1/sessions), log in once with actions and the session's session_id, then send the same session_id on every later scrape. The session keeps cookies, storage and one exit IP. A session runs one request at a time (ERR::SESSION::BUSY otherwise). Release it with POST /v1/sessions/{id}/release when done; releasing purges its cookies. See Sessions and logins.

Verify results before using them

A 200 from Spicrawl means the scrape ran. Check the result before the agent trusts it:

CheckWhereMeaning
Site statusX-Target-Status, or status in the envelope200 is the page. 404 and 410 are real answers (billed). Anything else was returned at 0 credits and means the site refused or failed.
Empty extractionempty_fields in the envelopeSelectors that matched nothing, usually a broken selector or an unrendered page.
WarningsX-Warning headers (CODE: message), warnings in the envelopeRENDER_DEGRADED means wait_for never appeared and the page is as it stood. FORMAT_COERCED means your format was switched to JSON.
Contentthe body or contentEmpty or a "Loading" placeholder means render (step 2).
CostX-Credits-ChargedWhat was billed.
BalanceX-Credits-RemainingMonthly allowance left after this request. Stop before it reaches the next step's max_cost, rather than discovering the ceiling as a 402. Absent if your organization has no monthly limit.

Reference tool: fetch_page(url)

Both implementations are a complete agent tool with no SDK. They request markdown, cap cost per step, check the site's status, retry only retryable errors, and escalate fetch → render → your own proxy (when MY_PROXY_URL is set). They return a small dict the agent can read, including the code, hint and request_id on failure.

import os, time, requests

API = "https://api.spicrawl.com/v1/scrape"
# Cheapest first. The second value is max_cost: the published price of that step.
LADDER = [({}, 1), ({"js_render": True}, 3)]
if os.environ.get("MY_PROXY_URL"):  # your own proxy; Spicrawl's managed pool is coming soon
    LADDER.append(({"js_render": True, "proxy": os.environ["MY_PROXY_URL"]}, 3))
NEW_EXIT = 2  # first step that changes the exit IP

def fetch_page(url: str) -> dict:
    """Agent tool: return a web page as markdown, escalating only as far as needed."""
    headers = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
    step, retries = 0, 0
    while step < len(LADDER):
        flags, price = LADDER[step]
        body = {"url": url, "response_format": "markdown", "max_cost": price, **flags}
        r = requests.post(API, json=body, headers=headers, timeout=180)
        if r.ok:
            target = int(r.headers.get("X-Target-Status") or 0)
            if target in (404, 410):
                return {"ok": False, "error": f"page does not exist (HTTP {target})"}
            if target == 200 and (step > 0 or len(r.text.strip()) >= 200):
                return {"ok": True, "markdown": r.text, "engine": r.headers.get("X-Engine"),
                        "credits": int(r.headers.get("X-Credits-Charged") or 0)}
            # 200 but near-empty on the fetch tier: render. Refused (403/429/503): new exit.
            step = step + 1 if target == 200 else max(step + 1, NEW_EXIT)
            continue
        try:
            p = r.json()
        except ValueError:
            p = {"code": f"HTTP_{r.status_code}", "retryable": r.status_code >= 500}
        code = p.get("code", "")
        if code == "ERR::UPSTREAM::CHALLENGE":
            step = max(step + 1, NEW_EXIT)
        elif p.get("retryable") and retries < 2:
            retries += 1
            time.sleep(p.get("retry_after_seconds") or 2 ** retries)
        else:
            return {"ok": False, "code": code, "detail": p.get("detail"),
                    "hint": (p.get("diagnostics") or {}).get("hint"),
                    "request_id": p.get("request_id")}
    return {"ok": False, "code": "ESCALATION_EXHAUSTED",
            "detail": "every step of the ladder was refused by the site"}

What these tools do and why:

  • max_cost per step. A step that would cost more than planned is refused before it runs, so the most one call can bill is the price of the step that succeeded (3 credits).
  • ERR::LIMIT::MAX_COST_EXCEEDED on a step means the deployment prices that step above the table (for example, rendering served by Chromium at 8 credits). The tool returns it with the detail, which states the real price; raise that step's max_cost if you accept it.
  • Near-empty threshold. 200 characters of markdown on the fetch tier is treated as an unrendered shell and moves to rendering. Tune it for your targets; a short page such as https://example.com renders the same on every step.
  • Target status. A refusal (403, 429, 503) comes back as HTTP 200 at 0 credits, so the tool reads X-Target-Status rather than trusting r.ok. Refusals skip rendering and go straight to your own proxy when one is set, because rendering does not fix IP reputation.
  • Retries. Only errors with retryable: true, at most twice per call, after retry_after_seconds. ERR::UPSTREAM::CHALLENGE escalates instead of repeating the same request.
  • Failure output. Non-retryable errors return code, detail, hint and request_id, so the agent can explain the failure or fix its request. Pass request_id to GET /v1/requests/{id} for the full trace.

Register the function as a tool with a description such as: "Fetch a web page and return its main content as markdown. Use for reading any public URL. Returns ok=false with an error code when the page cannot be fetched."

On this page