spicrawlspicrawlDocs

Errors

Every Spicrawl error code with its HTTP status, whether to retry, what it means and which parameter to change.

Every Spicrawl error is an application/problem+json body with a stable code of the form ERR::FAMILY::NAME. Switch on code, retry only when retryable is true, wait Retry-After seconds when it is present, and read diagnostics.hint before you retry: it names the parameter to change. Failed requests cost 0 credits.

502 ERR::UPSTREAM::CHALLENGE
{
  "type": "https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE",
  "title": "Target served a bot challenge",
  "status": 502,
  "code": "ERR::UPSTREAM::CHALLENGE",
  "detail": "Cloudflare served a bot challenge instead of the page, so no content could be fetched. Nothing was charged.",
  "retryable": true,
  "doc_url": "https://docs.spicrawl.com/errors#UPSTREAM_CHALLENGE",
  "request_id": "01M0HF5WFWE7PRE8KZHDTNETWN",
  "instance": "/v1/scrape",
  "target_status": 403,
  "diagnostics": {
    "failed_at": "fetch",
    "hint": "The verdict is per exit and per moment, so a retry often passes."
  }
}

doc_url points at this page. The fragment is the code without ERR::, with :: replaced by _, so ERR::PROXY::EXHAUSTED links to #PROXY_EXHAUSTED. Every fragment names one code. The same FAMILY_NAME form is the error_code in the request log. The short form (#EXHAUSTED) still resolves, for older links. Four short forms are shared by two codes each (NOT_FOUND, RATE_LIMITED, ERROR, UNAVAILABLE); their sections cover both codes, split by family.

Two statuses

The HTTP status is the platform's. The target site's status is target_status on an error, and X-Target-Status plus envelope status on a success.

  • A site that answered 403 or 500 is still a successful API call: you get HTTP 200, X-Target-Status: 403, the site's body, and 0 credits.
  • An error body with status: 404 and code: ERR::REQUEST::NOT_FOUND means a Spicrawl resource does not exist. It says nothing about the target URL.
  • target_status is null unless the failure came from the site.

original_status=true puts the target's status on the HTTP status line of a successful scrape. With it set, a target 404 and an API error look the same by status alone. Leave it off unless a client forces you to use it, and branch on Content-Type: application/problem+json if you do.

The problem body

FieldTypeAlways presentMeaning
typeURIyesLink to this code's section on this page.
titlestringyesShort fixed summary of the code. Do not parse it.
statusintegeryesThe platform's HTTP status for this response. Not the target site's.
codestringyesStable code, ERR::FAMILY::NAME. Switch on this.
detailstringnoWhat happened on this occurrence, in prose. Show it to a human; do not parse it.
retryablebooleanyesThe server's verdict on whether the same request can succeed on retry. Prefer it to guessing from status.
doc_urlURIyesSame link as type.
request_idstringnoULID of the failed request, also in X-Request-Id. Pass it to GET /v1/requests/{id} for the full trace, and quote it to support.
instancestringnoRequest path that produced the error, for example /v1/scrape.
target_statusinteger or nullyesThe target site's HTTP status when the failure came from the site, otherwise null.
retry_after_secondsintegernoMirrors the Retry-After header.
warningsstring[]noNon-fatal decisions the platform made for you, such as a flag ignored because proxy was set.
diagnosticsobjectnoWhere the request failed. See below.

diagnostics fields:

FieldMeaning
failed_atName of the stage that failed. Always the last row of timeline.
timeline[]Stages in order (stage, ms, ok, detail), stopping at the failure.
attempts[]Each execution attempt (n, engine, proxy, outcome, ms, error_code). Under mode=auto, or when a bot challenge moved the request to another exit or engine, these span every rung.
budget_ms, elapsed_msTime budget and time spent. elapsed_ms close to budget_ms means a timeout; much lower means a refusal.
engineEngine that served the final attempt.
proxyRedacted exit: source (custom, direct, or pool for the managed pool Coming soon), endpoint (an identifier, not an IP), country. Never contains credentials.
hintThe next thing to try, naming the parameter to change. Act on it before retrying.

Retry decision table

SituationWhat to do
retryable: true and Retry-After presentWait Retry-After seconds (or retry_after_seconds), then send the same request.
retryable: true, no Retry-AfterExponential backoff with jitter: about 1 s, 2 s, 4 s. Stop after 3 retries.
retryable: true and diagnostics.hint names a parameterRetry once unchanged. If it fails again, apply the hint (for example js_render=true).
retryable: false, HTTP 4xxDo not retry unchanged. Fix the request using code, detail and diagnostics.hint.
ERR::LIMIT::QUOTA_EXCEEDED (402)Stop. Retrying cannot succeed until the monthly credit ceiling is raised or it resets; the response names the exact reset date.
ERR::SESSION::EXPIRED or ERR::SESSION::RELEASED (410)Stop using that session. Create a new one and log in again.
Network error or client timeout with no responseRetry at most once. POST /v1/scrape is not idempotent, so if the first call completed on the server you are billed twice.

A retry of a failed request costs nothing extra: failures cost 0 credits. Only a success is billed.

retry.py
import os
import random
import time

import requests

API = "https://api.spicrawl.com/v1/scrape"
HEADERS = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}


def scrape(body: dict, max_retries: int = 3) -> requests.Response:
    for attempt in range(max_retries + 1):
        r = requests.post(API, json=body, headers=HEADERS, timeout=120)
        if r.ok:
            return r
        problem = r.json()
        if not problem.get("retryable") or attempt == max_retries:
            hint = (problem.get("diagnostics") or {}).get("hint")
            raise RuntimeError(f"{problem['code']}: {problem.get('detail')} hint={hint}")
        wait = int(r.headers.get("Retry-After", 0)) or 2**attempt
        time.sleep(wait + random.random())
    raise AssertionError("unreachable")

All codes

Find a code with Ctrl-F. Every failure listed here costs 0 credits.

CodeHTTPRetryable
ERR::AUTH::MISSING_KEY401no
ERR::AUTH::INVALID_KEY401no
ERR::AUTH::REVOKED_KEY401no
ERR::AUTH::EXPIRED_KEY401no
ERR::AUTH::FORBIDDEN403no
ERR::AUTH::INSUFFICIENT_SCOPE403no
ERR::AUTH::ENGINE_NOT_ENTITLED403no
ERR::AUTH::INVALID_TOKEN401no
ERR::AUTH::INVALID_CREDENTIALS401no
ERR::REQUEST::INVALID400no
ERR::REQUEST::MISSING_PARAMETER400no
ERR::REQUEST::INVALID_PARAMETER400no
ERR::REQUEST::INCOMPATIBLE_FLAGS400no
ERR::REQUEST::CAPABILITY_UNSUPPORTED400no
ERR::REQUEST::NOT_FOUND404no
ERR::REQUEST::BEYOND_RETENTION410no
ERR::REQUEST::CONFLICT409no
ERR::REQUEST::PAYLOAD_TOO_LARGE413no
ERR::SECURITY::SSRF_BLOCKED400no
ERR::SECURITY::URL_BLOCKED400no
ERR::SECURITY::INVALID_URL400no
ERR::SESSION::NOT_FOUND404no
ERR::SESSION::EXPIRED410no
ERR::SESSION::RELEASED410no
ERR::SESSION::BUSY409yes
ERR::SESSION::STATE_CORRUPT500no
ERR::LIMIT::RATE_LIMITED429yes
ERR::LIMIT::CONCURRENCY_EXCEEDED429yes
ERR::LIMIT::QUOTA_EXCEEDED402no
ERR::LIMIT::MAX_COST_EXCEEDED400no
ERR::LIMIT::SESSIONS_EXCEEDED429yes
ERR::UPSTREAM::TIMEOUT504yes
ERR::UPSTREAM::DNS_FAILED502yes
ERR::UPSTREAM::TLS_FAILED502yes
ERR::UPSTREAM::CONNECTION_RESET502yes
ERR::UPSTREAM::TOO_MANY_REDIRECTS502no
ERR::UPSTREAM::ERROR502yes
ERR::UPSTREAM::CHALLENGE502yes
ERR::PROXY::EXHAUSTED503yes
ERR::PROXY::UNREACHABLE502no
ERR::PROXY::AUTH_FAILED502no
ERR::PROXY::RATE_LIMITED502yes
ERR::ENGINE::RENDER_FAILED502yes
ERR::ENGINE::UNAVAILABLE503 (501 on /v1/browser)yes (no when 501)
ERR::EXTRACT::FAILED502no
ERR::EXTRACT::INVALID_RULES400no
ERR::INTERNAL::ERROR500yes
ERR::INTERNAL::UNAVAILABLE503yes

AUTH: who you are and what the key may do

ERR::AUTH::MISSING_KEY

401, not retryable, 0 credits. No API key reached the server.

  • Send Authorization: Bearer $SPICRAWL_API_KEY. Check that the variable is set in the process that makes the call.
  • /v1/scrape ignores ?apikey= and ?api_key=. Only /v1/browser reads a key from the query string. A scrape that relies on the query key gets this error.

ERR::AUTH::INVALID_KEY

401, not retryable, 0 credits. The key is malformed or does not exist. The two cases return the same answer on purpose.

  • A valid key is spicrawl_live_ or spicrawl_test_ followed by 52 characters from a-z and 2-7. Look for a truncated paste, quotes or a trailing newline.
  • Copy the key again from the dashboard at app.spicrawl.com.

ERR::AUTH::REVOKED_KEY

401, not retryable, 0 credits. Someone in your organization revoked this key. Create a replacement in the dashboard and update SPICRAWL_API_KEY.

ERR::AUTH::EXPIRED_KEY

401, not retryable, 0 credits. The key passed its expiry date. Create a replacement in the dashboard.

ERR::AUTH::FORBIDDEN

403, not retryable, 0 credits. The key is valid but the action is not permitted for its organization, for example because the organization is suspended. detail says which. Contact support with the request_id.

ERR::AUTH::INSUFFICIENT_SCOPE

403, not retryable, 0 credits. The key does not carry the scope this route needs. detail names the scope, for example "This API key does not carry the browser scope."

RouteScope
/v1/scrapescrape
/v1/batch/*batch
/v1/sessions/*sessions
/v1/browserbrowser
/v1/usage/*read
/v1/requests/*none beyond a valid key for the key's own project; read for all_projects=true (another project's single request answers 404 without it)

New keys get scrape, batch and sessions. Create a key with browser or read in the dashboard, or use the workspace default key, which carries all five. See Authentication.

ERR::AUTH::ENGINE_NOT_ENTITLED

403, not retryable, 0 credits. You pinned an engine your plan does not include. detail lists the engines your plan does include. Remove the engine pin and use js_render=true, or change plan.

ERR::AUTH::INVALID_TOKEN

401, not retryable, 0 credits. A dashboard session token is invalid or expired. This applies to dashboard sign-in, not to API keys. Sign in again.

ERR::AUTH::INVALID_CREDENTIALS

401, not retryable, 0 credits. Dashboard sign-in failed: the email or password is incorrect. API calls never return this code.

REQUEST: the request itself is wrong

ERR::REQUEST::INVALID

400, not retryable, 0 credits. The server could not read the request: the body is empty, is not valid JSON, or Content-Type is not application/json. Send a JSON object with Content-Type: application/json.

ERR::REQUEST::MISSING_PARAMETER

400, not retryable, 0 credits. A required parameter is absent. detail names it. The usual case is a scrape without url, or a batch with neither urls nor items.

ERR::REQUEST::INVALID_PARAMETER

400, not retryable, 0 credits. A parameter has a wrong type or value, or is not a parameter of this API. detail names it.

  • JSON bodies are decoded strictly. A misspelled or unknown field is refused, never ignored. The same holds for unknown query parameters on GET /v1/scrape.
  • Common causes: headers instead of custom_headers; a hop-by-hop header such as Host or Connection inside custom_headers; a string where a boolean or integer is expected.
  • Parameter names borrowed from other scraping APIs are refused with the Spicrawl equivalent named in detail.

ERR::REQUEST::INCOMPATIBLE_FLAGS

400, not retryable, 0 credits. Two parameters contradict each other, or one needs another that is missing. detail names both. Common pairs:

You sentFix
wait, wait_for, actions or block_resources without a browserAdd js_render=true (or mode=auto).
mode=auto with js_render or enginePick one: mode=auto alone, or an explicit engine flag.
proxy_verify without proxyAdd proxy, or remove proxy_verify.
wait_for_timeout without wait_forAdd wait_for.
cache=false with cache_ttl above 0Send only cache=false.
session_id with an engine that differs from the session's engineRemove engine; the session's engine is used.
Batch with both urls and itemsSend one of them.

ERR::REQUEST::CAPABILITY_UNSUPPORTED

400, not retryable, 0 credits. The engine that would run cannot do what you asked. detail names the capability.

  • screenshot=true or response_format=pdf needs a rasterising engine: set engine=chromium. mode=auto cannot satisfy a screenshot.
  • headless=false works only with engine=chromium.

ERR::REQUEST::NOT_FOUND and ERR::SESSION::NOT_FOUND

Both codes link here. Both are 404, not retryable, 0 credits, and both are about a Spicrawl resource, never about the target URL. A target that answers 404 is a successful scrape with X-Target-Status: 404.

ERR::REQUEST::NOT_FOUND: no batch job, batch task, request-log record or other resource with that id exists for this key's organization or project.

  • Check the id. Batch jobs and request-log records are visible only to keys in the project that created them.
  • Request-log records expire after the retention window; an expired cursor returns ERR::REQUEST::BEYOND_RETENTION instead.

ERR::SESSION::NOT_FOUND: no session with that session_id exists for this organization.

  • Check the session_id you passed to /v1/scrape or /v1/sessions/{sessionID}.
  • A malformed id, a deleted session and another organization's session all get this same answer.
  • On /v1/scrape the session is checked before any work, so nothing is fetched and nothing is charged.
  • If the session existed and was released or expired, you get ERR::SESSION::RELEASED or ERR::SESSION::EXPIRED (410) instead.
  • Create a new session with POST /v1/sessions.

ERR::REQUEST::BEYOND_RETENTION

410, not retryable, 0 credits. The records existed and have expired.

  • GET /v1/requests: the cursor (before, before_id) points past your request-log retention window. Stop paging and restart from the first page without a cursor. For durable totals use GET /v1/usage.
  • GET /v1/batch/{batchID}/results: the job's results_expire_at has passed (72 hours after submission by default). The results are gone; the job object stays readable with GET /v1/batch/{batchID}. Download results before results_expire_at next time.

ERR::REQUEST::CONFLICT

409, not retryable, 0 credits. The resource already exists or conflicts with current state. The common case: another live session in this project already holds the sticky_key you passed to POST /v1/sessions. Release that session or choose a different sticky_key.

ERR::REQUEST::PAYLOAD_TOO_LARGE

413, not retryable, 0 credits. The request body is over the size limit. For POST /v1/batch the limit is 1 MiB (1,048,576 bytes) and 10,000 items per call. Split the job, or create it with open: true and append items with POST /v1/batch/{batchID}/items.

SECURITY: refused before any request left Spicrawl

ERR::SECURITY::SSRF_BLOCKED

400, not retryable, 0 credits. The target url (or your proxy host) resolves to a private, loopback or reserved address. Only publicly routable hosts can be scraped. Use a public hostname; localhost, 10.0.0.0/8 and similar ranges are always refused.

ERR::SECURITY::URL_BLOCKED

400, not retryable, 0 credits. The target URL is blocked by platform policy. Changing parameters will not help. Contact support with the request_id if you believe the block is wrong.

ERR::SECURITY::INVALID_URL

400, not retryable, 0 credits. url is not a valid http or https URL with a host. Include the scheme (https://example.com/products/42, not example.com/products/42) and URL-encode it when you pass it as a query parameter.

SESSION: the named session cannot serve the request

ERR::SESSION::NOT_FOUND is covered with ERR::REQUEST::NOT_FOUND under NOT_FOUND.

ERR::SESSION::EXPIRED

410, not retryable, 0 credits. The session's TTL elapsed and its cookies and storage were purged. There is no recovery. Create a new session with POST /v1/sessions and log in again. To keep state, read GET /v1/sessions/{sessionID}/context before the session expires and store it yourself.

ERR::SESSION::RELEASED

410, not retryable, 0 credits. The session was released (POST /v1/sessions/{sessionID}/release), which purges its context and the browser's copy of its cookies. A /v1/scrape with a released session_id is refused before any work. Create a new session and log in again. (A deleted session is gone entirely and returns ERR::SESSION::NOT_FOUND.)

ERR::SESSION::BUSY

409, retryable, 0 credits. Another request is using the session right now. Wait Retry-After seconds and retry. Send one request at a time per session, or create one session per concurrent worker. To release or delete a busy session anyway, repeat the call with ?force=true.

ERR::SESSION::STATE_CORRUPT

500, not retryable, 0 credits. The stored session context cannot be read and will fail the same way on every attempt. Delete the session, create a new one and log in again. Quote the request_id to support.

LIMIT: over an allowance

ERR::LIMIT::RATE_LIMITED and ERR::PROXY::RATE_LIMITED

Both codes link here. Both are retryable and cost 0 credits, but they come from different places.

ERR::LIMIT::RATE_LIMITED (429): you exceeded your requests-per-second limit. On /v1/scrape the limit is per API key; on /v1/browser it counts session opens per key.

  • Wait Retry-After seconds, then retry. X-RateLimit-Reset carries the same wait.
  • Pace requests from X-RateLimit-Limit and X-RateLimit-Remaining so you do not hit the limit. Cache hits count against it too.
  • See Rate limits.

ERR::PROXY::RATE_LIMITED (502): the proxy exit that carried the request was rate limited. It is not your API limit.

  • Retry with backoff, or use a different proxy.

ERR::LIMIT::CONCURRENCY_EXCEEDED

429, retryable, 0 credits. Every concurrent slot for your organization is in use. The response carries Concurrency-Limit and Concurrency-Remaining: 0.

  • /v1/scrape: Retry-After: 1. Wait for an in-flight request to finish, and cap your client's parallelism at Concurrency-Limit.
  • /v1/browser: the limit is browser_concurrency (default 2) and Retry-After: 5. A slot is held until the WebSocket closes, so close browsers you are done with.

ERR::LIMIT::QUOTA_EXCEEDED

402, not retryable, 0 credits. Your organization's monthly credit allowance is used up. Retrying will fail until the ceiling is raised or the allowance resets. Each org's allowance month runs from the day the org was created and resets at 00:00 UTC on that day every month (the last day of a shorter month for an org created on the 29th, 30th or 31st); the scrape error message names the exact reset date, e.g. "It resets at 00:00 UTC on 2026-10-28."

  • Under mode=auto and in batch jobs, the dearest possible outcome is reserved against the allowance before work starts, so a request can be refused while some allowance remains.
  • On /v1/browser the whole session_ttl price is reserved (held) against the allowance before the session opens; a shorter session_ttl may fit. A held session is charged only for the minutes it actually used, and the unused hold is returned when it closes.

ERR::LIMIT::MAX_COST_EXCEEDED

400, not retryable, 0 credits. The dearest reachable outcome of the request costs more than your max_cost. Nothing ran. Raise max_cost, or choose a cheaper configuration: use js_render=true instead of engine=chromium, or replace mode=auto (checked against its top rung, 25 credits) with an explicit engine. See Credits.

ERR::LIMIT::SESSIONS_EXCEEDED

429, retryable, 0 credits. Your organization already holds 100 live sessions. Retry-After: 30. Release sessions you no longer need (POST /v1/sessions/{sessionID}/release) or wait for a TTL to elapse.

UPSTREAM: the target or the network to it failed

ERR::UPSTREAM::TIMEOUT

504, retryable, 0 credits. The target did not answer within the time budget. Compare diagnostics.elapsed_ms with diagnostics.budget_ms.

  • Without a browser: if the page builds its content with JavaScript, set js_render=true.
  • With a browser: if wait_for never matched, raise wait_for_timeout or fix the selector. If the site gates on anti-bot checks, see Anti-bot.

ERR::UPSTREAM::DNS_FAILED

502, retryable, 0 credits. The target hostname could not be resolved. Check the spelling of the host. Retry once for a transient resolver failure; a hostname that does not exist keeps failing.

ERR::UPSTREAM::TLS_FAILED

502, retryable, 0 credits. The TLS handshake with the target failed. Retry once. If it persists, the site's certificate or TLS setup is broken, or it rejects non-browser TLS: drop impersonate=false if you sent it (impersonate is on by default and free), or try js_render=true.

ERR::UPSTREAM::CONNECTION_RESET

502, retryable, 0 credits. The target closed the connection mid-request. Retry with backoff. Repeated resets from one site often mean it is blocking the exit: route through your own proxy.

ERR::UPSTREAM::TOO_MANY_REDIRECTS

502, not retryable, 0 credits. The target redirected too many times, usually a loop. Scrape the final URL directly. If the loop depends on a cookie, use a session (session_id) so cookies persist.

ERR::UPSTREAM::ERROR and ERR::INTERNAL::ERROR

Both codes link here. Both are retryable and cost 0 credits.

ERR::UPSTREAM::ERROR (502): the request to the target failed for a reason that matches no more specific UPSTREAM code. Read detail and diagnostics.failed_at, then retry with backoff.

ERR::INTERNAL::ERROR (500): a fault inside Spicrawl, not your request. Retry with backoff. If it repeats, quote the request_id to support.

ERR::UPSTREAM::CHALLENGE

502, retryable, 0 credits. An anti-bot vendor (Cloudflare, DataDome, Akamai, PerimeterX) served a challenge page instead of the content. Spicrawl never returns the "Just a moment" page as a success. target_status carries the challenge's status.

  1. Retry once. The verdict is per exit and per moment, and a different exit often passes.
  2. Set js_render=true if the challenge needs JavaScript to run.
  3. Route through your own proxy (for example a residential one). Datacenter IP ranges are challenged on sight. This fixes IP reputation, not rendering.

Stealth mode Coming soon will add a fourth option: rendering in the hardened Camoufox browser. Where Camoufox runs, a request through your own proxy or a premium_proxy Coming soon exit that pins no engine and does not use mode=auto moves to Camoufox by itself when challenged, so this error means every rung was challenged; it still costs 0. A fetch through a managed pool exit is also retried once on another exit first. See Anti-bot.

A bare 403, 429 or 503 from the site with no vendor marker is not a challenge. It is returned as a successful call with X-Target-Status set, at 0 credits.

ERR::UPSTREAM::STATUS

Request log only, 0 credits. No response carries this code. The target answered with a status that is not billable (a 5xx, or a 4xx other than 404/410 that you did not list in allowed_status_codes). You received the body with X-Target-Status set. The request log records the call as failed with error_code UPSTREAM_STATUS, so you can see why it cost nothing.

PROXY: the exit failed

ERR::PROXY::RATE_LIMITED is covered with ERR::LIMIT::RATE_LIMITED under RATE_LIMITED.

ERR::PROXY::EXHAUSTED

503, retryable, 0 credits. No healthy managed-pool exit matched the request. The managed proxy pool Coming soon is not available yet. If you sent a managed-pool flag, drop it and supply your own proxy instead.

ERR::PROXY::UNREACHABLE

502, not retryable, 0 credits. Your own proxy did not accept a connection. Check its host, port and scheme (http, https, socks5, socks5h). Set proxy_verify=true so a dead proxy fails fast before the engine runs.

ERR::PROXY::AUTH_FAILED

502, not retryable, 0 credits. The proxy rejected the credentials. For your own proxy, check the username and password in its userinfo (http://user:pass@proxy.example.net:8080) and URL-encode special characters.

ENGINE: the fetch or render engine failed

ERR::ENGINE::RENDER_FAILED

502, retryable, 0 credits. The browser could not render the page: navigation failed, the engine crashed, a script errored, a screenshot failed, or an actions step failed. diagnostics.failed_at names the stage.

  • For a failed action, check its selector. Set on_error: "skip" on steps that may legitimately find nothing.
  • Retry once for a crash. If it repeats on one site, try engine=chromium.

ERR::ENGINE::UNAVAILABLE and ERR::INTERNAL::UNAVAILABLE

Both codes link here. Both cost 0 credits.

ERR::ENGINE::UNAVAILABLE (503, retryable): no capacity on the engine you need. Back off and retry.

  • POST /v1/sessions returns this for an engine this deployment does not run, and creates nothing. Retrying will not help; the detail lists the engines you can use.
  • On /v1/browser (remote browser Coming soon), 501 with retryable: false means this deployment has no CDP backend. Do not retry.

ERR::INTERNAL::UNAVAILABLE (503, retryable): a Spicrawl dependency is temporarily unavailable. Wait Retry-After (10 on /v1/browser) and retry.

  • ai_extract Coming soon returns this until AI extraction launches (no extraction model is configured). Retrying will not help; use extract with selectors instead.

EXTRACT: extraction failed

ERR::EXTRACT::FAILED

502, not retryable, 0 credits. The document was fetched but could not be extracted. The same document and rules fail the same way again.

  • ai_extract Coming soon: the model failed. Simplify the prompt or schema, or switch to selector extract.
  • parse_pdf=false with a PDF target and response_format markdown or text: set parse_pdf=true or response_format=html for the raw bytes.

ERR::EXTRACT::INVALID_RULES

400, not retryable, 0 credits. The extract rules are invalid, or you set extract_preset. detail names the rule.

  • Fix the selector map or JSON Schema in extract. Each property of a schema needs a selector.
  • Extraction presets (extract_preset) Coming soon: until they launch, any value fails. Put the rules inline in extract.

INTERNAL: a Spicrawl fault

ERR::INTERNAL::ERROR is covered under ERROR and ERR::INTERNAL::UNAVAILABLE under UNAVAILABLE. Both are retryable and never billed.

On this page

Two statusesThe problem bodyRetry decision tableAll codesAUTH: who you are and what the key may doERR::AUTH::MISSING_KEYERR::AUTH::INVALID_KEYERR::AUTH::REVOKED_KEYERR::AUTH::EXPIRED_KEYERR::AUTH::FORBIDDENERR::AUTH::INSUFFICIENT_SCOPEERR::AUTH::ENGINE_NOT_ENTITLEDERR::AUTH::INVALID_TOKENERR::AUTH::INVALID_CREDENTIALSREQUEST: the request itself is wrongERR::REQUEST::INVALIDERR::REQUEST::MISSING_PARAMETERERR::REQUEST::INVALID_PARAMETERERR::REQUEST::INCOMPATIBLE_FLAGSERR::REQUEST::CAPABILITY_UNSUPPORTEDERR::REQUEST::NOT_FOUND and ERR::SESSION::NOT_FOUNDERR::REQUEST::BEYOND_RETENTIONERR::REQUEST::CONFLICTERR::REQUEST::PAYLOAD_TOO_LARGESECURITY: refused before any request left SpicrawlERR::SECURITY::SSRF_BLOCKEDERR::SECURITY::URL_BLOCKEDERR::SECURITY::INVALID_URLSESSION: the named session cannot serve the requestERR::SESSION::EXPIREDERR::SESSION::RELEASEDERR::SESSION::BUSYERR::SESSION::STATE_CORRUPTLIMIT: over an allowanceERR::LIMIT::RATE_LIMITED and ERR::PROXY::RATE_LIMITEDERR::LIMIT::CONCURRENCY_EXCEEDEDERR::LIMIT::QUOTA_EXCEEDEDERR::LIMIT::MAX_COST_EXCEEDEDERR::LIMIT::SESSIONS_EXCEEDEDUPSTREAM: the target or the network to it failedERR::UPSTREAM::TIMEOUTERR::UPSTREAM::DNS_FAILEDERR::UPSTREAM::TLS_FAILEDERR::UPSTREAM::CONNECTION_RESETERR::UPSTREAM::TOO_MANY_REDIRECTSERR::UPSTREAM::ERROR and ERR::INTERNAL::ERRORERR::UPSTREAM::CHALLENGEERR::UPSTREAM::STATUSPROXY: the exit failedERR::PROXY::EXHAUSTEDERR::PROXY::UNREACHABLEERR::PROXY::AUTH_FAILEDENGINE: the fetch or render engine failedERR::ENGINE::RENDER_FAILEDERR::ENGINE::UNAVAILABLE and ERR::INTERNAL::UNAVAILABLEEXTRACT: extraction failedERR::EXTRACT::FAILEDERR::EXTRACT::INVALID_RULESINTERNAL: a Spicrawl fault