spicrawlspicrawlDocs

Capture the page's own API calls

Record the XHR and fetch responses a page makes while it renders with network_capture, and read their JSON bodies from the network array.

Use this when the data you want arrives in the page's own API calls: a product grid loaded from /api/search, prices from a GraphQL endpoint, reviews fetched on scroll. That JSON is usually cleaner and more complete than anything you can pull out of the rendered DOM, and it survives redesigns that break selectors. network_capture records those responses while the browser renders the page and returns them under network. It needs a browser engine and adds no credits.

Minimal request

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/search?q=desk",
    "js_render": true,
    "network_capture": {"urls": ["/api/search"]}
  }'

What comes back

The JSON envelope, with one entry per recorded response under network:

{
  "url": "https://example.com/search?q=desk",
  "status": 200,
  "content": "<html>...</html>",
  "credits": 3,
  "engine": "obscura",
  "proxy_source": "direct",
  "warnings": [],
  "network": [
    {
      "url": "https://example.com/api/search?q=desk&page=1",
      "status": 200,
      "resource_type": "fetch",
      "content_type": "application/json",
      "duration_ms": 142,
      "from_cache": false,
      "body": { "total": 38, "items": [{ "id": 42, "name": "Walnut Desk", "price": 349.0 }] }
    },
    {
      "url": "https://example.com/api/search/facets?q=desk",
      "status": 200,
      "resource_type": "xhr",
      "content_type": "application/json",
      "duration_ms": 97,
      "body_truncated": true,
      "body_base64": "eyJmYWNldHMiOlt7Im5hbWUiOiJicmFuZCIs...",
      "encoding": "base64"
    }
  ]
}

How to read each entry:

  • body holds the parsed JSON, present only when the response was complete, valid JSON with a JSON content type.
  • body_base64 (with encoding: "base64") holds any other body: HTML, text, images, or JSON that was cut off.
  • body_truncated: true means the body hit max_body_bytes. Truncated JSON is always delivered as base64 because it no longer parses. Raise max_body_bytes to get it whole.
  • status is the status of that sub-request, not of the page. The page's own status is the top-level status.

network_capture forces the envelope. If you also set response_format=markdown, content holds the markdown and the response carries X-Warning: FORMAT_COERCED.

Options

All fields of network_capture are optional; {} records every XHR and fetch response.

FieldDefaultLimitsEffect
urls[] (everything)stringsA response is kept when its URL contains one of these substrings. A value starting with ^ is an anchored regular expression, e.g. "^https://api\\.example\\.com/v2/".
resource_types["xhr", "fetch"]document, stylesheet, image, media, font, script, xhr, fetch, websocket, otherTypes to record. Exact singular names; no aliases. An unknown value is 400.
include_bodiestruebooleanfalse returns only url, status, resource_type and timing: far smaller when you only need to discover endpoints.
max_body_bytes1048576 (1 MiB)0–2097152 (2 MiB); 0 means defaultPer-response body cap.
max_responses1000–100; 0 means defaultMaximum responses recorded.

Capturing image or font responses lifts the default image and font blocking for that render, so those requests are actually made.

Find the endpoint first

When you do not know which call carries the data, capture without bodies, then narrow:

{ "url": "https://example.com/search?q=desk", "js_render": true, "network_capture": { "include_bodies": false } }

Pick the URL whose name looks like the data, then rerun with urls set to it and bodies on. Some calls fire only after interaction; add wait_for or browser actions (for example a scroll) to trigger them before capture.

Requires a render

network_capture is browser-only. Set js_render: true or a browser engine (obscura, chromium). On the fetch engine there is no page making requests, and the request fails with 400 ERR::REQUEST::INCOMPATIBLE_FLAGS before anything is spent.

Failure modes

CodeHTTPWhat to do
ERR::REQUEST::INCOMPATIBLE_FLAGS400No browser engine. Add js_render: true.
ERR::REQUEST::INVALID_PARAMETER400An unknown resource_types value, max_body_bytes above 2097152, max_responses above 100, or an unknown key inside network_capture.
ERR::ENGINE::RENDER_FAILED502The browser could not render the page. Retryable.
ERR::UPSTREAM::TIMEOUT504The render ran out of time. Retry, or narrow wait_for.

An empty network array is not an error: the page made no matching calls. Widen urls, add fetch/xhr to resource_types, or trigger the calls with wait_for or actions.

Cost

network_capture adds 0 credits. You pay the render: 3 on obscura, 8 on chromium. Failures cost 0. See Credits.

On this page