spicrawlspicrawlDocs

Stay logged in with sessions

Create a session with POST /v1/sessions, pass its id as session_id on /v1/scrape to reuse cookies, storage and a pinned exit IP, and release it when you are done.

Use this when a site needs state across requests: you log in once and then read ten account pages, or a site sets a consent cookie you do not want to fight every time. A session is a sealed cookie jar and web-storage snapshot, pinned to one engine and, by default, one exit IP. You create it, pass its id as session_id on /v1/scrape, and cookies the page sets are written back to it after each render.

Requires an API key with the sessions scope to create and manage sessions, and scrape to use one.

Minimal flow

Create a session

curl https://api.spicrawl.com/v1/sessions \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"engine": "chromium", "ttl_seconds": 7200}'

The 201 response is the session object: id (a 26-character ULID), engine, status: "active", proxy, expires_at, hard_expires_at and a context size summary. It never contains cookies.

Log in with actions

{
  "url": "https://example.com/login",
  "session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD",
  "js_render": true,
  "actions": [
    {"fill": {"selector": "input[name=email]", "value": "ada@example.com"}},
    {"fill": {"selector": "input[name=password]", "value": "correct-horse-battery", "secret": true}},
    {"click": "button[type=submit]"},
    {"wait_for": ".account-menu"}
  ]
}

The session's engine is used automatically. Still set js_render: true when you send browser-only fields such as actions or wait_for; session_id alone does not satisfy that check. See Browser actions.

Scrape pages as the logged-in user

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/account/orders", "session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD", "js_render": true, "response_format": "markdown"}'

Release it

curl -X POST https://api.spicrawl.com/v1/sessions/01J9ZQ4M7R3T8VX2K5N6P0B1CD/release \
  -H "Authorization: Bearer $SPICRAWL_API_KEY"
# CLI: spicrawl sessions release 01J9ZQ4M7R3T8VX2K5N6P0B1CD

Rules that trip people up

One request at a time per session. A session is single-writer. While a render holds its lease, a second scrape with the same session_id gets 409 ERR::SESSION::BUSY with Retry-After. Serialise requests per session, or create one session per concurrent worker.

Managed exits for sessions. Coming soon Pinning a session to a Spicrawl-managed exit (sticky_key, rotate_ip) and session proxy tiers (premium_proxy, proxy_country, region_pool) are not available during the beta. Many sites challenge a login whose IP moves, so if you route through your own proxy, use one with a fixed exit.

Release destroys the state. POST /v1/sessions/{id}/release marks the session released and deletes its cookies and storage in the same transaction, along with the browser's own copy. It is not "give back the lease, keep the login". A later /v1/scrape with that session_id is refused with 410 ERR::SESSION::RELEASED at 0 credits. The session row stays readable with GET /v1/sessions/{id}; the context does not. DELETE /v1/sessions/{id} also purges the context and removes the row, so the id stops resolving (a 204 is returned even for an unknown id). Both answer 409 ERR::SESSION::BUSY while a render runs; add force=true to end the session anyway and fail the running task.

Sessions are never cached. A request with session_id bypasses the result cache (Cache-State: bypass) and is always fetched fresh.

The engine is fixed. Scrapes on the session run on its engine. Pinning a different engine on /v1/scrape is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. When you omit engine at creation, the default is the first of obscura, chromium that the deployment runs, else fetch; check engine in the response. An engine the deployment does not run is refused at creation with 503 ERR::ENGINE::UNAVAILABLE.

An unknown session_id is an error. /v1/scrape checks the session before doing anything. An unknown, malformed or deleted id, or another organization's, is 404 ERR::SESSION::NOT_FOUND at 0 credits. The request never runs without the session.

Use is counted. Each successful scrape adds 1 to usage_count, sets last_used_at and moves expires_at to that time plus ttl_seconds, never past hard_expires_at. Failed scrapes and /context reads do not count.

Save and restore a login

GET /v1/sessions/{id}/context is the only endpoint that returns the cookies, local storage, session storage and IndexedDB. It works only while the session is active, does not slide the TTL, and is sent with Cache-Control: no-store, private.

spicrawl sessions context 01J9ZQ4M7R3T8VX2K5N6P0B1CD > login.json

Store that object yourself before you release the session. To clone the login later, pass its session_context to POST /v1/sessions. The serialised context is limited to 1 MiB (413 ERR::REQUEST::PAYLOAD_TOO_LARGE).

Create options

FieldDefaultEffect
enginefirst of obscura, chromium deployedfetch, obscura or chromium. Fixed for the session's life.
ttl_seconds1800Sliding lifetime, minimum 30. Each successful scrape pushes expires_at forward, never past hard_expires_at. Above the organization maximum (604800, 7 days, by default) is a 400 that names the real maximum.
session_contextemptyInitial cookies and storage, e.g. from a previous /context call.

Failure modes

CodeHTTPWhat to do
ERR::SESSION::BUSY409Another render holds the session. Wait Retry-After (or until lease.expires_at), then retry.
ERR::SESSION::EXPIRED410The TTL passed and the context was purged. Do not retry: create a new session and log in again.
ERR::SESSION::RELEASED410The session was released and its context purged. Create a new one.
ERR::SESSION::NOT_FOUND404Unknown, deleted, or malformed id, or another organization's session.
ERR::ENGINE::UNAVAILABLE503On create: this deployment does not run that engine. Pick one from the detail.
ERR::LIMIT::SESSIONS_EXCEEDED429The organization already holds 100 live sessions. Release sessions you no longer need; Retry-After is 30.
ERR::SESSION::STATE_CORRUPT500The stored context cannot be read. It will not recover; create a new session.

A session a worker has taken out of service shows a retired field (expired, usage_exhausted or explicit); replace it.

Cost

Creating, reading and releasing sessions costs nothing. Each scrape on a session is billed like any other scrape at the session engine's price: fetch 1, obscura 3, chromium 8 credits. Failures, including a 409 or 410, cost 0. See Credits.

On this page