Stay logged in with sessions
Create a session with POST /v1/sessions, pass its id as session_id on /v1/scrape to reuse cookies, storage and a pinned exit IP, and release it when you are done.
Use this when a site needs state across requests: you log in once and then read ten account pages, or a site sets a consent cookie you do not want to fight every time. A session is a sealed cookie jar and web-storage snapshot, pinned to one engine and, by default, one exit IP. You create it, pass its id as session_id on /v1/scrape, and cookies the page sets are written back to it after each render.
Requires an API key with the sessions scope to create and manage sessions, and scrape to use one.
Minimal flow
Create a session
curl https://api.spicrawl.com/v1/sessions \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"engine": "chromium", "ttl_seconds": 7200}'The 201 response is the session object: id (a 26-character ULID), engine, status: "active", proxy, expires_at, hard_expires_at and a context size summary. It never contains cookies.
Log in with actions
{
"url": "https://example.com/login",
"session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD",
"js_render": true,
"actions": [
{"fill": {"selector": "input[name=email]", "value": "ada@example.com"}},
{"fill": {"selector": "input[name=password]", "value": "correct-horse-battery", "secret": true}},
{"click": "button[type=submit]"},
{"wait_for": ".account-menu"}
]
}The session's engine is used automatically. Still set js_render: true when you send browser-only fields such as actions or wait_for; session_id alone does not satisfy that check. See Browser actions.
Scrape pages as the logged-in user
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/account/orders", "session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD", "js_render": true, "response_format": "markdown"}'Release it
curl -X POST https://api.spicrawl.com/v1/sessions/01J9ZQ4M7R3T8VX2K5N6P0B1CD/release \
-H "Authorization: Bearer $SPICRAWL_API_KEY"
# CLI: spicrawl sessions release 01J9ZQ4M7R3T8VX2K5N6P0B1CDRules that trip people up
One request at a time per session. A session is single-writer. While a render holds its lease, a second scrape with the same session_id gets 409 ERR::SESSION::BUSY with Retry-After. Serialise requests per session, or create one session per concurrent worker.
Managed exits for sessions. Coming soon Pinning a session to a Spicrawl-managed exit (sticky_key, rotate_ip) and session proxy tiers (premium_proxy, proxy_country, region_pool) are not available during the beta. Many sites challenge a login whose IP moves, so if you route through your own proxy, use one with a fixed exit.
Release destroys the state. POST /v1/sessions/{id}/release marks the session released and deletes its cookies and storage in the same transaction, along with the browser's own copy. It is not "give back the lease, keep the login". A later /v1/scrape with that session_id is refused with 410 ERR::SESSION::RELEASED at 0 credits. The session row stays readable with GET /v1/sessions/{id}; the context does not. DELETE /v1/sessions/{id} also purges the context and removes the row, so the id stops resolving (a 204 is returned even for an unknown id). Both answer 409 ERR::SESSION::BUSY while a render runs; add force=true to end the session anyway and fail the running task.
Sessions are never cached. A request with session_id bypasses the result cache (Cache-State: bypass) and is always fetched fresh.
The engine is fixed. Scrapes on the session run on its engine. Pinning a different engine on /v1/scrape is 400 ERR::REQUEST::INCOMPATIBLE_FLAGS. When you omit engine at creation, the default is the first of obscura, chromium that the deployment runs, else fetch; check engine in the response. An engine the deployment does not run is refused at creation with 503 ERR::ENGINE::UNAVAILABLE.
An unknown session_id is an error. /v1/scrape checks the session before doing anything. An unknown, malformed or deleted id, or another organization's, is 404 ERR::SESSION::NOT_FOUND at 0 credits. The request never runs without the session.
Use is counted. Each successful scrape adds 1 to usage_count, sets last_used_at and moves expires_at to that time plus ttl_seconds, never past hard_expires_at. Failed scrapes and /context reads do not count.
Save and restore a login
GET /v1/sessions/{id}/context is the only endpoint that returns the cookies, local storage, session storage and IndexedDB. It works only while the session is active, does not slide the TTL, and is sent with Cache-Control: no-store, private.
spicrawl sessions context 01J9ZQ4M7R3T8VX2K5N6P0B1CD > login.jsonStore that object yourself before you release the session. To clone the login later, pass its session_context to POST /v1/sessions. The serialised context is limited to 1 MiB (413 ERR::REQUEST::PAYLOAD_TOO_LARGE).
Create options
| Field | Default | Effect |
|---|---|---|
engine | first of obscura, chromium deployed | fetch, obscura or chromium. Fixed for the session's life. |
ttl_seconds | 1800 | Sliding lifetime, minimum 30. Each successful scrape pushes expires_at forward, never past hard_expires_at. Above the organization maximum (604800, 7 days, by default) is a 400 that names the real maximum. |
session_context | empty | Initial cookies and storage, e.g. from a previous /context call. |
Failure modes
| Code | HTTP | What to do |
|---|---|---|
ERR::SESSION::BUSY | 409 | Another render holds the session. Wait Retry-After (or until lease.expires_at), then retry. |
ERR::SESSION::EXPIRED | 410 | The TTL passed and the context was purged. Do not retry: create a new session and log in again. |
ERR::SESSION::RELEASED | 410 | The session was released and its context purged. Create a new one. |
ERR::SESSION::NOT_FOUND | 404 | Unknown, deleted, or malformed id, or another organization's session. |
ERR::ENGINE::UNAVAILABLE | 503 | On create: this deployment does not run that engine. Pick one from the detail. |
ERR::LIMIT::SESSIONS_EXCEEDED | 429 | The organization already holds 100 live sessions. Release sessions you no longer need; Retry-After is 30. |
ERR::SESSION::STATE_CORRUPT | 500 | The stored context cannot be read. It will not recover; create a new session. |
A session a worker has taken out of service shows a retired field (expired, usage_exhausted or explicit); replace it.
Cost
Creating, reading and releasing sessions costs nothing. Each scrape on a session is billed like any other scrape at the session engine's price: fetch 1, obscura 3, chromium 8 credits. Failures, including a 409 or 410, cost 0. See Credits.
Related
- Browser actions for the login steps
- Proxies and geo for using your own proxy
- CDP browser Coming soon
- Caching
Browser actions
Drive a rendered page with the actions array: wait_for, wait_for_navigation, click, fill, select, scroll, evaluate and screenshot steps, up to 50 per request, before the page is captured.
Batch jobs
Submit up to 10,000 URLs per call to POST /v1/batch, poll the job until it is terminal, and page the JSONL results before they expire after 72 hours.