# Stay logged in with sessions

> Create a session with POST /v1/sessions, pass its id as session_id on /v1/scrape to reuse cookies, storage and a pinned exit IP, and release it when you are done.

Source: https://docs.spicrawl.com/guides/sessions-and-logins

Use this when a site needs state across requests: you log in once and then read ten account pages, or a site sets a consent cookie you do not want to fight every time. A session is a sealed cookie jar and web-storage snapshot, pinned to one engine and, by default, one exit IP. You create it, pass its `id` as `session_id` on `/v1/scrape`, and cookies the page sets are written back to it after each render.

Requires an API key with the `sessions` scope to create and manage sessions, and `scrape` to use one.

## Minimal flow

**Step 1: Create a session**

```bash title="curl"
curl https://api.spicrawl.com/v1/sessions \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"engine": "chromium", "ttl_seconds": 7200}'
```

```python title="Python"
import os, requests

API = "https://api.spicrawl.com"
H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}

s = requests.post(f"{API}/v1/sessions", headers=H,
                  json={"engine": "chromium", "ttl_seconds": 7200}, timeout=30)
s.raise_for_status()
session_id = s.json()["id"]
```

```typescript title="TypeScript"
const API = "https://api.spicrawl.com";
const H = { Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`, "Content-Type": "application/json" };

const s = await fetch(`${API}/v1/sessions`, {
  method: "POST",
  headers: H,
  body: JSON.stringify({ engine: "chromium", ttl_seconds: 7200 }),
});
if (!s.ok) throw new Error(JSON.stringify(await s.json()));
const { id: sessionId } = await s.json();
```

```bash title="CLI"
spicrawl sessions create --engine chromium --ttl 7200
```

The `201` response is the session object: `id` (a 26-character ULID), `engine`, `status: "active"`, `proxy`, `expires_at`, `hard_expires_at` and a `context` size summary. It never contains cookies.

**Step 2: Log in with actions**

```json
{
  "url": "https://example.com/login",
  "session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD",
  "js_render": true,
  "actions": [
    {"fill": {"selector": "input[name=email]", "value": "ada@example.com"}},
    {"fill": {"selector": "input[name=password]", "value": "correct-horse-battery", "secret": true}},
    {"click": "button[type=submit]"},
    {"wait_for": ".account-menu"}
  ]
}
```

The session's engine is used automatically. Still set `js_render: true` when you send browser-only fields such as `actions` or `wait_for`; `session_id` alone does not satisfy that check. See [Browser actions](https://docs.spicrawl.com/guides/browser-actions.md).

**Step 3: Scrape pages as the logged-in user**

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/account/orders", "session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD", "js_render": true, "response_format": "markdown"}'
```

```python title="Python"
r = requests.post(f"{API}/v1/scrape", headers=H, json={
    "url": "https://example.com/account/orders",
    "session_id": session_id, "js_render": True, "response_format": "markdown",
}, timeout=120)
```

```typescript title="TypeScript"
const r = await fetch(`${API}/v1/scrape`, {
  method: "POST",
  headers: H,
  body: JSON.stringify({
    url: "https://example.com/account/orders",
    session_id: sessionId,
    js_render: true,
    response_format: "markdown",
  }),
});
```

```bash title="CLI"
spicrawl scrape https://example.com/account/orders --session 01J9ZQ4M7R3T8VX2K5N6P0B1CD --render --format markdown
```

**Step 4: Release it**

```bash
curl -X POST https://api.spicrawl.com/v1/sessions/01J9ZQ4M7R3T8VX2K5N6P0B1CD/release \
  -H "Authorization: Bearer $SPICRAWL_API_KEY"
# CLI: spicrawl sessions release 01J9ZQ4M7R3T8VX2K5N6P0B1CD
```

## Rules that trip people up

**One request at a time per session.** A session is single-writer. While a render holds its lease, a second scrape with the same `session_id` gets `409 ERR::SESSION::BUSY` with `Retry-After`. Serialise requests per session, or create one session per concurrent worker.

**Managed exits for sessions.** (coming soon) Pinning a session to a Spicrawl-managed exit (`sticky_key`, `rotate_ip`) and session proxy tiers (`premium_proxy`, `proxy_country`, `region_pool`) are not available during the beta. Many sites challenge a login whose IP moves, so if you route through your own proxy, use one with a fixed exit.

**Release destroys the state.** `POST /v1/sessions/{id}/release` marks the session `released` and deletes its cookies and storage in the same transaction, along with the browser's own copy. It is not "give back the lease, keep the login". A later `/v1/scrape` with that `session_id` is refused with `410 ERR::SESSION::RELEASED` at 0 credits. The session row stays readable with `GET /v1/sessions/{id}`; the context does not. `DELETE /v1/sessions/{id}` also purges the context and removes the row, so the id stops resolving (a 204 is returned even for an unknown id). Both answer `409 ERR::SESSION::BUSY` while a render runs; add `force=true` to end the session anyway and fail the running task.

**Sessions are never cached.** A request with `session_id` bypasses the result cache (`Cache-State: bypass`) and is always fetched fresh.

**The engine is fixed.** Scrapes on the session run on its engine. Pinning a different `engine` on `/v1/scrape` is `400 ERR::REQUEST::INCOMPATIBLE_FLAGS`. When you omit `engine` at creation, the default is the first of `obscura`, `chromium` that the deployment runs, else `fetch`; check `engine` in the response. An `engine` the deployment does not run is refused at creation with `503 ERR::ENGINE::UNAVAILABLE`.

**An unknown `session_id` is an error.** `/v1/scrape` checks the session before doing anything. An unknown, malformed or deleted id, or another organization's, is `404 ERR::SESSION::NOT_FOUND` at 0 credits. The request never runs without the session.

**Use is counted.** Each successful scrape adds 1 to `usage_count`, sets `last_used_at` and moves `expires_at` to that time plus `ttl_seconds`, never past `hard_expires_at`. Failed scrapes and `/context` reads do not count.

## Save and restore a login

`GET /v1/sessions/{id}/context` is the only endpoint that returns the cookies, local storage, session storage and IndexedDB. It works only while the session is `active`, does not slide the TTL, and is sent with `Cache-Control: no-store, private`.

```bash
spicrawl sessions context 01J9ZQ4M7R3T8VX2K5N6P0B1CD > login.json
```

Store that object yourself before you release the session. To clone the login later, pass its `session_context` to `POST /v1/sessions`. The serialised context is limited to 1 MiB (`413 ERR::REQUEST::PAYLOAD_TOO_LARGE`).

## Create options

| Field             | Default                                 | Effect                                                                                                                                                                                                            |
| ----------------- | --------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `engine`          | first of `obscura`, `chromium` deployed | `fetch`, `obscura` or `chromium`. Fixed for the session's life.                                                                                                                                                   |
| `ttl_seconds`     | `1800`                                  | Sliding lifetime, minimum 30. Each successful scrape pushes `expires_at` forward, never past `hard_expires_at`. Above the organization maximum (604800, 7 days, by default) is a 400 that names the real maximum. |
| `session_context` | empty                                   | Initial cookies and storage, e.g. from a previous `/context` call.                                                                                                                                                |

## Failure modes

| Code                            | HTTP | What to do                                                                                                  |
| ------------------------------- | ---- | ----------------------------------------------------------------------------------------------------------- |
| `ERR::SESSION::BUSY`            | 409  | Another render holds the session. Wait `Retry-After` (or until `lease.expires_at`), then retry.             |
| `ERR::SESSION::EXPIRED`         | 410  | The TTL passed and the context was purged. Do not retry: create a new session and log in again.             |
| `ERR::SESSION::RELEASED`        | 410  | The session was released and its context purged. Create a new one.                                          |
| `ERR::SESSION::NOT_FOUND`       | 404  | Unknown, deleted, or malformed id, or another organization's session.                                       |
| `ERR::ENGINE::UNAVAILABLE`      | 503  | On create: this deployment does not run that `engine`. Pick one from the detail.                            |
| `ERR::LIMIT::SESSIONS_EXCEEDED` | 429  | The organization already holds 100 live sessions. Release sessions you no longer need; `Retry-After` is 30. |
| `ERR::SESSION::STATE_CORRUPT`   | 500  | The stored context cannot be read. It will not recover; create a new session.                               |

A session a worker has taken out of service shows a `retired` field (`expired`, `usage_exhausted` or `explicit`); replace it.

## Cost

Creating, reading and releasing sessions costs nothing. Each scrape on a session is billed like any other scrape at the session engine's price: `fetch` 1, `obscura` 3, `chromium` 8 credits. Failures, including a `409` or `410`, cost 0. See [Credits](https://docs.spicrawl.com/credits.md).

## Related

* [Browser actions](https://docs.spicrawl.com/guides/browser-actions.md) for the login steps
* [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md) for using your own proxy
* [CDP browser](https://docs.spicrawl.com/guides/cdp-browser.md) (coming soon)
* [Caching](https://docs.spicrawl.com/guides/caching.md)
