Sessions and cloud browsers
Keep cookies and engine across scrapes with spicrawl sessions. Cloud browsers (spicrawl browser url) are coming soon.
A session keeps one engine and one cookie jar across many spicrawl scrape calls, so a login or a cart survives between requests. Cloud browsers Coming soon, a live Chromium you drive yourself over the Chrome DevTools Protocol (CDP), are not available yet.
SID=$(spicrawl sessions create --engine chromium | jq -r .id)
spicrawl scrape https://example.com/account --session "$SID" --format markdown
spicrawl sessions release "$SID"Sessions
spicrawl sessions create
Creates a session (POST /v1/sessions). Only flags you set are sent; the API applies its defaults (engine: the first of obscura, chromium available; TTL 1800 s).
| Flag | API field | Meaning |
|---|---|---|
--engine E | engine | fetch, obscura or chromium. Any other value exits 2. |
--ttl S | ttl_seconds | Sliding lifetime in seconds, minimum 30, default 1800. Each use extends it. |
--session-context FILE | session_context | Seed the session from a saved context: a JSON object with cookies, local_storage, session_storage, indexed_db, or the output of spicrawl sessions context ID --json (its session_context member is sent). - reads stdin. A missing or non-object file is exit 2. |
# A chromium session for two hours
spicrawl sessions create --engine chromium --ttl 7200
# Import a saved login instead of logging in again
spicrawl sessions context 01J9ZQ4M7R3T8VX2K5N6P0B1CD --json > login.json
spicrawl sessions create --engine chromium --session-context login.jsonJSON mode prints the API's Session object unchanged (id, engine, status, proxy, created_at, expires_at, hard_expires_at, usage_count, ...). Human mode prints key/value lines and use it with: spicrawl scrape <url> --session <id> on stderr. Creation warnings go to stderr as warning: ....
Managed exit flags for sessions Coming soon: --country, --premium-proxy, --region-pool, --rotate-ip and --sticky-key.
Pass the id to any scrape with --session. The session's engine is used; pinning a different --engine on the scrape fails with ERR::REQUEST::INCOMPATIBLE_FLAGS. Requests with a session are never cached.
Other session commands
| Command | API | Does |
|---|---|---|
sessions list | GET /v1/sessions | Newest first. --status active|released|expired, --engine E, --limit N (1-200, default 50), --all to follow next_cursor. |
sessions get <id> | GET /v1/sessions/{id} | One session's metadata. |
sessions context <id> | GET /v1/sessions/{id}/context | The live cookies and storage, as JSON. |
sessions release <id> | POST /v1/sessions/{id}/release | Ends the session and purges its cookies and storage. It cannot be revived; the metadata stays readable. --force ends it even while a render holds its lease (that render fails). |
sessions delete <id> | DELETE /v1/sessions/{id} | Deletes the session and everything stored with it. Needs --yes without a terminal. --force as for release. |
spicrawl sessions list --status active --engine chromium --all
spicrawl sessions list --json | jq -r '.sessions[].id'
spicrawl sessions context 01J9ZQ4M7R3T8VX2K5N6P0B1CD > ctx.json
spicrawl sessions delete 01J9ZQ4M7R3T8VX2K5N6P0B1CD --yessessions list filters --status after the API reads a page, so a page can be short or empty and still have more behind it. Use --all when you need every match. JSON output is {"sessions": [...], "next_cursor": "..."}, with next_cursor only when more pages remain.
sessions context output is a credential: it can replay a login. Treat the file like a password. Pass it back as session_context when creating a session through the API to carry a login over.
sessions delete on a terminal asks Delete session <id> and its cookies and storage? [y/N]. Without a terminal and without --yes it refuses with exit code 2 and sends nothing.
Session errors
| Code | Exit | Do next |
|---|---|---|
ERR::SESSION::NOT_FOUND | 9 | Wrong id or deleted; create a new session. |
ERR::SESSION::EXPIRED | 9 | The TTL ran out; create a new session. |
ERR::SESSION::RELEASED | 9 | Released sessions cannot be reused; create a new one. |
ERR::SESSION::BUSY | 9 | Another render holds the session; retry after it finishes, or run requests on one session serially. |
ERR::LIMIT::SESSIONS_EXCEEDED | 5 | Release sessions you no longer need. |
See sessions and logins for login flows.
Cloud browsers Coming soon
Coming soon
spicrawl browser url will mint a connect URL for a browser hosted by Spicrawl, so you can drive it with Puppeteer or Playwright. Remote browsers are not available yet. See CDP browser.