# Quickstart

> Get a Spicrawl API key, scrape your first page as markdown with curl, Python, TypeScript or the CLI, and read the response headers that matter.

Source: https://docs.spicrawl.com/quickstart

You need a Spicrawl account and one of: `curl`, Python 3 with `requests`, Node.js 18 or later, or the [`spicrawl` CLI](https://docs.spicrawl.com/cli/install.md). For Node.js there is also a typed client, `npm install @spicrawl/sdk`. The first scrape below costs 1 credit.

**Step 1: Get an API key**

Sign in at [app.spicrawl.com](https://app.spicrawl.com) and open [**API Keys**](https://app.spicrawl.com/dashboard/keys). Every workspace has a live key named **Default** that you can reveal and copy, or you can create a new key. A new key's secret is shown once, so copy it now.

Keys look like `spicrawl_live_…` (spends credits) or `spicrawl_test_…` (can never spend live credits). See [Authentication](https://docs.spicrawl.com/authentication.md).

**Step 2: Set SPICRAWL_API_KEY**

Every example in these docs, the CLI and the MCP server read the key from this variable.

```bash
export SPICRAWL_API_KEY="spicrawl_live_…"
```

Keep the key out of source code and logs. In production, load it from your secret manager.

**Step 3: Scrape a page as markdown**

Send the URL and `response_format: "markdown"`. The response body is the page's main content as markdown.

```bash title="curl"
curl -sS -D - "https://api.spicrawl.com/v1/scrape" \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/products/42", "response_format": "markdown"}'
```

```python title="quickstart.py"
import os

import requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={"url": "https://example.com/products/42", "response_format": "markdown"},
    timeout=120,
)
if not r.ok:
    problem = r.json()
    raise SystemExit(f"{problem['code']}: {problem.get('detail')}")

print("site status:", r.headers.get("X-Target-Status"))
print("credits:", r.headers["X-Credits-Charged"])
print(r.text)
```

```typescript title="quickstart.ts"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url: "https://example.com/products/42", response_format: "markdown" }),
});
if (!r.ok) {
  const problem = await r.json();
  throw new Error(`${problem.code}: ${problem.detail}`);
}

console.log("site status:", r.headers.get("X-Target-Status"));
console.log("credits:", r.headers.get("X-Credits-Charged"));
console.log(await r.text());
```

```bash title="CLI"
# --meta prints engine, credits, cache state and target status to stderr.
spicrawl scrape https://example.com/products/42 --format markdown --meta
```

Run the TypeScript file with `npx tsx quickstart.ts`.

**Step 4: Read the headers that matter**

A successful call looks like this:

```text
HTTP/2 200
content-type: text/markdown; charset=utf-8
x-request-id: 01M0HF5WFWE7PRE8KZHDTNETWN
x-target-status: 200
x-engine: fetch
x-credits-charged: 1
x-request-cost: 1
x-credits-remaining: 999
cache-state: miss
```

| Header                | Check it because                                                                                                                                               |
| --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `X-Target-Status`     | The HTTP `200` is Spicrawl's. This is what the **site** answered. `403` or `503` here means you got a block page, charged 0; see [anti-bot](https://docs.spicrawl.com/guides/anti-bot.md). |
| `X-Credits-Charged`   | What this request billed. `0` on any failure, cache hit, or site status other than `200`, `404`, `410`.                                                        |
| `X-Credits-Remaining` | Credits left in your monthly allowance after this request. Absent if your organization has no monthly limit.                                                   |
| `X-Engine`            | Which engine ran: `fetch` (no JavaScript, 1 credit), `obscura` (`js_render`, 3), `chromium` (8).                                                               |
| `Cache-State`         | `hit` means a stored result was served at 0 credits. `miss` means it was fetched now. `bypass` means the request was not cacheable.                            |
| `X-Request-Id`        | Quote it to support, or pass it to `GET /v1/requests/{id}` for the full trace.                                                                                 |

All headers are listed in [Response headers](https://docs.spicrawl.com/response-headers.md).

## If it failed

Errors are `application/problem+json` with a stable `code`, a `retryable` flag and often a `diagnostics.hint` naming the parameter to change. Failures cost 0 credits.

| Code                                    | Fix                                                                                                                                          |
| --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `ERR::AUTH::MISSING_KEY` (401)          | `SPICRAWL_API_KEY` is empty in this shell, or the header is missing.                                                                         |
| `ERR::REQUEST::INVALID_PARAMETER` (400) | A field is misspelled or unknown. `detail` names it.                                                                                         |
| `ERR::UPSTREAM::CHALLENGE` (502)        | The site served a bot challenge. Retry once, then add `js_render: true` or route through your own `proxy`. See [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md). |
| `ERR::UPSTREAM::TIMEOUT` (504)          | If the page is built by JavaScript, add `js_render: true`.                                                                                   |

Every code is in [Errors](https://docs.spicrawl.com/errors.md).

## Next steps

- [Render JavaScript](https://docs.spicrawl.com/guides/javascript-rendering.md): Add `js_render: true` when the content is built in the browser. 3 credits.

- [Extract data](https://docs.spicrawl.com/guides/structured-data.md): Get JSON with CSS selectors (`extract`, free). Extraction with a model (`ai_extract`, +4) (coming soon).

- [Run a batch](https://docs.spicrawl.com/guides/batch.md): Queue up to 10,000 URLs in one job and page through the results.
