# Scrape job postings from job boards

> Job scraping API cookbook: scrape job boards, infinite-scroll listings and public job APIs with Spicrawl. Tested requests, credit costs and limits.

Source: https://docs.spicrawl.com/use-cases/job-postings

Send each listing URL to `POST /v1/scrape` with the cheapest request that returns real listings: a plain fetch (1 credit) for job APIs and feeds, `js_render: true` (3) for JavaScript pages, and scroll `actions` (8) for infinite-scroll pages. In October 2026 tests, 15 of 16 job sources returned listings; one geo-gated board did not. Failed requests cost 0 credits.

 

## Which request should I use for a job board, and what does it cost?

Start with a plain fetch; if the body has no job cards add `js_render: true`; if it is still a few hundred bytes add scroll `actions`.

* **Public JSON API or RSS feed:** plain fetch with `response_format: "html"`, 1 credit.
* **JavaScript-rendered listing or search page:** `js_render: true`, 3 credits.
* **Infinite-scroll page** (a plain request returns a tiny shell): `js_render: true` plus scroll `actions`, 8 credits.
* **Detail page served as plain HTML:** plain fetch, 1 credit.
* **Board that answers a direct request with 403:** some returned full listings once rendered with `js_render: true`, 3 credits.

10 infinite-scroll pages cost 80 credits; 100 plain-HTML detail pages cost 103 (one 3-credit search plus 100 fetches at 1). Cache hits are billed like the fetch; `max_cost` caps spend.

## How do I scrape and extract job postings?

### How do I scrape JavaScript-rendered job listing pages?

Set `js_render: true`: Spicrawl runs the page in a real browser and returns the rendered job list for 3 credits.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/jobs",
    "js_render": true,
    "cache": false
  }'
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={
        "url": "https://example.com/jobs",
        "js_render": True,
        "cache": False,
    },
    timeout=180,
)
r.raise_for_status()
print(r.headers["X-Credits-Charged"], len(r.content), "bytes")
html = r.text
```

```typescript title="TypeScript"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    url: "https://example.com/jobs",
    js_render: true,
    cache: false,
  }),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
const html = await r.text();
console.log(r.headers.get("X-Credits-Charged"), html.length, "chars");
```

```bash title="CLI"
spicrawl scrape https://example.com/jobs --render --no-cache --meta -o jobs.html
```

The same request works for any JavaScript-rendered job board; use the board's own search URL. Add `"response_format": "markdown"` for LLM-ready output, or `"wait_for": "<selector>"` if the list loads late. See [JavaScript rendering](https://docs.spicrawl.com/guides/javascript-rendering.md).

### How do I scrape infinite-scroll job listings?

Many boards fill in job cards as you scroll, so add two scroll actions: pages that returned a 60 to 113 byte shell without actions returned 31 to 68 KB with them, for 8 credits.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://jobs.example.com/search?q=data+engineer&location=Remote",
    "js_render": true,
    "wait": 2000,
    "actions": [
      {"scroll": {"to_bottom": true}},
      {"scroll": {"to_bottom": true}}
    ]
  }'
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={
        "url": "https://jobs.example.com/search?q=data+engineer&location=Remote",
        "js_render": True,
        "wait": 2000,
        "actions": [
            {"scroll": {"to_bottom": True}},
            {"scroll": {"to_bottom": True}},
        ],
    },
    timeout=180,
)
r.raise_for_status()
if len(r.content) < 2000:  # a tiny shell is still HTTP 200
    raise RuntimeError(f"thin page: {len(r.content)} bytes")
print(r.headers["X-Credits-Charged"], len(r.content), "bytes")
```

```typescript title="TypeScript"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    url: "https://jobs.example.com/search?q=data+engineer&location=Remote",
    js_render: true,
    wait: 2000,
    actions: [{ scroll: { to_bottom: true } }, { scroll: { to_bottom: true } }],
  }),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
const html = await r.text();
if (html.length < 2000) throw new Error(`thin page: ${html.length} chars`); // a shell is still HTTP 200
console.log(r.headers.get("X-Credits-Charged"), html.length, "chars");
```

```bash title="CLI"
spicrawl scrape "https://jobs.example.com/search?q=data+engineer&location=Remote" --render --wait 2000 \
  --actions '[{"scroll":{"to_bottom":true}},{"scroll":{"to_bottom":true}}]' --meta -o jobs.html
```

Any request with `actions` runs on full-browser rendering (8 credits) and is never cached. See [Browser actions](https://docs.spicrawl.com/guides/browser-actions.md).

### Can I scrape job postings without logging in?

Yes, for public listings. A public search page cost 3 credits and a public plain-HTML detail endpoint 1 credit per posting in our tests, with no login, cookies or session.

**Step 1: Collect the job IDs from the search page**

Fetch the search page with `js_render: true` and pull the numeric ID off every posting link.

> **Pattern to adapt:** The search and detail requests were tested on real boards; this ID extraction was not. It assumes links shaped like `/listing/<id>`, so adjust it to the markup you receive.

```bash title="curl"
curl -s https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://jobs.example.com/search?q=data+engineer&location=Remote", "js_render": true, "cache": false}' \
  | grep -oE '/listing/[0-9]+' | grep -oE '[0-9]+$' | sort -u > ids.txt
```

```python title="Python"
import os, re, requests

API = "https://api.spicrawl.com/v1/scrape"
H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}

search = requests.post(API, headers=H, json={
    "url": "https://jobs.example.com/search?q=data+engineer&location=Remote",
    "js_render": True,
    "cache": False,
}, timeout=180)
search.raise_for_status()

# Posting links end in a numeric id: /listing/123
ids = list(dict.fromkeys(re.findall(r"/listing/(\d+)", search.text)))
print(len(ids), ids[:5])
```

```typescript title="TypeScript"
const API = "https://api.spicrawl.com/v1/scrape";
const H = { Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`, "Content-Type": "application/json" };

const search = await fetch(API, {
  method: "POST",
  headers: H,
  body: JSON.stringify({
    url: "https://jobs.example.com/search?q=data+engineer&location=Remote",
    js_render: true,
    cache: false,
  }),
});
if (!search.ok) throw new Error(JSON.stringify(await search.json()));

// Posting links end in a numeric id: /listing/123
const html = await search.text();
const ids = [...new Set([...html.matchAll(/\/listing\/(\d+)/g)].map((m) => m[1]))];
console.log(ids.length, ids.slice(0, 5));
```

```bash title="CLI"
spicrawl scrape "https://jobs.example.com/search?q=data+engineer&location=Remote" --render --no-cache \
  | jq -r .content | grep -oE '/listing/[0-9]+' | grep -oE '[0-9]+$' | sort -u > ids.txt
```

**Step 2: Fetch each posting's detail page**

If the board serves detail pages as plain HTML, request them without `js_render` (1 credit). Check `X-Target-Status`: a site that refuses you still answers HTTP 200 at 0 credits. The 1,000-byte floor is a guess; tune it.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://jobs.example.com/listing/123", "cache": false}'
```

```python title="Python"
import time

posts = {}
for job_id in ids[:25]:
    p = requests.post(API, headers=H, json={
        "url": f"https://jobs.example.com/listing/{job_id}",
        "cache": False,
    }, timeout=60)
    if p.ok and p.headers.get("X-Target-Status") == "200" and len(p.content) > 1000:
        posts[job_id] = p.text
    time.sleep(2)  # pace the requests
print(len(posts), "postings")
```

```typescript title="TypeScript"
const posts: Record<string, string> = {};
for (const id of ids.slice(0, 25)) {
  const p = await fetch(API, {
    method: "POST",
    headers: H,
    body: JSON.stringify({
      url: `https://jobs.example.com/listing/${id}`,
      cache: false,
    }),
  });
  const body = await p.text();
  if (p.ok && p.headers.get("X-Target-Status") === "200" && body.length > 1000) posts[id] = body;
  await new Promise((r) => setTimeout(r, 2000)); // pace the requests
}
console.log(Object.keys(posts).length, "postings");
```

```bash title="CLI"
sed 's#^#https://jobs.example.com/listing/#' ids.txt \
  | spicrawl scrape - --no-cache --concurrency 2 > posts.jsonl
```

### How do I get remote job listings from JSON APIs and RSS feeds?

Send the API or feed URL with `"response_format": "html"`, the raw-bytes format even for JSON and XML: a public jobs API returned 100 items for 1 credit. `markdown` and `text` fail on JSON or RSS with `ERR::EXTRACT::FAILED` (HTTP 502, 0 credits).

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://api.example.com/jobs", "response_format": "html", "cache": false}'
```

```python title="Python"
import json, os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={"url": "https://api.example.com/jobs", "response_format": "html", "cache": False},
    timeout=60,
)
r.raise_for_status()
jobs = json.loads(r.text)  # field names depend on the API
for job in jobs[:3]:
    print(job["title"], "|", job["company"], "|", job["location"], "|", job["url"])
```

```typescript title="TypeScript"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url: "https://api.example.com/jobs", response_format: "html", cache: false }),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
const jobs = JSON.parse(await r.text()); // field names depend on the API
for (const j of jobs.slice(0, 3)) console.log(j.title, "|", j.company, "|", j.location, "|", j.url);
```

```bash title="CLI"
spicrawl scrape https://api.example.com/jobs --format html --no-cache -o jobs.json
```

RSS and Atom feeds work the same way: swap the `url` and parse the XML. Public endpoints can also be called directly for free. Check the API's terms first, as some ask you to credit them as the source.

### How do I extract title, company, location and salary from a job post?

Add `"autoparse": true` to get the page's schema.org `JobPosting` block under `data.json_ld` at no extra credits. Not every board embeds it, and salary is often missing.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/jobs/42", "js_render": true, "autoparse": true}'
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={"url": "https://example.com/jobs/42", "js_render": True, "autoparse": True},
    timeout=180,
)
r.raise_for_status()
env = r.json()  # autoparse returns the JSON envelope


def job_postings(env):
    for block in env["data"].get("json_ld", []):
        if block.get("@type") != "JobPosting":
            continue
        place = block.get("jobLocation") or {}
        place = place[0] if isinstance(place, list) and place else place
        pay = (block.get("baseSalary") or {}).get("value") or {}
        yield {
            "title": block.get("title"),
            "company": (block.get("hiringOrganization") or {}).get("name"),
            "location": (place.get("address") or {}).get("addressLocality"),
            "salary_min": pay.get("minValue") if isinstance(pay, dict) else None,
            "url": block.get("url") or env["final_url"],
        }


print(list(job_postings(env)))
```

```typescript title="TypeScript"
const r = await fetch("https://api.spicrawl.com/v1/scrape", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url: "https://example.com/jobs/42", js_render: true, autoparse: true }),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
const env = await r.json(); // autoparse returns the JSON envelope
const postings = (env.data.json_ld ?? []).filter((b: { "@type"?: string }) => b["@type"] === "JobPosting");
console.log(postings);
```

```bash title="CLI"
spicrawl scrape https://example.com/jobs/42 --render --autoparse | jq '.data.json_ld'
```

Without embedded data, name the fields with CSS selectors (the ones below are placeholders). `empty_fields` lists every selector that matched nothing, so alert on it: it is the first sign a board changed its markup. See [Structured data](https://docs.spicrawl.com/guides/structured-data.md).

```json
{
  "url": "https://example.com/jobs/42",
  "js_render": true,
  "extract": {
    "type": "object",
    "properties": {
      "title":    { "type": "string", "selector": "h1" },
      "company":  { "type": "string", "selector": ".company-name" },
      "location": { "type": "string", "selector": ".job-location" },
      "salary":   { "type": "string", "selector": ".salary" },
      "url":      { "type": "string", "selector": "a.apply", "attribute": "href" }
    },
    "required": ["title", "company"]
  }
}
```

## How do I scale up, and what are the limits?

### How do I scrape thousands of job listings for a job aggregator?

`POST /v1/batch` takes up to 10,000 URLs per job, runs them server-side and keeps JSON Lines results for 72 hours. An aggregator batches the detail URLs, de-duplicates on job ID and re-scrapes listing pages with `cache: false`. Batch items are plain fetches or renders only: `actions`, `extract`, `autoparse` and `session_id` are refused with a `400`, so loop over `/v1/scrape` for the infinite-scroll recipe.

This job fetches detail pages three at a time and keeps each job ID as `external_id`:

```bash title="curl"
curl https://api.spicrawl.com/v1/batch \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "job-postings-2026-10",
    "concurrency": 3,
    "items": [
      {"url": "https://jobs.example.com/listing/123", "external_id": "123"},
      {"url": "https://jobs.example.com/listing/124", "external_id": "124"}
    ]
  }'
```

```python title="Python"
import json, os, time, requests

API = "https://api.spicrawl.com"
H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}

job = requests.post(f"{API}/v1/batch", headers=H, json={
    "name": "job-postings-2026-10",
    "concurrency": 3,
    "items": [
        {"url": f"https://jobs.example.com/listing/{i}", "external_id": i}
        for i in ids
    ],
}, timeout=60).json()

while job["status"] not in ("completed", "failed", "cancelled"):
    time.sleep(10)
    job = requests.get(f"{API}/v1/batch/{job['id']}", headers=H, timeout=30).json()

postings, cursor = {}, None
while True:
    r = requests.get(f"{API}/v1/batch/{job['id']}/results", headers=H,
                     params={"cursor": cursor} if cursor else {}, timeout=120)
    r.raise_for_status()
    for line in r.text.splitlines():
        item = json.loads(line)
        if item["status"] == "succeeded" and "content" in item.get("result", {}):
            postings[item["external_id"]] = json.loads(item["result"]["content"])  # content is a JSON string literal
    cursor = r.headers.get("X-Next-Cursor")
    if not cursor:
        break
print(len(postings), "postings")
```

```typescript title="TypeScript"
const API = "https://api.spicrawl.com";
const H = { Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`, "Content-Type": "application/json" };

let job = await (await fetch(`${API}/v1/batch`, {
  method: "POST",
  headers: H,
  body: JSON.stringify({
    name: "job-postings-2026-10",
    concurrency: 3,
    items: ids.map((id) => ({
      url: `https://jobs.example.com/listing/${id}`,
      external_id: id,
    })),
  }),
})).json();

while (!["completed", "failed", "cancelled"].includes(job.status)) {
  await new Promise((r) => setTimeout(r, 10000));
  job = await (await fetch(`${API}/v1/batch/${job.id}`, { headers: H })).json();
}

const postings: Record<string, string> = {};
let cursor: string | null = null;
do {
  const r = await fetch(`${API}/v1/batch/${job.id}/results${cursor ? `?cursor=${cursor}` : ""}`, { headers: H });
  for (const line of (await r.text()).split("\n").filter(Boolean)) {
    const item = JSON.parse(line);
    if (item.status === "succeeded" && item.result?.content) postings[item.external_id] = JSON.parse(item.result.content);
  }
  cursor = r.headers.get("X-Next-Cursor");
} while (cursor);
console.log(Object.keys(postings).length, "postings");
```

```bash title="CLI"
jq -Rc '{url: ("https://jobs.example.com/listing/" + .), external_id: .}' ids.txt > items.jsonl
spicrawl batch submit items.jsonl --name job-postings-2026-10 --concurrency 3 --wait --max-wait 20m
spicrawl batch results <job-id> --all -o results.jsonl
```

Submitting is free but holds the dearest possible cost of every item against your monthly allowance; each item is charged on success only. `concurrency` (default 10, maximum 50) is also how you pace a site. For infinite-scroll pages, loop over `/v1/scrape` (`scroll.json` holds the two scroll steps):

```bash
spicrawl scrape - --render --wait 2000 --actions @scroll.json --concurrency 2 < urls.txt > pages.jsonl
```

See [Batch jobs](https://docs.spicrawl.com/guides/batch.md).

### How do I keep job data fresh and complete?

* **Fresh:** send `"cache": false`. Results are cached for 48 hours by default, and a hit is billed like the fetch, so skipping the cache costs nothing extra. See [Caching](https://docs.spicrawl.com/guides/caching.md).
* **Location:** the tested sources worked on the default exit. For location-correct results, pass your own `proxy` and `proxy_country`, and keep one exit while paging a result set. See [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md).
* **Thin pages:** an empty shell comes back as HTTP 200 (60 and 113 bytes in our tests, against 14 to 68 KB for real pages). Check the body size, `X-Target-Status` and `X-Warning` before you store it, then retry with scroll `actions`, a `wait` or your own proxy.
* **Pacing:** leave 3 to 5 seconds between pages of one site (our conservative default, not a site limit), use `--concurrency 2` or a batch `concurrency` of 2 to 3, and honour `Retry-After` on a `429`. See [Rate limits](https://docs.spicrawl.com/rate-limits.md).

### What can't Spicrawl scrape yet?

These are point-in-time observations; job boards change their markup and bot protection without notice.

* **Geo-gated boards:** one returned an empty 403 and needs an exit in its own country, which we have not tested.
* **Login-only pages**, such as company reviews, returned only a login-wall shell. A login-once `session_id` is the documented pattern but unverified on job boards; see [Sessions and logins](https://docs.spicrawl.com/guides/sessions-and-logins.md), and use it only with an account you may use under the site's terms.
* **DataDome** may still block. On DataDome-protected non-job pages the request returned no content at 0 credits; try your own residential `proxy` and see [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md).
* **Deep listings** on category overview pages may need pagination beyond the first page we tested.

## Frequently asked questions

**How do I scrape job postings from job boards?**

Send each listing URL to `POST /v1/scrape` with the cheapest request that returns real listings: a plain fetch (1 credit) for JSON and RSS feeds, `js_render: true` (3 credits) for JavaScript pages, and scroll `actions` (8 credits) for infinite-scroll pages. In October 2026 tests, 15 of 16 job sources returned listings.

**Can I scrape job postings without logging in?**

Yes, for public listings. In October 2026 tests Spicrawl scraped a public job search page for 3 credits and a public plain-HTML detail endpoint for 1 credit per posting, with no login, cookies or session.

**Why does a job board return an empty page when I scrape it?**

Many job boards fill in their job cards as you scroll, so a request without scroll actions can return a tiny shell (60 to 113 bytes in our tests) with HTTP 200. Adding `js_render: true` and two `{"scroll": {"to_bottom": true}}` actions returned the full listings for 8 credits.

**How much does it cost to scrape 100 job postings?**

Between 1 and 103 credits, depending on the source: one public API request that returns about 100 jobs costs 1 credit, while 100 detail pages from a plain-HTML endpoint cost 103 (one 3-credit search page plus 100 fetches at 1).

**Can I scrape a public job API or RSS feed?**

Yes. Public JSON APIs and RSS feeds returned structured jobs in October 2026 tests. Fetch them through Spicrawl at 1 credit per request with `response_format: "html"` to get the raw JSON or XML.

**Can I scrape job pages that need a login, such as company reviews?**

Public listings yes, login-only pages not without a login. In October 2026 tests, company review pages returned only a login-wall shell, while job listings on the same board rendered for 3 credits.

## Related

* [JavaScript rendering](https://docs.spicrawl.com/guides/javascript-rendering.md), [Browser actions](https://docs.spicrawl.com/guides/browser-actions.md) and [Structured data](https://docs.spicrawl.com/guides/structured-data.md)
* [Batch jobs](https://docs.spicrawl.com/guides/batch.md), [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md) and [Sessions and logins](https://docs.spicrawl.com/guides/sessions-and-logins.md)
* [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md), [Caching](https://docs.spicrawl.com/guides/caching.md) and [Credits](https://docs.spicrawl.com/credits.md)

 

## Disclaimer

> **For educational purposes:** The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and `robots.txt`, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.
