spicrawlspicrawlDocs

Scrape job postings from job boards

Job scraping API cookbook: scrape job boards, infinite-scroll listings and public job APIs with Spicrawl. Tested requests, credit costs and limits.

Send each listing URL to POST /v1/scrape with the cheapest request that returns real listings: a plain fetch (1 credit) for job APIs and feeds, js_render: true (3) for JavaScript pages, and scroll actions (8) for infinite-scroll pages. In October 2026 tests, 15 of 16 job sources returned listings; one geo-gated board did not. Failed requests cost 0 credits.

Which request should I use for a job board, and what does it cost?

Start with a plain fetch; if the body has no job cards add js_render: true; if it is still a few hundred bytes add scroll actions.

  • Public JSON API or RSS feed: plain fetch with response_format: "html", 1 credit.
  • JavaScript-rendered listing or search page: js_render: true, 3 credits.
  • Infinite-scroll page (a plain request returns a tiny shell): js_render: true plus scroll actions, 8 credits.
  • Detail page served as plain HTML: plain fetch, 1 credit.
  • Board that answers a direct request with 403: some returned full listings once rendered with js_render: true, 3 credits.

10 infinite-scroll pages cost 80 credits; 100 plain-HTML detail pages cost 103 (one 3-credit search plus 100 fetches at 1). Cache hits are billed like the fetch; max_cost caps spend.

How do I scrape and extract job postings?

How do I scrape JavaScript-rendered job listing pages?

Set js_render: true: Spicrawl runs the page in a real browser and returns the rendered job list for 3 credits.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/jobs",
    "js_render": true,
    "cache": false
  }'

The same request works for any JavaScript-rendered job board; use the board's own search URL. Add "response_format": "markdown" for LLM-ready output, or "wait_for": "<selector>" if the list loads late. See JavaScript rendering.

How do I scrape infinite-scroll job listings?

Many boards fill in job cards as you scroll, so add two scroll actions: pages that returned a 60 to 113 byte shell without actions returned 31 to 68 KB with them, for 8 credits.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://jobs.example.com/search?q=data+engineer&location=Remote",
    "js_render": true,
    "wait": 2000,
    "actions": [
      {"scroll": {"to_bottom": true}},
      {"scroll": {"to_bottom": true}}
    ]
  }'

Any request with actions runs on full-browser rendering (8 credits) and is never cached. See Browser actions.

Can I scrape job postings without logging in?

Yes, for public listings. A public search page cost 3 credits and a public plain-HTML detail endpoint 1 credit per posting in our tests, with no login, cookies or session.

Collect the job IDs from the search page

Fetch the search page with js_render: true and pull the numeric ID off every posting link.

Pattern to adapt

The search and detail requests were tested on real boards; this ID extraction was not. It assumes links shaped like /listing/<id>, so adjust it to the markup you receive.

curl -s https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://jobs.example.com/search?q=data+engineer&location=Remote", "js_render": true, "cache": false}' \
  | grep -oE '/listing/[0-9]+' | grep -oE '[0-9]+$' | sort -u > ids.txt

Fetch each posting's detail page

If the board serves detail pages as plain HTML, request them without js_render (1 credit). Check X-Target-Status: a site that refuses you still answers HTTP 200 at 0 credits. The 1,000-byte floor is a guess; tune it.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://jobs.example.com/listing/123", "cache": false}'

How do I get remote job listings from JSON APIs and RSS feeds?

Send the API or feed URL with "response_format": "html", the raw-bytes format even for JSON and XML: a public jobs API returned 100 items for 1 credit. markdown and text fail on JSON or RSS with ERR::EXTRACT::FAILED (HTTP 502, 0 credits).

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://api.example.com/jobs", "response_format": "html", "cache": false}'

RSS and Atom feeds work the same way: swap the url and parse the XML. Public endpoints can also be called directly for free. Check the API's terms first, as some ask you to credit them as the source.

How do I extract title, company, location and salary from a job post?

Add "autoparse": true to get the page's schema.org JobPosting block under data.json_ld at no extra credits. Not every board embeds it, and salary is often missing.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/jobs/42", "js_render": true, "autoparse": true}'

Without embedded data, name the fields with CSS selectors (the ones below are placeholders). empty_fields lists every selector that matched nothing, so alert on it: it is the first sign a board changed its markup. See Structured data.

{
  "url": "https://example.com/jobs/42",
  "js_render": true,
  "extract": {
    "type": "object",
    "properties": {
      "title":    { "type": "string", "selector": "h1" },
      "company":  { "type": "string", "selector": ".company-name" },
      "location": { "type": "string", "selector": ".job-location" },
      "salary":   { "type": "string", "selector": ".salary" },
      "url":      { "type": "string", "selector": "a.apply", "attribute": "href" }
    },
    "required": ["title", "company"]
  }
}

How do I scale up, and what are the limits?

How do I scrape thousands of job listings for a job aggregator?

POST /v1/batch takes up to 10,000 URLs per job, runs them server-side and keeps JSON Lines results for 72 hours. An aggregator batches the detail URLs, de-duplicates on job ID and re-scrapes listing pages with cache: false. Batch items are plain fetches or renders only: actions, extract, autoparse and session_id are refused with a 400, so loop over /v1/scrape for the infinite-scroll recipe.

This job fetches detail pages three at a time and keeps each job ID as external_id:

curl https://api.spicrawl.com/v1/batch \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "job-postings-2026-10",
    "concurrency": 3,
    "items": [
      {"url": "https://jobs.example.com/listing/123", "external_id": "123"},
      {"url": "https://jobs.example.com/listing/124", "external_id": "124"}
    ]
  }'

Submitting is free but holds the dearest possible cost of every item against your monthly allowance; each item is charged on success only. concurrency (default 10, maximum 50) is also how you pace a site. For infinite-scroll pages, loop over /v1/scrape (scroll.json holds the two scroll steps):

spicrawl scrape - --render --wait 2000 --actions @scroll.json --concurrency 2 < urls.txt > pages.jsonl

See Batch jobs.

How do I keep job data fresh and complete?

  • Fresh: send "cache": false. Results are cached for 48 hours by default, and a hit is billed like the fetch, so skipping the cache costs nothing extra. See Caching.
  • Location: the tested sources worked on the default exit. For location-correct results, pass your own proxy and proxy_country, and keep one exit while paging a result set. See Proxies and geo.
  • Thin pages: an empty shell comes back as HTTP 200 (60 and 113 bytes in our tests, against 14 to 68 KB for real pages). Check the body size, X-Target-Status and X-Warning before you store it, then retry with scroll actions, a wait or your own proxy.
  • Pacing: leave 3 to 5 seconds between pages of one site (our conservative default, not a site limit), use --concurrency 2 or a batch concurrency of 2 to 3, and honour Retry-After on a 429. See Rate limits.

What can't Spicrawl scrape yet?

These are point-in-time observations; job boards change their markup and bot protection without notice.

  • Geo-gated boards: one returned an empty 403 and needs an exit in its own country, which we have not tested.
  • Login-only pages, such as company reviews, returned only a login-wall shell. A login-once session_id is the documented pattern but unverified on job boards; see Sessions and logins, and use it only with an account you may use under the site's terms.
  • DataDome may still block. On DataDome-protected non-job pages the request returned no content at 0 credits; try your own residential proxy and see Anti-bot.
  • Deep listings on category overview pages may need pagination beyond the first page we tested.

Frequently asked questions

Start scraping in minutes

1,000 free credits every month. One API key, one request.

Disclaimer

For educational purposes

The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and robots.txt, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.

On this page