spicrawlspicrawlDocs

Lead generation and market research from the web

Scrape company directories, startup lists and listing sites into structured leads (name, website, location, size), in batches, exported to CSV or a CRM.

To generate B2B leads from the web, collect profile URLs from a company directory with links: true, call POST /v1/scrape with extract on each profile to get the company name, website, location and size as JSON, then write the rows to CSV or your CRM. A profile costs 1 credit (3 with JavaScript rendering), extract adds nothing, and a failed request costs 0.

How do you collect profile URLs from a directory or startup list?

Scrape the listing with links: true and keep the profile links. Every selector on this page is a placeholder for example.com: use the classes your target really has.

curl -s https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/directory/robotics?page=1",
    "js_render": true,
    "wait_for": ".company-card",
    "links": true
  }' | jq -r '.links[] | select(test("/companies/"))' | sort -u > profile-urls.txt

For ?page=2 style lists, generate the URLs and stop when a page returns no profile links. For "Load more" or infinite scroll, use browser actions (full-browser rendering, 8 credits per listing page):

{
  "url": "https://example.com/directory/robotics",
  "js_render": true,
  "links": true,
  "actions": [
    {"scroll": {"to_bottom": true}, "timeout_ms": 20000},
    {"wait_for": ".company-card"}
  ]
}

How do you extract name, website, location and size from a profile page?

Send the profile URL with an extract selector map. Each key becomes a field under data, and @href returns an attribute instead of text.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/companies/acme-robotics",
    "js_render": true,
    "wait_for": "h1",
    "max_cost": 3,
    "extract": {
      "name": "h1",
      "website": "a.company-website@href",
      "location": ".company-location",
      "size": ".company-size"
    }
  }'
{
  "url": "https://example.com/companies/acme-robotics",
  "final_url": "https://example.com/companies/acme-robotics",
  "status": 200,
  "credits": 3,
  "data": {
    "name": "Acme Robotics",
    "website": "https://acmerobotics.example",
    "location": "Pittsburgh, PA, USA",
    "size": "11-50 employees"
  },
  "empty_fields": []
}

status is the directory's status. Keep a row only when it is 200 and your required fields are present, and alert when a field goes empty on every profile: the site changed its markup. Resolve relative @href values against final_url.

TierRequestUse it when
autoparse"autoparse": trueThe profile publishes schema.org Organization markup; data.json_ld holds it, no selectors needed.
Selector map"extract": {"name": "h1"}You know the markup, as above.
JSON Schema"extract": {"type": "object", "properties": {...}}You want typed, validated values; failures land in empty_fields.

All three add 0 credits. See Structured data.

How do you scrape thousands of companies?

Stream the profile URLs through the CLI, which supports extract. A batch job runs server-side but returns pages, not fields, so parse its HTML yourself.

echo '{"name":"h1","website":"a.company-website@href","location":".company-location","size":".company-size"}' > fields.json

spicrawl scrape - --render --wait-for h1 --max-cost 3 --concurrency 4 --retry 3 \
  --extract @fields.json < profile-urls.txt > companies.jsonl

# Failures and why, to retry later
jq -r 'select(.error) | [.url, .error.code] | @tsv' companies.jsonl

Keep --concurrency at or below your Concurrency-Limit (Rate limits); beyond it you get a retryable 429 that costs 0. For larger unattended runs, use a batch job:

id=$(spicrawl batch submit profile-urls.txt --render --max-cost 3 --credit-budget 1500 | jq -r .id)
spicrawl batch wait "$id"
spicrawl batch results "$id" --status succeeded --all | jq -r '.result.content | fromjson' > profiles.html

A batch submission reserves the dearest possible cost of every item against your monthly allowance and fails with 402 ERR::LIMIT::QUOTA_EXCEEDED before anything runs if that does not fit. Estimated costs, from the price table:

RunCredits
500 profiles, plain fetch (1)500
500 profiles, JavaScript rendering (3)1,500
20 scroll-action listing pages (8) plus 500 JavaScript-rendered profiles (3)1,660

What if the directory blocks scrapers or shows an empty page?

Climb one step at a time; each costs 0 until a request returns the real page.

SymptomWhat to do
status 200 but every field is in empty_fieldsThe page is built by JavaScript or your selectors are wrong: add js_render: true and a wait_for on an element only the finished page has.
A 60 to 110 byte pageA single-page app returned its shell: add actions that scroll and wait, then check the length of content.
X-Target-Status: 403 with HTTP 200 and 0 creditsThe site refused you: try js_render: true.
ERR::UPSTREAM::CHALLENGE (502, retryable)A bot challenge: retry with your own residential proxy in proxy (8 credits, about 18 to 20 seconds in our tests). See Cloudflare.
A login wallUse sessions only for accounts you are entitled to use.

Some sites stay out of reach: in our October 2026 tests, listing pages behind DataDome stayed on the interstitial at 0 credits. Look for the owner's export or API instead. See Anti-bot and Proxies and geo.

How do you export leads to CSV or a CRM?

Flatten each result to a row, drop duplicates by website domain, then write one CSV or send one record per company to your CRM. Most CRMs import a CSV, which is the quickest first load.

import csv, os, requests
from urllib.parse import urlparse

H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
FIELDS = {"name": "h1", "website": "a.company-website@href", "location": ".company-location", "size": ".company-size"}
profiles = [u.strip() for u in open("profile-urls.txt") if u.strip()]

def domain(site):
    host = urlparse(site or "").hostname or ""
    return host.removeprefix("www.")  # Python 3.9+

unique, failed = {}, []
for url in profiles:  # sequential; add a thread pool up to your Concurrency-Limit for speed
    r = requests.post("https://api.spicrawl.com/v1/scrape", headers=H, timeout=180,
                      json={"url": url, "js_render": True, "wait_for": "h1", "max_cost": 3, "extract": FIELDS})
    if r.status_code != 200:
        failed.append((url, r.json().get("code")))
        continue
    env = r.json()
    data = env.get("data") or {}
    if env["status"] != 200 or not data.get("name"):
        failed.append((url, "EMPTY"))
        continue
    unique[domain(data.get("website")) or url] = {**data, "domain": domain(data.get("website")), "source_url": url}

with open("leads.csv", "w", newline="") as f:
    w = csv.DictWriter(f, fieldnames=["name", "website", "domain", "location", "size", "source_url"], extrasaction="ignore")
    w.writeheader()
    w.writerows(unique.values())
print(len(unique), "leads,", len(failed), "failed")

# Optional: push each lead to your CRM's REST API (placeholder URL and token).
for lead in unique.values():
    res = requests.post("https://crm.example.com/api/companies", timeout=30,
                        headers={"Authorization": f"Bearer {os.environ['CRM_API_TOKEN']}"},
                        json={k: lead.get(k) for k in ("name", "domain", "location", "size")})
    if not res.ok:
        print("CRM rejected", lead["domain"], res.status_code)

For market research, aggregate the rows instead of contacting them, for example a market map by location: jq -r 'select(.data) | .data.location' companies.jsonl | sort | uniq -c | sort -rn | head.

FAQ

Related: Structured data · Batch jobs · Browser actions · Cloudflare · Anti-bot · Proxies and geo · Credits · Rate limits

Start scraping in minutes

1,000 free credits every month. One API key, one request.

For educational purposes

The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and robots.txt, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.

On this page