Lead generation and market research from the web
Scrape company directories, startup lists and listing sites into structured leads (name, website, location, size), in batches, exported to CSV or a CRM.
To generate B2B leads from the web, collect profile URLs from a company directory with links: true, call POST /v1/scrape with extract on each profile to get the company name, website, location and size as JSON, then write the rows to CSV or your CRM. A profile costs 1 credit (3 with JavaScript rendering), extract adds nothing, and a failed request costs 0.
How do you collect profile URLs from a directory or startup list?
Scrape the listing with links: true and keep the profile links. Every selector on this page is a placeholder for example.com: use the classes your target really has.
curl -s https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/directory/robotics?page=1",
"js_render": true,
"wait_for": ".company-card",
"links": true
}' | jq -r '.links[] | select(test("/companies/"))' | sort -u > profile-urls.txtFor ?page=2 style lists, generate the URLs and stop when a page returns no profile links. For "Load more" or infinite scroll, use browser actions (full-browser rendering, 8 credits per listing page):
{
"url": "https://example.com/directory/robotics",
"js_render": true,
"links": true,
"actions": [
{"scroll": {"to_bottom": true}, "timeout_ms": 20000},
{"wait_for": ".company-card"}
]
}How do you extract name, website, location and size from a profile page?
Send the profile URL with an extract selector map. Each key becomes a field under data, and @href returns an attribute instead of text.
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/companies/acme-robotics",
"js_render": true,
"wait_for": "h1",
"max_cost": 3,
"extract": {
"name": "h1",
"website": "a.company-website@href",
"location": ".company-location",
"size": ".company-size"
}
}'{
"url": "https://example.com/companies/acme-robotics",
"final_url": "https://example.com/companies/acme-robotics",
"status": 200,
"credits": 3,
"data": {
"name": "Acme Robotics",
"website": "https://acmerobotics.example",
"location": "Pittsburgh, PA, USA",
"size": "11-50 employees"
},
"empty_fields": []
}status is the directory's status. Keep a row only when it is 200 and your required fields are present, and alert when a field goes empty on every profile: the site changed its markup. Resolve relative @href values against final_url.
| Tier | Request | Use it when |
|---|---|---|
autoparse | "autoparse": true | The profile publishes schema.org Organization markup; data.json_ld holds it, no selectors needed. |
| Selector map | "extract": {"name": "h1"} | You know the markup, as above. |
| JSON Schema | "extract": {"type": "object", "properties": {...}} | You want typed, validated values; failures land in empty_fields. |
All three add 0 credits. See Structured data.
How do you scrape thousands of companies?
Stream the profile URLs through the CLI, which supports extract. A batch job runs server-side but returns pages, not fields, so parse its HTML yourself.
echo '{"name":"h1","website":"a.company-website@href","location":".company-location","size":".company-size"}' > fields.json
spicrawl scrape - --render --wait-for h1 --max-cost 3 --concurrency 4 --retry 3 \
--extract @fields.json < profile-urls.txt > companies.jsonl
# Failures and why, to retry later
jq -r 'select(.error) | [.url, .error.code] | @tsv' companies.jsonlKeep --concurrency at or below your Concurrency-Limit (Rate limits); beyond it you get a retryable 429 that costs 0. For larger unattended runs, use a batch job:
id=$(spicrawl batch submit profile-urls.txt --render --max-cost 3 --credit-budget 1500 | jq -r .id)
spicrawl batch wait "$id"
spicrawl batch results "$id" --status succeeded --all | jq -r '.result.content | fromjson' > profiles.htmlA batch submission reserves the dearest possible cost of every item against your monthly allowance and fails with 402 ERR::LIMIT::QUOTA_EXCEEDED before anything runs if that does not fit. Estimated costs, from the price table:
| Run | Credits |
|---|---|
| 500 profiles, plain fetch (1) | 500 |
| 500 profiles, JavaScript rendering (3) | 1,500 |
| 20 scroll-action listing pages (8) plus 500 JavaScript-rendered profiles (3) | 1,660 |
What if the directory blocks scrapers or shows an empty page?
Climb one step at a time; each costs 0 until a request returns the real page.
| Symptom | What to do |
|---|---|
status 200 but every field is in empty_fields | The page is built by JavaScript or your selectors are wrong: add js_render: true and a wait_for on an element only the finished page has. |
| A 60 to 110 byte page | A single-page app returned its shell: add actions that scroll and wait, then check the length of content. |
X-Target-Status: 403 with HTTP 200 and 0 credits | The site refused you: try js_render: true. |
ERR::UPSTREAM::CHALLENGE (502, retryable) | A bot challenge: retry with your own residential proxy in proxy (8 credits, about 18 to 20 seconds in our tests). See Cloudflare. |
| A login wall | Use sessions only for accounts you are entitled to use. |
Some sites stay out of reach: in our October 2026 tests, listing pages behind DataDome stayed on the interstitial at 0 credits. Look for the owner's export or API instead. See Anti-bot and Proxies and geo.
How do you export leads to CSV or a CRM?
Flatten each result to a row, drop duplicates by website domain, then write one CSV or send one record per company to your CRM. Most CRMs import a CSV, which is the quickest first load.
import csv, os, requests
from urllib.parse import urlparse
H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
FIELDS = {"name": "h1", "website": "a.company-website@href", "location": ".company-location", "size": ".company-size"}
profiles = [u.strip() for u in open("profile-urls.txt") if u.strip()]
def domain(site):
host = urlparse(site or "").hostname or ""
return host.removeprefix("www.") # Python 3.9+
unique, failed = {}, []
for url in profiles: # sequential; add a thread pool up to your Concurrency-Limit for speed
r = requests.post("https://api.spicrawl.com/v1/scrape", headers=H, timeout=180,
json={"url": url, "js_render": True, "wait_for": "h1", "max_cost": 3, "extract": FIELDS})
if r.status_code != 200:
failed.append((url, r.json().get("code")))
continue
env = r.json()
data = env.get("data") or {}
if env["status"] != 200 or not data.get("name"):
failed.append((url, "EMPTY"))
continue
unique[domain(data.get("website")) or url] = {**data, "domain": domain(data.get("website")), "source_url": url}
with open("leads.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=["name", "website", "domain", "location", "size", "source_url"], extrasaction="ignore")
w.writeheader()
w.writerows(unique.values())
print(len(unique), "leads,", len(failed), "failed")
# Optional: push each lead to your CRM's REST API (placeholder URL and token).
for lead in unique.values():
res = requests.post("https://crm.example.com/api/companies", timeout=30,
headers={"Authorization": f"Bearer {os.environ['CRM_API_TOKEN']}"},
json={k: lead.get(k) for k in ("name", "domain", "location", "size")})
if not res.ok:
print("CRM rejected", lead["domain"], res.status_code)For market research, aggregate the rows instead of contacting them, for example a market map by location: jq -r 'select(.data) | .data.location' companies.jsonl | sort | uniq -c | sort -rn | head.
FAQ
Some can be. In an October 2026 test, a startup job board that returned HTTP 403 to a direct request returned its listings through a js_render Spicrawl request for 3 credits. Company profile pages were not tested, so try one before a full run. Check each site's robots.txt and terms first.
Collect profile URLs with links: true, then request each one with extract to get the name, website, location and size as JSON. Try autoparse first when the profile publishes schema.org Organization markup. Spicrawl has no site-crawl or search endpoint, so you supply the listing URLs.
No. A batch item returns html, markdown, text or json, and a submission that sets extract, autoparse or links is refused with a 400 at 0 credits. Run extract per URL, or fetch with a batch job and parse the HTML yourself.
Write the rows to CSV and use your CRM's CSV import, or post each company to its API from the same script. Deduplicate on the website domain first so a re-run updates records instead of copying them.
Yes, with your own residential proxy in proxy: Spicrawl moves a challenged request to full-browser rendering on the same exit and bills 8 credits. Success is likely, not guaranteed.
No. Spicrawl fetches the URLs you send, so checking robots.txt, the site's terms and your request rate is your responsibility. Python's urllib.robotparser can filter a URL list.
Related: Structured data · Batch jobs · Browser actions · Cloudflare · Anti-bot · Proxies and geo · Credits · Rate limits
Start scraping in minutes
1,000 free credits every month. One API key, one request.
For educational purposes
The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and robots.txt, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.
AI knowledge base
Scrape websites and docs sites to clean markdown for RAG and AI agents: batch ingestion, chunking, the MCP server, llms.txt and scheduled refreshes.
Price monitoring
Price monitoring API cookbook: extract product title, price, availability and rating, re-scrape with cache false, batch catalogs, compare country prices.