# Lead generation and market research from the web

> Scrape company directories, startup lists and listing sites into structured leads (name, website, location, size), in batches, exported to CSV or a CRM.

Source: https://docs.spicrawl.com/use-cases/lead-generation

To generate B2B leads from the web, collect profile URLs from a company directory with `links: true`, call `POST /v1/scrape` with `extract` on each profile to get the company name, website, location and size as JSON, then write the rows to CSV or your CRM. A profile costs 1 credit (3 with JavaScript rendering), `extract` adds nothing, and a failed request costs 0.

 

## How do you collect profile URLs from a directory or startup list?

Scrape the listing with `links: true` and keep the profile links. Every selector on this page is a placeholder for `example.com`: use the classes your target really has.

```bash title="curl"
curl -s https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/directory/robotics?page=1",
    "js_render": true,
    "wait_for": ".company-card",
    "links": true
  }' | jq -r '.links[] | select(test("/companies/"))' | sort -u > profile-urls.txt
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={
        "url": "https://example.com/directory/robotics?page=1",
        "js_render": True,
        "wait_for": ".company-card",
        "links": True,
    },
    timeout=180,
)
r.raise_for_status()
profiles = sorted({u for u in r.json().get("links", []) if "/companies/" in u})
```

```typescript title="TypeScript"
import { Spicrawl } from "@spicrawl/sdk";

const spicrawl = new Spicrawl(); // reads SPICRAWL_API_KEY

const listing = await spicrawl.scrape({
  url: "https://example.com/directory/robotics?page=1",
  js_render: true,
  wait_for: ".company-card",
  links: true,
});
const profiles = [...new Set((listing.links ?? []).filter((u) => u.includes("/companies/")))];
```

```bash title="CLI"
spicrawl scrape "https://example.com/directory/robotics?page=1" --render --wait-for .company-card --links \
  | jq -r '.links[] | select(test("/companies/"))' | sort -u > profile-urls.txt
```

For `?page=2` style lists, generate the URLs and stop when a page returns no profile links. For "Load more" or infinite scroll, use [browser actions](https://docs.spicrawl.com/guides/browser-actions.md) (full-browser rendering, 8 credits per listing page):

```json
{
  "url": "https://example.com/directory/robotics",
  "js_render": true,
  "links": true,
  "actions": [
    {"scroll": {"to_bottom": true}, "timeout_ms": 20000},
    {"wait_for": ".company-card"}
  ]
}
```

## How do you extract name, website, location and size from a profile page?

Send the profile URL with an `extract` selector map. Each key becomes a field under `data`, and `@href` returns an attribute instead of text.

```bash title="curl"
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/companies/acme-robotics",
    "js_render": true,
    "wait_for": "h1",
    "max_cost": 3,
    "extract": {
      "name": "h1",
      "website": "a.company-website@href",
      "location": ".company-location",
      "size": ".company-size"
    }
  }'
```

```python title="Python"
import os, requests

r = requests.post(
    "https://api.spicrawl.com/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
    json={
        "url": "https://example.com/companies/acme-robotics",
        "js_render": True,
        "wait_for": "h1",
        "max_cost": 3,
        "extract": {
            "name": "h1",
            "website": "a.company-website@href",
            "location": ".company-location",
            "size": ".company-size",
        },
    },
    timeout=180,
)
r.raise_for_status()
env = r.json()
print(env["data"], env.get("empty_fields"))
```

```typescript title="TypeScript"
import { Spicrawl } from "@spicrawl/sdk";

const spicrawl = new Spicrawl(); // reads SPICRAWL_API_KEY

const page = await spicrawl.scrape({
  url: "https://example.com/companies/acme-robotics",
  js_render: true,
  wait_for: "h1",
  max_cost: 3,
  extract: {
    name: "h1",
    website: "a.company-website@href",
    location: ".company-location",
    size: ".company-size",
  },
});
console.log(page.data, page.empty_fields);
```

```bash title="CLI"
spicrawl scrape https://example.com/companies/acme-robotics --render --wait-for h1 --max-cost 3 \
  --extract '{"name":"h1","website":"a.company-website@href","location":".company-location","size":".company-size"}'
```

```json
{
  "url": "https://example.com/companies/acme-robotics",
  "final_url": "https://example.com/companies/acme-robotics",
  "status": 200,
  "credits": 3,
  "data": {
    "name": "Acme Robotics",
    "website": "https://acmerobotics.example",
    "location": "Pittsburgh, PA, USA",
    "size": "11-50 employees"
  },
  "empty_fields": []
}
```

`status` is the directory's status. Keep a row only when it is 200 and your required fields are present, and alert when a field goes empty on every profile: the site changed its markup. Resolve relative `@href` values against `final_url`.

| Tier         | Request                                              | Use it when                                                                                           |
| ------------ | ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------- |
| `autoparse`  | `"autoparse": true`                                  | The profile publishes schema.org `Organization` markup; `data.json_ld` holds it, no selectors needed. |
| Selector map | `"extract": {"name": "h1"}`                          | You know the markup, as above.                                                                        |
| JSON Schema  | `"extract": {"type": "object", "properties": {...}}` | You want typed, validated values; failures land in `empty_fields`.                                    |

All three add 0 credits. See [Structured data](https://docs.spicrawl.com/guides/structured-data.md).

## How do you scrape thousands of companies?

Stream the profile URLs through the CLI, which supports `extract`. A batch job runs server-side but returns pages, not fields, so parse its HTML yourself.

```bash
echo '{"name":"h1","website":"a.company-website@href","location":".company-location","size":".company-size"}' > fields.json

spicrawl scrape - --render --wait-for h1 --max-cost 3 --concurrency 4 --retry 3 \
  --extract @fields.json < profile-urls.txt > companies.jsonl

# Failures and why, to retry later
jq -r 'select(.error) | [.url, .error.code] | @tsv' companies.jsonl
```

Keep `--concurrency` at or below your `Concurrency-Limit` ([Rate limits](https://docs.spicrawl.com/rate-limits.md)); beyond it you get a retryable `429` that costs 0. For larger unattended runs, use a [batch job](https://docs.spicrawl.com/guides/batch.md):

```bash
id=$(spicrawl batch submit profile-urls.txt --render --max-cost 3 --credit-budget 1500 | jq -r .id)
spicrawl batch wait "$id"
spicrawl batch results "$id" --status succeeded --all | jq -r '.result.content | fromjson' > profiles.html
```

A batch submission reserves the dearest possible cost of every item against your monthly allowance and fails with `402 ERR::LIMIT::QUOTA_EXCEEDED` before anything runs if that does not fit. Estimated costs, from the [price table](https://docs.spicrawl.com/credits.md#price-table):

| Run                                                                          | Credits |
| ---------------------------------------------------------------------------- | ------- |
| 500 profiles, plain fetch (1)                                                | 500     |
| 500 profiles, JavaScript rendering (3)                                       | 1,500   |
| 20 scroll-action listing pages (8) plus 500 JavaScript-rendered profiles (3) | 1,660   |

## What if the directory blocks scrapers or shows an empty page?

Climb one step at a time; each costs 0 until a request returns the real page.

| Symptom                                            | What to do                                                                                                                                                |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `status` 200 but every field is in `empty_fields`  | The page is built by JavaScript or your selectors are wrong: add `js_render: true` and a `wait_for` on an element only the finished page has.             |
| A 60 to 110 byte page                              | A single-page app returned its shell: add `actions` that scroll and wait, then check the length of `content`.                                             |
| `X-Target-Status: 403` with HTTP 200 and 0 credits | The site refused you: try `js_render: true`.                                                                                                              |
| `ERR::UPSTREAM::CHALLENGE` (502, retryable)        | A bot challenge: retry with your own residential proxy in `proxy` (8 credits, about 18 to 20 seconds in our tests). See [Cloudflare](https://docs.spicrawl.com/guides/cloudflare.md). |
| A login wall                                       | Use [sessions](https://docs.spicrawl.com/guides/sessions-and-logins.md) only for accounts you are entitled to use.                                                                    |

Some sites stay out of reach: in our October 2026 tests, listing pages behind DataDome stayed on the interstitial at 0 credits. Look for the owner's export or API instead. See [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md) and [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md).

## How do you export leads to CSV or a CRM?

Flatten each result to a row, drop duplicates by website domain, then write one CSV or send one record per company to your CRM. Most CRMs import a CSV, which is the quickest first load.

```python
import csv, os, requests
from urllib.parse import urlparse

H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
FIELDS = {"name": "h1", "website": "a.company-website@href", "location": ".company-location", "size": ".company-size"}
profiles = [u.strip() for u in open("profile-urls.txt") if u.strip()]

def domain(site):
    host = urlparse(site or "").hostname or ""
    return host.removeprefix("www.")  # Python 3.9+

unique, failed = {}, []
for url in profiles:  # sequential; add a thread pool up to your Concurrency-Limit for speed
    r = requests.post("https://api.spicrawl.com/v1/scrape", headers=H, timeout=180,
                      json={"url": url, "js_render": True, "wait_for": "h1", "max_cost": 3, "extract": FIELDS})
    if r.status_code != 200:
        failed.append((url, r.json().get("code")))
        continue
    env = r.json()
    data = env.get("data") or {}
    if env["status"] != 200 or not data.get("name"):
        failed.append((url, "EMPTY"))
        continue
    unique[domain(data.get("website")) or url] = {**data, "domain": domain(data.get("website")), "source_url": url}

with open("leads.csv", "w", newline="") as f:
    w = csv.DictWriter(f, fieldnames=["name", "website", "domain", "location", "size", "source_url"], extrasaction="ignore")
    w.writeheader()
    w.writerows(unique.values())
print(len(unique), "leads,", len(failed), "failed")

# Optional: push each lead to your CRM's REST API (placeholder URL and token).
for lead in unique.values():
    res = requests.post("https://crm.example.com/api/companies", timeout=30,
                        headers={"Authorization": f"Bearer {os.environ['CRM_API_TOKEN']}"},
                        json={k: lead.get(k) for k in ("name", "domain", "location", "size")})
    if not res.ok:
        print("CRM rejected", lead["domain"], res.status_code)
```

For market research, aggregate the rows instead of contacting them, for example a market map by location: `jq -r 'select(.data) | .data.location' companies.jsonl | sort | uniq -c | sort -rn | head`.

## FAQ

**Can I scrape directories that return 403 to a plain request?**

Some can be. In an October 2026 test, a startup job board that returned HTTP 403 to a direct request returned its listings through a `js_render` Spicrawl request for 3 credits. Company profile pages were not tested, so try one before a full run. Check each site's `robots.txt` and terms first.

**What is the best B2B lead scraping API approach for company data?**

Collect profile URLs with `links: true`, then request each one with `extract` to get the name, website, location and size as JSON. Try `autoparse` first when the profile publishes schema.org `Organization` markup. Spicrawl has no site-crawl or search endpoint, so you supply the listing URLs.

**Can a batch job extract company fields?**

No. A batch item returns `html`, `markdown`, `text` or `json`, and a submission that sets `extract`, `autoparse` or `links` is refused with a `400` at 0 credits. Run `extract` per URL, or fetch with a batch job and parse the HTML yourself.

**How do I export scraped leads to a CRM?**

Write the rows to CSV and use your CRM's CSV import, or post each company to its API from the same script. Deduplicate on the website domain first so a re-run updates records instead of copying them.

**Can I scrape directories behind Cloudflare?**

Yes, with your own residential proxy in `proxy`: Spicrawl moves a challenged request to full-browser rendering on the same exit and bills 8 credits. Success is likely, not guaranteed.

**Does Spicrawl follow robots.txt for me?**

No. Spicrawl fetches the URLs you send, so checking `robots.txt`, the site's terms and your request rate is your responsibility. Python's `urllib.robotparser` can filter a URL list.

**Related:** [Structured data](https://docs.spicrawl.com/guides/structured-data.md) · [Batch jobs](https://docs.spicrawl.com/guides/batch.md) · [Browser actions](https://docs.spicrawl.com/guides/browser-actions.md) · [Cloudflare](https://docs.spicrawl.com/guides/cloudflare.md) · [Anti-bot](https://docs.spicrawl.com/guides/anti-bot.md) · [Proxies and geo](https://docs.spicrawl.com/guides/proxies-and-geo.md) · [Credits](https://docs.spicrawl.com/credits.md) · [Rate limits](https://docs.spicrawl.com/rate-limits.md)

 

> **For educational purposes:** The examples on this page are for educational purposes only; the URLs are placeholders, and Spicrawl is not affiliated with any site you scrape. Check each site's terms and `robots.txt`, respect rate limits, and follow the laws that apply to you, including data-protection laws such as the GDPR and CCPA when pages contain personal data. You are responsible for how you use Spicrawl and the data you collect. This is not legal advice.
