Website Contacts scraper artworkEmails, phones & socials

Website Contacts Scraper.

Start with a list of company websites. Collect the contact channels each company publishes and group its social profiles by network, so your team has more to work with than a domain name.

Run on Apify

Ways to use this scraper.

Try these examples for marketing, growth and competitor research.

01

Add contact details to a partnership shortlist

If you've already researched potential stockists, collect their public company channels. Review which channel fits your partnership enquiry.

Keep a table by domain with the published contact details and social links.

02

Fill in an agency's account research

Start with a category-specific domain list from Google Search. Check the websites, collect their social profiles, and review each company's audience and positioning.

Add source pages and your own research notes to the account list.

03

Check your own sites' contact details

For a business with several brands, check whether each site publishes the right emails, phone numbers and social links. Follow up with the website owners on missing details or outdated links.

Keep a contact audit with links to the pages you checked.

FROM FIRST RUN TO REPEATABLE WORKFLOW

Your first run, step by step.

  1. Open the Actor on Apify, select its Input tab and switch to the JSON editor if you want to paste a configuration.
  2. Paste two confirmed website URLs into startUrls. Replace the example sites with companies from your own list. Avoid social profile URLs when you already know the company’s main website.
  3. Start with eight pages per site and maxResults 2. Keep skipSitesWithoutContacts false so you can see sites with no contacts too. No proxy is needed by default; add one only when the source requires it.
  4. Start the run and watch the log. Check which targets and filters were actually read before assuming an empty result means nothing exists.
  5. Open Storage → Dataset, inspect several records and export JSON for nested data or CSV for a first spreadsheet review.
First-run configuration
{
  "startUrls": ["https://apify.com", "https://www.iana.org"],
  "maxPagesPerSite": 8,
  "maxResults": 2,
  "skipSitesWithoutContacts": false
}

Paste this into the Actor’s JSON input editor. Replace the example targets with yours before running.

Check the current input form on Apify

Choose how to search.

The scraper visits a website and prioritizes pages likely to contain public contact details. Give it confirmed website URLs when possible. Company-name shortcuts can guess a domain, so they need an extra identity check.

Search methodWhen to use itWhat changes
Website domainsWhen to use itYou know the exact entity you want to collect.What changesWebsite URLs or bare domains, as strings or objects with a url field. Each site is scanned independently. Company names and social profiles can become guessed domains; use verified websites to avoid mistaken identity.
Company-name shortcutsWhen to use itYou want to find relevant results before choosing specific targets.What changesOne website alongside or instead of startUrls. A business email can be read as its domain, but a verified website is the clearest input.
Crawl depthWhen to use itYou want broader coverage and can allow a larger run.What changes1-100 pages per website, default 8. The crawler prioritizes contact, imprint, about, team and support pages rather than crawling every link. Raise for large sites, not for a broader domain list.

The difference that matters

Cap delivered website records, not email addresses or crawled pages. Each domain normally produces one record containing lists of contacts. Empty processes the whole input list.

Compare the related scraper

Every input, explained.

Use the exact field names below in JSON. In Apify’s form, enter list items separately, choose filters, and keep numbers and booleans in their proper types.

Default and prefill are different. A default applies when you omit a setting; a prefill is an example already entered in Apify’s form. Review prefilled targets and limits before every run. Some settings have no schema default. You still need to supply at least one supported target.

Targets and search inputs2
startUrls
ListForm prefill: ["https://apify.com","https://www.iana.org"]

Website URLs or bare domains, as strings or objects with a url field. Each site is scanned independently. Company names and social profiles can become guessed domains; use verified websites to avoid mistaken identity.

url
Text

One website alongside or instead of startUrls. A business email can be read as its domain, but a verified website is the clearest input.

Limits, details and proxies7
maxPagesPerSite
IntegerDefault: 8

1-100 pages per website, default 8. The crawler prioritizes contact, imprint, about, team and support pages rather than crawling every link. Raise for large sites, not for a broader domain list.

maxResults
Integer

Cap delivered website records, not email addresses or crawled pages. Each domain normally produces one record containing lists of contacts. Empty processes the whole input list.

skipSitesWithoutContacts
True or falseDefault: false

true omits sites that were read but had no contacts; false keeps empty-contact records for coverage review. Unreachable or blocked sites are separate failures, not proof that a company has no contacts.

maxConcurrency
IntegerDefault: 5

1-20 websites at once, default 5. It parallelizes different hosts, not more pages on the same site. Keep a per-site pause when increasing the number of websites.

delayMs
IntegerDefault: 300

Pause between requests to the same website, in milliseconds. It does not pause unrelated websites running in parallel. Default 300 is a useful starting point.

requestTimeoutSecs
IntegerDefault: 20

5-120 seconds to wait for a page, default 20. Raise for slow sites; lower for fast list checks. It is a page timeout, not the total Actor runtime.

proxyConfiguration
ObjectDefault: {"useApifyProxy":false}

Optional for ordinary public business sites; the default uses no proxy. Try a proxy for larger lists or refused pages, and inspect failures rather than labeling blocked sites as empty.

Advanced settings and recovery3
resume
True or falseDefault: true

Saves progress about every 30 seconds so an Apify restart or migration can continue the current run. Leave it on for normal use.

continueFromLastRun
True or falseDefault: false

Continues unfinished work from the previous run with matching input. Earlier results remain in that run’s dataset. Keep false for a fresh collection or a recurring snapshot.

impersonate
TextDefault: chrome131

The browser identity used for requests. Keep the Actor’s default unless you are diagnosing immediate blocks. It changes the request fingerprint, not the data you ask for.

This reference follows the Actor’s published input fields. Check the live form before changing a production workflow. Check the current input form on Apify.

Configurations you can copy.

Each example is a separate run. Start small, inspect the results, then increase coverage. Update the targets, countries and dates to match your question.

Check a confirmed account list

Collect public contacts from four known domains, up to eight pages each. Keep empty rows so you can audit coverage. Replace these example sites with your approved research list before running.

Check a confirmed account list
{
  "startUrls": [
    "https://apify.com",
    "https://www.iana.org",
    "https://www.python.org",
    "https://www.getdbt.com"
  ],
  "maxPagesPerSite": 8,
  "maxResults": 4,
  "skipSitesWithoutContacts": false
}

Look deeper into a few websites

Raise the page budget to 20 for two domains when contact details are spread across several pages. The Actor prioritizes likely contact pages; a larger budget is still a bounded crawl, not every URL on the site.

Look deeper into a few websites
{
  "startUrls": ["https://apify.com", "https://www.getdbt.com"],
  "maxPagesPerSite": 20,
  "maxResults": 2,
  "skipSitesWithoutContacts": false,
  "maxConcurrency": 2,
  "delayMs": 500
}

Export only sites with contacts

Set skipSitesWithoutContacts true when you need a contact-bearing list. Review a run with the switch false first: filtering empty rows makes a convenient export but hides part of your coverage picture.

Export only sites with contacts
{
  "startUrls": [
    "https://apify.com",
    "https://www.iana.org",
    "https://www.python.org",
    "https://www.getdbt.com"
  ],
  "maxPagesPerSite": 8,
  "maxResults": 4,
  "skipSitesWithoutContacts": true
}

Run, check, export, repeat.

Expect one record per website with arrays of public emails, phone numbers and social profiles. Empty contact arrays can mean no published contacts, a blocked page or JavaScript-only content that this HTTP scraper cannot render. Check logs and source pages before labeling a company unreachable. Verify guessed domains before joining records to your account list.

  1. Check the dataset and the run’s SUMMARY record. Compare the number collected with your cap, inspect failed or skipped inputs, and verify a few original source links.
  2. Keep the original IDs and add collected_at and run_id when saving results. Export CSV for flat columns; retain JSON when arrays or nested details matter.
  3. Save the tested configuration as an Apify Task and schedule it. For repeated snapshots, leave continueFromLastRun false. Deduplicate new records by source ID while retaining each observation date.
  4. In Make or n8n, wait for a successful run, fetch its dataset and map fields into Sheets or your warehouse. Send records to Looker Studio through a reporting table; use dbt to flatten and test warehouse models.
  5. The same JSON works with Apify’s Actor API. In Claude with Apify MCP, name this Actor, ask it to inspect the live schema, and give explicit targets, markets and result limits before it runs.

Resume is not a fresh snapshot

resume protects the current run if Apify restarts it. continueFromLastRun continues an earlier run with the same configuration; earlier records stay in the earlier dataset. Combine both datasets for the complete collection, and raise a previously reached result cap when continuing. Start fresh when you want to see what changed today.

Follow the Sheets, Claude, Looker and BigQuery setup guides
Run this Actor from the API

Save one configuration above as input.json. Set APIFY_TOKEN to your Apify API token in your terminal, then send the file as the request body.

Start the run
curl --fail-with-body --request POST \
  --url "https://api.apify.com/v2/actors/jmlp~web-contact-scraper/runs" \
  --header "Authorization: Bearer $APIFY_TOKEN" \
  --header "Content-Type: application/json" \
  --data-binary @input.json

The response contains a run ID and defaultDatasetId, not finished results. Wait for the run to succeed, set DATASET_ID to that dataset ID, then fetch its items. For large datasets, use limit and offset to page through the export.

Fetch the dataset
curl --fail-with-body \
  --url "https://api.apify.com/v2/datasets/$DATASET_ID/items?format=json" \
  --header "Authorization: Bearer $APIFY_TOKEN"

Apify’s run and export API reference
Dataset export options

When the results look wrong.

Change one setting at a time, keep a small cap, and check the run summary before scaling up.

A company name led to the wrong website.

Website URLs or bare domains, as strings or objects with a url field. Each site is scanned independently. Company names and social profiles can become guessed domains; use verified websites to avoid mistaken identity.

A site has contact details, but the result is empty.

Expect one record per website with arrays of public emails, phone numbers and social profiles. Empty contact arrays can mean no published contacts, a blocked page or JavaScript-only content that this HTTP scraper cannot render. Check logs and source pages before labeling a company unreachable. Verify guessed domains before joining records to your account list.

Some input sites are missing from the dataset.

true omits sites that were read but had no contacts; false keeps empty-contact records for coverage review. Unreachable or blocked sites are separate failures, not proof that a company has no contacts.

The run succeeded but returned nothing

Success means the Actor finished handling the request, not that the source returned data. Check SUMMARY.inputProblem, SUMMARY.problem and the log for missing targets, unsupported filters or refused requests. Test one known target with fewer filters.

Fewer records than expected

Check the global limit, per-search or per-page limits, platform coverage and deduplication. Several searches can find the same record. A source’s headline count can include records the public endpoint does not return. Review unfinished jobs before treating the dataset as complete.

Use the results in your tools.

Google Sheets

Join the results on normalized domain. Put contact arrays in dedicated columns and add notes on the preferred contact route, account relevance and source pages.

Read the setup
Claude + MCP

Ask Claude to organize the public channels by network and flag missing or ambiguous records. A general inbox doesn't identify a named person, and a published email isn't a deliverability check.

Read the setup
Looker Studio

Show which domains have an email, phone number or LinkedIn profile. Keep collection errors in the report so you can tell missing data from missing contact channels.

Read the setup
BigQuery + dbt

Save one domain record per run with the collection time. UNNEST emails, phones and social profiles into separate tables for joins, keeping the source-page evidence.

Read the setup
Copy a prompt for Claude
Prompt for Claude + Apify MCP
Inspect jmlp/web-contact-scraper and run it for https://apify.com with at most eight pages per site. Organize the public emails, phones and social profiles, cite the collected pages, and flag errors or missing channels. Do not claim email deliverability is verified.

The fields you’ll get.

Keep the collection time and original IDs with your records. You’ll need them to check where a result came from or compare it with a later run.

Before you draw conclusions

The scraper collects contact details that a website publishes. It can't reveal private details or check whether an email will be delivered. Some sites return no contacts. Review errors and source pages before choosing a company channel for relevant outreach.

domain / url
The deduplication key and the resolved website address.
emails / phones
Publicly listed addresses and phone numbers.
socials / socials_by_network
Public social profiles, including a network grouping.
has_contact
Whether at least one public contact channel was found.
pages / pages_crawled
Collection coverage and the pages visited.
errors
Collection issues that need review.

Common questions.

Does it verify email deliverability?

No. It extracts addresses published on a website. Deliverability checking is a separate process.

Can I connect it to Google Search results?

Yes. Validate the discovered websites, normalize domains, and use that reviewed list as input. Check reconstructed search URLs before passing them to another crawler.

Source and current product details: JMLP’s Website Contacts Actor on Apify.

Other scrapers
you might use.

Let’s talk about
your project.

Tell me what you need to collect or understand. I can help with a custom scraper, a pipeline or the analysis.

Start a project