Google Maps Lead Generation
Generate high-quality B2B leads from Google Maps with deep contact enrichment.
Overview
This pipeline scrapes Google Maps for businesses, then enriches each result by:
- Scraping their website (main page + up to 5 contact pages)
- Searching DuckDuckGo for additional contact info
- Using Claude to extract structured contact data from all sources
Tested at scale: 50+ leads per run, 68 total leads across plumbers, electricians, HVAC, and roofing contractors.
When to Use
- Building outbound sales lists for local service businesses
- Generating leads for B2B services (contractors, medical, legal, etc.)
- Researching businesses in a specific geographic area
- Creating prospecting lists with verified contact info
Inputs
| Parameter |
Required |
Description |
--search |
Yes |
Search query (e.g., "plumbers in Austin TX") |
--limit |
No |
Max results to scrape (default: 10) |
--location |
No |
Additional location filter |
--sheet-url |
No |
Existing Google Sheet to append to |
--sheet-name |
No |
Name for new sheet if creating |
--workers |
No |
Parallel workers for enrichment (default: 3) |
Execution
# Basic usage - creates new sheet
python3 execution/gmaps_lead_pipeline.py --search "plumbers in Austin TX" --limit 10
# Append to existing sheet (recommended for building lead database)
python3 execution/gmaps_lead_pipeline.py --search "dentists in Miami FL" --limit 25 \
--sheet-url "https://docs.google.com/spreadsheets/d/..."
# Higher volume run
python3 execution/gmaps_lead_pipeline.py --search "roofing contractors in Austin TX" \
--limit 50 --workers 5
Output Schema (36 fields)
Business Basics (from Google Maps)
business_name, category, address, city, state, zip_code, country
phone, website, google_maps_url, place_id
rating, review_count, price_level
Extracted Contacts (from website + web search + Claude)
emails - All email addresses found (comma-separated)
additional_phones - Phone numbers from website
business_hours - Operating hours
Social Media
facebook, twitter, linkedin, instagram, youtube, tiktok
Owner/Key Person Info
owner_name, owner_title, owner_email, owner_phone, owner_linkedin
Team Contacts
team_contacts - JSON array of team members with name, title, email, phone, linkedin
Metadata
lead_id - Unique identifier (MD5 hash of name|address, for deduplication)
scraped_at - ISO timestamp
search_query - Original search term used
pages_scraped - Number of pages fetched (1 main + up to 5 contact pages)
search_enriched - Whether DuckDuckGo search was used (yes/no)
enrichment_status - success/partial/error
Pipeline Steps
- Google Maps Scrape - Apify
compass/crawler-google-places actor returns business listings with basic info
- Website Scraping - Fetches main page + up to 5 prioritized contact pages (/contact, /about, /team, etc.)
- Web Search Enrichment - DuckDuckGo search for
"{business}" owner email contact + scrapes first relevant result
- Claude Extraction - Claude 3.5 Haiku extracts structured contacts from all gathered content
- Google Sheet Sync - Appends new leads, automatically deduplicates by
lead_id
Contact Page Patterns (22 total, priority-ordered)
High priority: /contact, /about, /team, /contact-us, /about-us, /our-team
Medium: /staff, /people, /meet-the-team, /leadership, /management, /founders, /who-we-are
Lower: /company, /meet-us, /our-story, /the-team, /employees, /directory, /locations, /offices
Cost Considerations
| Component |
Cost per lead |
| Apify Google Maps |
~$0.01-0.02 |
| Claude Haiku extraction |
~$0.002 |
| DuckDuckGo search |
Free |
| HTTP requests (6-7 pages) |
Free |
| Google Sheets |
Free |
| Total |
~$0.012-0.022 |
For 100 leads: ~$1.50-2.50 total
The pipeline maximizes value per Apify dollar by scraping 6+ pages + web search per business.
Dependencies
apify-client
httpx
html2text
anthropic
gspread
google-auth
google-auth-oauthlib
python-dotenv
Files
execution/gmaps_lead_pipeline.py - Main orchestration script
execution/scrape_google_maps.py - Google Maps scraper (standalone)
execution/extract_website_contacts.py - Website contact extractor (standalone)
Troubleshooting
"No businesses found"
- Check search query is valid
- Include location in query (e.g., "plumbers in Austin, TX" not just "plumbers")
403 Forbidden errors
- ~10-15% of sites block scrapers with 403/503 errors
- These are handled gracefully and marked as errors in
enrichment_status
- The lead is still saved with Google Maps data (phone, address, etc.)
"Could not fetch website"
- Some sites have broken DNS or are offline
- Marked as
error in enrichment_status
- Reduce
--workers if seeing many timeouts
"APIFY_API_TOKEN not found"
- Ensure
.env file has valid Apify token
- Check token hasn't expired at apify.com
Google Sheet auth issues
- Delete
token.json and re-authenticate
- Ensure
credentials.json is valid OAuth client
Duplicate detection
- Pipeline uses
lead_id (MD5 of name|address) to skip existing leads
- Running same search twice will show "No new leads to add (all duplicates)"
Learnings
- Google Maps actor returns
website field directly - no need to scrape for it
- Contact pages commonly use /contact, /about, /team URL patterns
- Claude Haiku is sufficient for extraction and costs 10x less than Sonnet
- ~10-15% of business websites return 403/503 errors - this is normal
- Facebook URLs always fail with 400 errors (blocks scrapers)
- Some sites have broken DNS - handled gracefully as errors
- DuckDuckGo HTML search is free and doesn't block (unlike Google)
stringify_value() helper needed because Claude sometimes returns dicts instead of strings
- Deduplication by lead_id prevents re-adding existing businesses across runs
- 50 leads takes ~3-4 minutes with 3 workers
Production Sheet
Active lead database: https://docs.google.com/spreadsheets/d/1ATrOiq3wfph8Or5BE8VCybgvqK5gh7hVPWiSlgb3QiU
Contains: plumbers, electricians, HVAC contractors, roofing contractors (Austin TX)
1---2name: gmaps-lead-generation3description: Google Maps Lead Generation4---5# Google Maps Lead Generation67Generate high-quality B2B leads from Google Maps with deep contact enrichment.89## Overview1011This pipeline scrapes Google Maps for businesses, then enriches each result by:121. Scraping their website (main page + up to 5 contact pages)132. Searching DuckDuckGo for additional contact info143. Using Claude to extract structured contact data from all sources1516**Tested at scale**: 50+ leads per run, 68 total leads across plumbers, electricians, HVAC, and roofing contractors.1718## When to Use1920- Building outbound sales lists for local service businesses21- Generating leads for B2B services (contractors, medical, legal, etc.)22- Researching businesses in a specific geographic area23- Creating prospecting lists with verified contact info2425## Inputs2627| Parameter | Required | Description |28|-----------|----------|-------------|29| `--search` | Yes | Search query (e.g., "plumbers in Austin TX") |30| `--limit` | No | Max results to scrape (default: 10) |31| `--location` | No | Additional location filter |32| `--sheet-url` | No | Existing Google Sheet to append to |33| `--sheet-name` | No | Name for new sheet if creating |34| `--workers` | No | Parallel workers for enrichment (default: 3) |3536## Execution3738```bash39# Basic usage - creates new sheet40python3 execution/gmaps_lead_pipeline.py --search "plumbers in Austin TX" --limit 104142# Append to existing sheet (recommended for building lead database)43python3 execution/gmaps_lead_pipeline.py --search "dentists in Miami FL" --limit 25 \44 --sheet-url "https://docs.google.com/spreadsheets/d/..."4546# Higher volume run47python3 execution/gmaps_lead_pipeline.py --search "roofing contractors in Austin TX" \48 --limit 50 --workers 549```5051## Output Schema (36 fields)5253### Business Basics (from Google Maps)54- `business_name`, `category`, `address`, `city`, `state`, `zip_code`, `country`55- `phone`, `website`, `google_maps_url`, `place_id`56- `rating`, `review_count`, `price_level`5758### Extracted Contacts (from website + web search + Claude)59- `emails` - All email addresses found (comma-separated)60- `additional_phones` - Phone numbers from website61- `business_hours` - Operating hours6263### Social Media64- `facebook`, `twitter`, `linkedin`, `instagram`, `youtube`, `tiktok`6566### Owner/Key Person Info67- `owner_name`, `owner_title`, `owner_email`, `owner_phone`, `owner_linkedin`6869### Team Contacts70- `team_contacts` - JSON array of team members with name, title, email, phone, linkedin7172### Metadata73- `lead_id` - Unique identifier (MD5 hash of name|address, for deduplication)74- `scraped_at` - ISO timestamp75- `search_query` - Original search term used76- `pages_scraped` - Number of pages fetched (1 main + up to 5 contact pages)77- `search_enriched` - Whether DuckDuckGo search was used (yes/no)78- `enrichment_status` - success/partial/error7980## Pipeline Steps81821. **Google Maps Scrape** - Apify `compass/crawler-google-places` actor returns business listings with basic info832. **Website Scraping** - Fetches main page + up to 5 prioritized contact pages (/contact, /about, /team, etc.)843. **Web Search Enrichment** - DuckDuckGo search for `"{business}" owner email contact` + scrapes first relevant result854. **Claude Extraction** - Claude 3.5 Haiku extracts structured contacts from all gathered content865. **Google Sheet Sync** - Appends new leads, automatically deduplicates by `lead_id`8788## Contact Page Patterns (22 total, priority-ordered)8990High priority: `/contact`, `/about`, `/team`, `/contact-us`, `/about-us`, `/our-team`91Medium: `/staff`, `/people`, `/meet-the-team`, `/leadership`, `/management`, `/founders`, `/who-we-are`92Lower: `/company`, `/meet-us`, `/our-story`, `/the-team`, `/employees`, `/directory`, `/locations`, `/offices`9394## Cost Considerations9596| Component | Cost per lead |97|-----------|---------------|98| Apify Google Maps | ~$0.01-0.02 |99| Claude Haiku extraction | ~$0.002 |100| DuckDuckGo search | Free |101| HTTP requests (6-7 pages) | Free |102| Google Sheets | Free |103| **Total** | **~$0.012-0.022** |104105**For 100 leads**: ~$1.50-2.50 total106107The pipeline maximizes value per Apify dollar by scraping 6+ pages + web search per business.108109## Dependencies110111```112apify-client113httpx114html2text115anthropic116gspread117google-auth118google-auth-oauthlib119python-dotenv120```121122## Files123124- `execution/gmaps_lead_pipeline.py` - Main orchestration script125- `execution/scrape_google_maps.py` - Google Maps scraper (standalone)126- `execution/extract_website_contacts.py` - Website contact extractor (standalone)127128## Troubleshooting129130### "No businesses found"131- Check search query is valid132- Include location in query (e.g., "plumbers in Austin, TX" not just "plumbers")133134### 403 Forbidden errors135- ~10-15% of sites block scrapers with 403/503 errors136- These are handled gracefully and marked as errors in `enrichment_status`137- The lead is still saved with Google Maps data (phone, address, etc.)138139### "Could not fetch website"140- Some sites have broken DNS or are offline141- Marked as `error` in enrichment_status142- Reduce `--workers` if seeing many timeouts143144### "APIFY_API_TOKEN not found"145- Ensure `.env` file has valid Apify token146- Check token hasn't expired at apify.com147148### Google Sheet auth issues149- Delete `token.json` and re-authenticate150- Ensure `credentials.json` is valid OAuth client151152### Duplicate detection153- Pipeline uses `lead_id` (MD5 of name|address) to skip existing leads154- Running same search twice will show "No new leads to add (all duplicates)"155156## Learnings157158- Google Maps actor returns `website` field directly - no need to scrape for it159- Contact pages commonly use /contact, /about, /team URL patterns160- Claude Haiku is sufficient for extraction and costs 10x less than Sonnet161- ~10-15% of business websites return 403/503 errors - this is normal162- Facebook URLs always fail with 400 errors (blocks scrapers)163- Some sites have broken DNS - handled gracefully as errors164- DuckDuckGo HTML search is free and doesn't block (unlike Google)165- `stringify_value()` helper needed because Claude sometimes returns dicts instead of strings166- Deduplication by lead_id prevents re-adding existing businesses across runs167- 50 leads takes ~3-4 minutes with 3 workers168169## Production Sheet170171Active lead database: https://docs.google.com/spreadsheets/d/1ATrOiq3wfph8Or5BE8VCybgvqK5gh7hVPWiSlgb3QiU172173Contains: plumbers, electricians, HVAC contractors, roofing contractors (Austin TX)