CRM Lookalikes
Your real ICP is written in your closed-won list, not in your pitch deck. Extract the pattern, then search 70M+ companies for more of it.
Works for PEOPLE too, not only companies: search_sql_users has a lookalike graph —
similar_to: [<best customer contact aliases>] (tight) / also_viewed (loose) plus
normal filters. Same discipline: the user confirms the seed set, candidates get scored.
Flow
1. Collect the seed set
crm_query_records(object_type="companies", list_id=<customers list> | search=...,
properties=[record_id, name, domain, industry, <size/stage if mapped>])
Need the user's help to identify "best": a customers list, a lifecycle/status field, or an explicit pick of 10–30 names. Fewer than ~8 seeds → warn that the pattern will be weak.
2. Profile the seeds
Resolve each seed to structured firmographics — exact verification is mandatory on every
resolve (the website search is substring match and can return only look-alike domains;
a wrong seed poisons the whole ICP pattern downstream):
execute linkedin/search/search_sql_companies {website: "seed1.com", count: 5} # per seed
# batched variant allowed, but: any seed without an exact match must be re-queried
# individually. query_cache filters the WHOLE cached set; `limit` (default 10) caps only
# how many rows come back — pass one when a batch should return more than 10 matches.
query_cache {conditions: [{"field": "website", "op": "=", "value": "seed1.com"}], limit: 50}
A seed with no exact website match is NOT dropped yet — resolve it via the site itself
(webparser/parse {url, extract_minimal: true} → top-level title + own linkedin.com/company
URL in links[] → linkedin/company), or via crunchbase → contacts.linkedin_url. Only a
seed that survives neither is excluded from profiling, and say which ones.
Plus crunchbase/company for stage/funding on a subset (venture-relevant seeds only).
Derive the pattern in-session and SHOW it:
Industries: X (60%), Y (25%) · Size: 11-200 dominant · Geo: US+UK 80%
Stage: seed-B · Common traits: has API docs page, hiring in data roles, ...
Build the size band from employee_count, never from employee_count_range — the two can
contradict each other in one record (verified: employee_count: 1465 with
employee_count_range: "201-500"), and a wrong band here propagates into every search
below. Bucket the exact counts yourself.
The user confirms/edits the pattern — it's their ICP, the data only proposes it.
3. Search for lookalikes
execute linkedin/search/search_sql_companies— industry_name/keywords DSL from the pattern, employee_count band, country filter, count up to 1000.execute crunchbase/db/db_search— when stage matters (last_funding_type,last_funding_date_after);crunchbase/searchlive forhiring: trueorshares_investors_with: [<seed investors>](a strong hidden-similarity filter).- Niche supplements per pattern:
yc/search/search_companies(early-stage),builtin(US tech hubs),producthunt(product-led).
Search wide, profile narrow: the searches themselves are cheap even at count 1000, but do NOT enrich every candidate — score on the fields the search already returned, and fetch extra evidence (crunchbase lookups etc.) only for the top ~50. State the credit estimate before any per-candidate enrichment.
4. Score and dedup
Score candidates against the confirmed pattern (same rubric discipline as
anysite-crm-score — weighted criteria, evidence per company, no guessed values).
Dedup against the CRM by domain (crm_query_records) — existing accounts drop out or get
flagged "already in CRM, unworked".
5. Hand off
Output: top-N table (name, domain, why-it-matches, score) + the confirmed ICP pattern for
reuse. Pushing to CRM → anysite-crm-prospect (its dedup/create/working-list rules apply);
finding people at these companies → same skill. This skill itself writes nothing.