Find Influencers
Core contract
Turn the user's request into a confirmed rubric, research with read-only Webcmd commands, qualify only creators who meet every required criterion, and return every evaluated creator in one evidence-backed CSV output. Never follow, subscribe, like, message, post, reply, comment, or contact a creator.
Prerequisites
- Run
node --version; Node.js 20 or newer is required. - Run
webcmd --version. If Webcmd is missing, ask before installing it globally withnpm install -g @agentrhq/webcmd. - Load
webcmd:webcmd-usagebefore live Webcmd work and follow its complete filtered-registry discovery order. - Verify requested commands with
webcmd youtube --helpandwebcmd twitter --help. - Run
webcmd doctor, thenwebcmd youtube whoami -f jsonand/orwebcmd twitter whoami -f json. Use the matching login command and human handoff when authentication is required; rerunwhoamiafter the user reports completion.
Workflow
1. Build the rubric
Extract what the user already supplied and ask only for missing information, one question at a time:
- topic, niche, product, or campaign;
- YouTube, X/Twitter, or both;
- qualified target count;
- numeric and qualitative criteria;
- whether each criterion is required or preferred;
- acceptable evidence and confidence for estimates;
- exclusions and seed accounts.
YouTube-specific inputs
When YouTube is in scope, collect only missing YouTube-specific inputs: subscriber range with optional minimum and maximum; recent-view metric (median recommended, average, minimum on every sampled video, or a confirmed pass-count); recent-view range with optional floor and ceiling; exact eligible-video sample size; eligible formats (long-form, Shorts, livestreams, or a confirmed mixture); optional upload window or posting frequency; optional likes metric (median or average) or like-to-view ratio; optional public business email; and any user-requested faceless or on-camera requirement. Do not apply these requirements to X/Twitter.
Faceless or on-camera evaluation is opt-in. Add it to the required or preferred rubric only when the user requests it; otherwise do not ask for it as missing information. Do not run frame capture when the confirmed rubric has no faceless or on-camera criterion.
Treat subscriber and view ceilings as a campaign-fit proxy, not verified sponsorship pricing. State that limitation in the confirmed rubric whenever affordability motivates a ceiling. Do not infer sponsorship rates, costs, or willingness from those ceilings without accepted direct evidence. Recommend median views because one viral upload distorts it less than an average, but never force that default without confirmation.
Classify every user-supplied constraint, including topic or content fit, as required or preferred. Do not leave a constraint outside those lists.
Default to 20 qualified creators. Unless the user changes it, set the discovery budget to max(20, five times the qualified target), capped at 100 unique creators total. The budget counts resolved, deduplicated creators rather than raw search rows.
Rubric confirmation
Restate the complete rubric, including search scope, discovery budget, required criteria, preferred criteria, evidence rules, estimates, exclusions, seed accounts (none when absent), and target count, and wait for confirmation before searching. When the user asks to see the planned workflow, include it with the next missing-information question and explicitly preview numeric gating before qualitative enrichment, unknown preservation, and the one all-candidate CSV. A changed rubric requires reconfirmation. Refuse protected-characteristic or sensitive-trait targeting and request content- or audience-relevance alternatives. Never loosen a confirmed rubric because too few creators qualify; never loosen it for a shortage of matches.
2. Discover candidates
Build several focused topic searches without silently expanding beyond the confirmed scope. Derive semantic query variants from the user's topic rather than repeating one phrasing. For a GitHub/open-source request, cover relevant variants such as GitHub repos, GitHub repositories, open-source tools, and trending GitHub; preserve the user's topic terms in every variant. When the rubric depends on current or recent performance, add recency-oriented queries using the current year, recent-period wording, or supported upload filters.
- YouTube:
webcmd youtube search "<query>" --type video --limit <N> -f json, then resolve each result withwebcmd youtube video <url> -f jsonto obtain its stable channel ID. - X/Twitter:
webcmd twitter search "<query>" --product top --limit <N> -f jsonand collect author handles. - Add exact user-supplied seed handles or URLs.
Deduplicate after every query. If duplicate channels consume at least half of a query's rows, run another unused semantic or recency variant instead of treating those duplicates as discovery coverage. Do not stop merely because the qualified target has been reached; stop at the confirmed budget or exhausted query variants. Mark surfaced creators that were not resolved because of the budget as not-evaluated, which is distinct from not-qualified. Report query strings, per-query limits, raw rows, unique resolved creators, duplicates, and not-evaluated coverage; never imply an exhaustive platform search.
3. Resolve identities
Deduplicate YouTube by channel ID and X by normalized handle. Merge platforms only when an official profile, bio, website, or other explicit source links them, and record that source. Single-platform creators remain eligible and are not penalized for lacking the other platform.
4. Numeric gate
Collect inexpensive numeric evidence first:
webcmd youtube channel <channelId> --limit 30 -f jsonwebcmd youtube video <url> -f jsonwebcmd twitter profile <handle> -f jsonwebcmd twitter tweets <handle> --limit 20 -f json
Normalize human-readable counts while preserving source values. Treat webcmd youtube channel recent-video rows as candidate uploads, not guaranteed chronological truth.
For YouTube only, when a rubric refers to the latest N videos: resolve each candidate latest video with webcmd youtube video <url> -f json; normalize exact publication timestamps as publishedAt; sort publishedAt descending; apply confirmed format exclusions; deduplicate by stable video ID; then take the first confirmed number of eligible uploads. Resolve extra candidates whenever channel rows omit views or dates, contain collaborations, duplicates, shelf contamination, or suspicious ordering. If the true exact eligible-video sample cannot be established, mark that criterion unknown rather than using a popularity-ordered proxy. Require the confirmed sample size; if fewer eligible videos can be established, mark the criterion unknown and never calculate a passing statistic from a smaller sample. Use the same verified eligible-video sample for views and likes unless the rubric explicitly says otherwise. Compute only confirmed metrics: average_views = sum(views) / count(views); median views using the conventional sorted-sample median; minimum and maximum views when required by a range; average_likes = sum(likes) / count(likes); median likes using the conventional sorted-sample median; and like_view_rate = sum(likes) / sum(views). Hidden likes or unavailable values remain unknown and cannot satisfy a hard requirement.
For YouTube only, regardless of sampling metric: Derive activity only from exact publish dates. Include a public business email only when requested and plainly public in adapter-returned channel or video descriptions; record not_public when verified absent and unknown when unavailable. Do not bypass sign-in gates or CAPTCHA, use personal-data brokers, or guess email patterns.
Evaluate required numeric criteria before qualitative enrichment. Record explicit numeric failures in the CSV and skip their expensive enrichment.
5. Qualitative enrichment
For numeric survivors, collect only evidence required by the rubric. By default inspect up to five recent or relevant content items and up to twenty comments or replies per selected item:
webcmd youtube transcript <url> --mode grouped -f jsonwebcmd youtube comments <url> --limit 20 -f json- When the confirmed rubric includes faceless or on-camera evidence:
webcmd youtube frames <url> --count 5 -f json webcmd twitter tweets <handle> --limit 20 -f jsonwebcmd twitter thread <url> --limit 20 -f jsonwebcmd twitter followers <handle> --limit 50 -f json
Use channel/profile bios, descriptions, transcripts, post text, sampled thumbnails or media posters, comments, replies, and follower bios. Record exact source URLs for criterion-relevant claims.
For a confirmed faceless or on-camera criterion, capture five frames for every video in the confirmed qualitative sample and inspect every returned PNG. Record the video URL, requested and actual timestamps, PNG paths, and capture status. A video is not-faceless when the creator or presenter is visibly on camera in any inspected frame. Faces of incidental people, actors, gameplay characters, photographs, or people in the material being discussed do not establish that the creator or presenter is on camera. A video is faceless-in-sample only when all five frames were captured and inspected without the creator or presenter appearing. Use unknown when any frame failed, the person cannot be identified as the creator or presenter, or the visual evidence is ambiguous.
At channel level, any not-faceless video makes a required faceless criterion not-met; mark it met only when every video in the confirmed sample is faceless-in-sample, with at most medium confidence. Apply the inverse statuses for an on-camera criterion. Never use thumbnails or media posters as a substitute for requested frame evidence; frame sampling does not prove what appears throughout a full video. Public comments, replies, and follower bios may support a confidence-labeled audience geography estimate only when the rubric permits estimates. Describe the sample and label confidence high|medium|low|unknown; label it as an estimate, not platform analytics.
6. Decide and rank
For each criterion record met|not-met|unknown, concise evidence, high|medium|low|unknown confidence, and source URLs.
qualified: every required criterion is met at an accepted confidence.not-qualified: at least one required criterion is explicitly not met.unknown: no required criterion failed, but one or more cannot be established.
Use not-evaluated only for discovery rows outside the confirmed resolution budget; it is a coverage label, not a qualification decision.
Rank only qualified creators by preferred-criterion matches, using evidence confidence to break ties. Do not use a default 100-point score or replace criterion-level reasoning with an unexplained aggregate score.
7. Return the CSV
Return one UTF-8 CSV inline in the final response by default, using a fenced csv code block. Do not create a CSV file unless the user explicitly requests one. Include every evaluated creator exactly once and use proper CSV quoting.
Stable columns, in this exact order: rank, decision, creator_name, platform_availability, youtube_handle, youtube_profile_url, x_handle, x_profile_url, cross_platform_merge_evidence, youtube_subscribers, youtube_sampled_view_statistics, x_followers, x_sampled_engagement_statistics, audience_region_estimate, audience_region_sample, audience_region_confidence, face_signal, face_signal_confidence, overall_reason, unresolved_evidence, source_urls.
For every confirmed criterion, convert its name to a stable snake_case slug and add <criterion>_status, <criterion>_evidence, and <criterion>_confidence in rubric order. Use the Python standard library's csv module with io.StringIO to validate the generated CSV in memory: verify the exact header sequence, row count, and one row per merged creator before reporting completion.
If the user explicitly requests a file path, write the verified CSV there, reopen it with the csv module, and report the absolute path only when a file was requested and written. Otherwise report the rubric, actual discovery and enrichment coverage, qualified and unknown totals, top matches, and limitations, followed by the inline CSV. For an inline YouTube response, name the exact eligible-video sample and disclose any affordability proxies in the final summary.
Evidence record
{
"creatorKey": "stable merged or platform ID",
"platforms": {"youtube": {}, "x": {}},
"decision": "qualified|not-qualified|unknown",
"criteria": {
"criterion-slug": {
"status": "met|not-met|unknown",
"evidence": "criterion-relevant observation",
"confidence": "high|medium|low|unknown",
"sources": ["exact URL"]
}
},
"reason": "concise decision explanation"
}
Stop conditions
- Missing rubric confirmation: do not search.
- Entire requested platform unavailable: ask whether to continue with reduced scope.
- Truncated discovery or enrichment output: do not imply full coverage or infer absent evidence.
- Private, rate-limited, unavailable, or malformed evidence: mark affected criteria unknown.
- Zero matches: report zero without changing the rubric; offer a separately confirmed revision.
- Interrupted run: return a partial inline CSV only when unfinished candidates and coverage limits are explicit.
Current Webcmd limitations
- YouTube channel-filtered search does not expose useful structured channel rows; discover videos and resolve their channel IDs.
- YouTube channel recent-video rows can omit views for collaborations or mix shelves and ordering; resolve video metadata directly for latest-N gates.
- Faceless conclusions are limited to successfully inspected frames; frame sampling is not full-video proof.
- Comments, replies, follower bios, and public interactions are samples, not audience analytics.
Common mistakes
- Searching before rubric confirmation.
- Treating creators outside the discovery budget as not-qualified instead of not-evaluated.
- Reusing one search phrase without semantic query variants or duplicate-aware expansion.
- Assuming YouTube channel rows are chronologically ordered when applying a latest-video gate.
- Treating unknown evidence as a failed criterion.
- Deeply enriching candidates who already failed a required numeric gate.
- Merging accounts by similar names without an explicit linking source.
- Running frame capture when faceless or on-camera evidence is absent from the confirmed rubric.
- Treating incidental faces as proof that the creator presents on camera.
- Exporting only qualified candidates instead of every evaluated creator.
- Writing a CSV file when the user did not explicitly request one.