seo-master
Evidence first. Separate observed facts, inferences, and proposals. Never present
technical eligibility as a promise of ranking, citation, recommendation, or
traffic.
0. Quick start
Use the smallest workflow that answers the request:
Audit [site] for technical SEO + GEO. Read-only. Sample [N] URLs.
Fix on-page + truthful schema for [URL]. Preserve brand and analytics.
Measure AI visibility for [brand] across [platforms] with [query set].
Compare SEO vs GEO opportunities for [market, locale, device].
Before work, state:
- target host, locale, audience, site type, and conversion goal;
- read-only audit or authorized edits;
- access available: repository, browser, logs, GSC, GA4, rank data, AI tools;
- sample or full crawl, URL cap, and time window;
- house property: yes/no. Apply section 13 only when yes.
If required access is missing, continue with observable evidence and label the
unobservable layer UNVERIFIED; do not fill gaps with estimates.
1. Scope & when-not-to
Do run when
- The site has a public discovery or acquisition goal.
- The user asks about rankings, citations, AI mentions, crawlability, content,
schema, performance, links, analytics, or search reporting.
- A deployable path exists and the requested access is authorized.
Do NOT run (anti-patterns)
| Situation |
Action |
| Private-by-design sites (e.g. red-flowers: noindex + WAF CN) |
Skip entirely — never SEO-refine |
| Dead / abandoned projects with no publish path |
Report kill criteria; do not polish |
| Spam tactics (PBNs, cloaking, link farms, doorway pages) |
Refuse; propose white-hat alternative |
| Keyword stuffing / hidden text |
Strip; rewrite pain-first |
| Fabricating GSC verification tokens / fake metrics |
REPORT presence only; never invent |
| Brand/asset redesign under SEO pretext |
Out of scope (see House rules) |
projects/_templates/**, docs/AUDIT.md, payment code |
Do not touch |
Scope questions
Ask only questions that materially alter execution:
- Which canonical host, markets, languages, and priority conversions?
- Audit only, recommendations, or code/content changes?
- Which evidence sources are authorized and available?
- Which AI surfaces matter to the audience: ChatGPT, Gemini, Google AI
Overviews/AI Mode, Perplexity, Doubao, Claude, or another engine?
- What must remain untouched?
Record unresolved choices as assumptions. Never let a missing optional data
source block checks that can be performed safely.
2. Technical SEO audit
Audit in dependency order: reachability → renderability → crawlability →
indexability → canonicalization → discovery → internal architecture.
2.1 Evidence sample
Include at minimum, when in scope:
- homepage;
- one URL per indexable template;
- priority conversion pages;
- one paginated, filtered, localized, and JavaScript-heavy URL where present;
- redirect, 404, soft-404 candidate, and canonical duplicate samples;
- URLs reported by GSC or logs, if access exists.
For every finding record URL, timestamp, method, observed value, expected value,
evidence, confidence, impact, and affected scope. Do not extrapolate a sampled
finding to the whole site without a crawl or corroborating data.
2.2 Reachability, rendering, and crawlability
- Record DNS/TLS failures, final GET status, redirect chain, latency, and host.
- Compare raw HTML with rendered DOM for canonical, robots meta, headings,
links, primary copy, and structured data.
- Parse
robots.txt for the tested user agent. A Disallow controls crawling;
it is not a page-level noindex.
- Check WAF/CDN responses separately from robots policy. A spoofed user-agent
request does not authenticate a vendor crawler.
- Test robots and sitemap files on every host/subdomain that serves sampled URLs,
redirects to or from the canonical host, or appears in hreflang/sitemaps.
- Flag infinite spaces: parameters, faceted paths, calendars, session IDs,
internal search, and duplicate sort orders.
2.3 Indexability
For each sampled URL classify:
INDEXABLE, EXCLUDED-AS-INTENDED, BLOCKED-BY-TECHNICAL-CONFLICT, or
UNVERIFIED-IN-INDEX.
Check:
- final status is indexable content, not redirect/error/soft error;
meta robots and X-Robots-Tag;
- whether robots permits the tested crawler to fetch page-level directives; if
blocked, classify the directive's crawler visibility as
UNVERIFIED;
- canonical target status, indexability, locale, and content equivalence;
- GSC URL Inspection/indexing reports when authorized;
- orphan status and internal-link discoverability.
Never claim “indexed” from a 200 response, sitemap entry, or site: query alone.
2.4 Canonical, host, redirects, and hreflang
- Treat owner intent, production bindings, redirects, canonicals, sitemaps,
internal links, and analytics configuration as evidence. Resolve conflicts;
do not assume one signal is universally authoritative.
- Require one-hop permanent redirects for retired canonical URLs where safe.
- Reject loops, chains, protocol/host oscillation, mass redirects to irrelevant
pages, and query loss that changes meaning.
- Require self-referential canonical on canonical pages unless a documented
cross-domain or variant strategy says otherwise.
- For localized equivalents, require reciprocal valid language-region codes,
indexable canonical targets, and
x-default only when a true fallback exists.
2.5 Sitemaps and architecture
- Parse XML; do not count matching lines.
- Validate sitemap-index recursion, XML syntax, HTTP status, canonical host,
indexable URLs,
lastmod truthfulness, and URL limits.
- Exclude redirects, errors, duplicates, blocked/noindex URLs, internal search,
and noncanonical parameters.
- Confirm priority pages receive descriptive internal links from the topic,
product, category, or locale hubs that contain them in the site architecture.
- Report crawl depth as observed distribution, not a universal pass/fail number.
2.6 Technical scorecard
Use statuses, not a pseudo-scientific score:
| Layer |
Status |
Evidence |
Scope |
Confidence |
Next action |
| Reach/render |
PASS/FIX/BLOCK/UNVERIFIED |
URL/log |
sample/all |
H/M/L |
... |
| Crawl/index |
PASS/FIX/BLOCK/UNVERIFIED |
directive/GSC |
sample/all |
H/M/L |
... |
| Canonical/redirect |
PASS/FIX/BLOCK/UNVERIFIED |
chain/HTML |
sample/all |
H/M/L |
... |
| Sitemap/links |
PASS/FIX/BLOCK/UNVERIFIED |
parsed XML/crawl |
sample/all |
H/M/L |
... |
3. On-page SEO
Optimize the page for one primary intent and the jobs needed to satisfy it.
Format-vs-intent gate (first filter)
Classify primary intent, then reject format mismatch before copy polish.
| Intent |
Fit format |
Reject |
| Explore / shortlist |
Listicle, category roundup |
Decision page forced into “N alternatives” |
| Decide / choose (X vs Y, alt to X) |
Comparison or alternative page |
Listicle that never answers “which should I pick” |
| Learn / how-to |
Explainer, procedure |
Comparison shell with no decision criteria |
Decision-intent pages fail this gate if they are listicles. Fix format first.
Comparison/alternative page skeleton
- Match title, visible heading, opening, sections, media, and CTA to intent.
- Make titles and descriptions specific and nonduplicative; judge truncation
from rendered SERPs when available, not a fixed character quota.
- Use a clear page topic and logical heading hierarchy. Do not enforce “exactly
one H1” as a ranking rule; flag confusing document structure.
- Put the useful answer before biography, history, or promotional filler.
- Add internal links where they advance the task, using descriptive anchors.
- Give images meaningful alternatives when informative; use empty
alt for
decorative images.
- Check visible copy against canonical, schema, Open Graph, and product facts.
- Remove doorway duplication, hidden text, keyword stuffing, and templated
filler.
Output: KEEP, REWRITE, MERGE, REDIRECT, NOINDEX, or DELETE for each
page, with evidence and destination where applicable.
4. Content & E-E-A-T
E-E-A-T is a qualitative quality lens, not a published Google score.
4.1 Evidence gate
For consequential claims, require:
- a named source, first-party evidence, or clearly described methodology;
- author/reviewer identity and demonstrated competence for the claim where an
incorrect answer could affect health, safety, rights, or material decisions;
- publication and material-update dates when freshness matters;
- ownership, editorial, contact, correction, privacy, and commercial disclosure
required by the site's function, jurisdiction, risk, or commercial incentives;
- claims that match the cited source and do not exceed it.
For YMYL topics, escalate unsupported medical, legal, financial, or safety advice.
Never create credentials, reviews, test results, customers, or outcomes.
4.2 CORE-EEAT review
【Proposal】 Use this internal rubric for diagnosis only:
| Dimension |
Test |
| Coverage |
Does the page complete the user's task and answer necessary follow-ups? |
| Originality |
Is there first-hand evidence, analysis, data, or a useful synthesis? |
| Recency |
Are time-sensitive claims reviewed and dated? |
| Evidence |
Can a reviewer trace consequential claims to reliable sources? |
| Experience |
Are first-hand methods, constraints, and outcomes demonstrated? |
| Expertise |
Does the author/reviewer demonstrate credentials or first-hand competence for the claim? |
| Authority |
Do independent sources in the same field recognize the entity or work? |
| Trust |
Are identity, incentives, limitations, and corrections transparent? |
Assign PASS, WEAK, FAIL, or UNVERIFIED per dimension. Do not average the
labels into a claimed Google score.
4.3 Semantic gap and refresh
- Build an entity/attribute/question map from the query set, SERP, customer
language, and first-party data.
- Add missing concepts only when they help the task; avoid synonym padding.
- Preserve URLs with earned value unless evidence favors merge/redirect.
- Refresh changed facts and examples; do not change dates without material work.
- Prefer one useful new fact, example, test, or decision aid over generic length.
5. Keyword and demand research
Treat keywords as evidence of demand and language, not insertion targets.
Intent-first table
| Query cluster |
Intent |
Audience/job |
Evidence source |
Business value |
Existing URL |
Action |
| ... |
learn/compare/buy/navigate |
... |
GSC/SERP/research |
H/M/L |
... |
keep/create/merge |
Workflow
- Seed from products, problems, customer interviews, site search, sales/support,
GSC, paid-search terms, and competitor/category language.
- Expand variants by task, entity, comparison, alternative, constraint, locale,
and funnel stage.
- Inspect current SERPs by locale/device/date; classify dominant intent and
result type.
- Cluster by shared intent and answer requirements, not an arbitrary overlap
threshold.
- Map one canonical page or page family per cluster; identify cannibalization.
- Prioritize by evidenced demand, conversion fit, strategic value, feasibility,
authority gap, and maintenance cost.
- Label volume, difficulty, and trend data with provider, geography, and date.
【Proposal】 If comparable volume data is unavailable, use ordinal H/M/L inputs
with written reasons. Do not invent numeric opportunity scores.
6. Structured data / JSON-LD
Structured data describes visible truth; it does not guarantee rankings, rich
results, or AI citations.
Type selection
- Use the most specific Schema.org type supported by page content.
Organization/LocalBusiness: real identity, sameAs, contact, location.
WebSite: site identity; add actions only when the action truly works.
Article/NewsArticle: author, dates, headline, image, publisher.
Product/SoftwareApplication: real product attributes; include Offer only
with current price, currency, availability, and URL.
BreadcrumbList: visible navigational hierarchy.
VideoObject, Event, JobPosting, Recipe, or other types only when page
content and current provider eligibility support them.
- FAQ or How-to markup may describe visible content, but never promise a Google
rich result or AI extraction; check current provider documentation first.
Validation
- Extract every JSON-LD block from raw and rendered HTML.
- Parse as JSON; resolve duplicate/conflicting entities and stable
@id links.
- Match every claim to visible content and current facts.
- Validate vocabulary with Schema.org tooling.
- Validate provider eligibility with the provider's current rich-result docs
and test when that provider supports the page's structured-data type.
- Save errors/warnings with URL and timestamp. Never “validate mentally.”
- No fake
price: 0 on paid products; omit Offer if price unknown.
7. Performance / Core Web Vitals
Use field data at the 75th percentile when available; use lab data to diagnose.
Google's current “good” thresholds are LCP ≤2.5 s, INP ≤200 ms, and CLS ≤0.1
(official reference). Re-check the reference
at runtime.
Record:
- source: CrUX, PageSpeed Insights, GSC CWV, RUM, or lab;
- URL-level or origin-level aggregation;
- device, locale, connection/CPU profile, sample window, and collection date;
- field status and lab reproduction separately.
Prioritize the measured bottleneck:
- LCP: server response, critical resource discovery, image/font priority.
- INP: long tasks, main-thread work, hydration, event handlers.
- CLS: unsized media/ads/embeds, injected content, font swaps.
Do not claim a pass from one Lighthouse run. Re-test representative templates
under fixed conditions and verify field movement after the reporting window.
8. GEO / AEO — generative recommendation and citation
GEO here means increasing the probability that a brand, product, fact, or page
is retrieved, mentioned, recommended, or cited in a generated answer. It is not
“traditional SEO plus an llms.txt file.”
8.1 Source-selection model
Diagnose three distinct paths:
| Path |
What can happen |
Observable levers |
Measurement limit |
| Model memory |
A model recalls learned associations without live retrieval |
durable entity/fact consistency; broad legitimate recognition |
training corpus and attribution usually unobservable |
| Search/retrieval |
Engine retrieves an index or live sources for the prompt |
crawl/index eligibility, relevance, answer fit, authority, freshness |
retrieval/reranking systems are platform-specific black boxes |
| User-directed fetch |
A user action causes the engine to fetch a page |
fetcher access, page availability, renderability, answer clarity |
one fetch does not imply future indexing or citation |
This separation is supported by vendors publishing different training, search,
and user-fetch agents (OpenAI and Anthropic official crawler docs). Never use a
bot hit, training opt-in, or organic rank as proof of generated inclusion.
For every observed answer, record whether search/retrieval was visibly enabled.
If the interface does not disclose it, mark the path UNKNOWN.
8.2 Platform matrix
Verify current behavior at execution time; interfaces and controls change.
| Surface |
Safe current statement |
Primary control/evidence |
Do not infer |
| ChatGPT search |
OAI-SearchBot is used to surface sites in ChatGPT search; its control is independent of GPTBot training control |
robots policy, official IP ranges, answer citations |
allowed means ranked/cited |
| Google AI Overviews / AI Mode |
Standard SEO eligibility applies; Google says there are no extra requirements or special optimizations |
Googlebot; indexability; nosnippet, data-nosnippet, max-snippet, noindex |
Google-Extended controls these Search features |
| Gemini outside Search |
Product and grounding behavior must be tested in the named Gemini surface |
visible citations, official product docs; Google-Extended for specified training/grounding uses |
Google Search behavior equals every Gemini product |
| Perplexity |
PerplexityBot surfaces and links sites in search results; Perplexity-User serves user-triggered requests |
robots policy for bot, official IP lists, visible citations |
crawler access guarantees recommendation |
| Claude search |
Claude-SearchBot supports search quality; Claude-User supports user-directed retrieval; ClaudeBot is for potential training data |
robots policy, official docs, visible citations |
ClaudeBot is the citation crawler |
| Doubao |
No first-party public crawler/citation control was verified for this V6 research pass |
repeatable manual tests, visible citations, consented referral/log data |
Bytespider purpose or access guarantees Doubao inclusion |
Official references:
8.3 Crawler and retrieval eligibility
- Read robots groups for the exact agent; account for precedence and each
subdomain.
- Check page status, canonical, robots meta,
X-Robots-Tag, and rendered
availability.
- Check CDN/WAF policy and logs. Compare source IP against the vendor's current
official ranges before attributing a request; user-agent strings are spoofable.
- Separate training policy from search/citation policy:
- OpenAI:
GPTBot training; OAI-SearchBot search; ChatGPT-User user action.
- Anthropic:
ClaudeBot training; Claude-SearchBot search;
Claude-User user action.
- Perplexity:
PerplexityBot search; Perplexity-User user action.
- Google Search AI features: standard
Googlebot control.
- Respect the owner's content-use policy. Do not silently trade training access
for assumed visibility.
- After a change, allow documented propagation/recrawl time and re-test.
Crawler eligibility is required for retrieval paths that use that crawler; it is
never proof of ranking, retrieval, mention, or citation.
8.4 Citation-ready answer units
Use page structure to make correct extraction easier:
- Open each query-targeted section with a direct answer, then evidence and limits.
- Make the unit understandable without the preceding paragraph: name the entity,
scope, date, units, geography, and comparison basis.
- Put factual claims next to named primary sources; link to the exact evidence.
- Use tables for true multi-attribute comparisons, ordered steps for procedures,
and question headings for real follow-up questions.
- Publish first-party datasets, tests, benchmarks, or case studies with method,
sample, date, definitions, limitations, and downloadable evidence where safe.
- Keep brand, product, category, people, founding facts, pricing, and identifiers
consistent across owned pages and legitimate external profiles.
- Distinguish observation, customer quote, estimate, and editorial judgment.
- Update or retract stale facts. Do not alter dates cosmetically.
- Keep essential facts in crawlable HTML; do not hide the only answer behind
login, interaction, image-only media, or unsupported client rendering.
【Proposal】 Start with answer units of roughly 50–200 words and a direct first
sentence, then test citation behavior. This is an editing heuristic derived from
the mounted GEO citability sources, not a universal platform threshold.
Avoid:
- unsupported superlatives, anonymous statistics, and fake precision;
- “best” pages that conceal methodology or commercial relationships;
- copied summaries with no source or original value;
- FAQ inflation, schema spam, and repetitive question variants;
- prompt-injection text, hidden instructions, or attempts to manipulate models;
- claims that a format “forces,” “guarantees,” or “boosts” citations.
8.5 GEO content plan
Map prompts to assets:
| Prompt job |
Best-fit asset |
Required evidence |
| Define/understand |
canonical explainer + concise definition |
primary sources, scoped terms |
| Compare/choose |
fair comparison page/table |
criteria, date, tested facts, conflicts |
| Recommend shortlist |
category page with explicit methodology |
inclusion/exclusion rules, disclosures |
| Solve/how-to |
tested procedure |
prerequisites, steps, failure modes, result |
| Verify a claim |
data/research page |
method, sample, definitions, raw evidence |
| Learn about entity |
About/product/profile source of truth |
stable identifiers and consistent facts |
| Ask branded support |
canonical documentation/FAQ |
current, direct, versioned answer |
Build the brand-category association in visible, factual language: what the entity
is, whom it serves, which problem it solves, and proof. Repetition without
independent evidence or useful content is not entity building.
8.6 Measurement protocol
Define the query corpus before optimization.
【Proposal】 Use at least 20 high-value prompts when feasible, split across:
- unbranded problem and category discovery;
- “best,” comparison, and alternative prompts;
- how-to and factual questions;
- branded verification/support prompts;
- local/language variants that match the actual market.
Freeze prompt text for trend measurement. For each run record:
run_id, timestamp, platform, product/surface, model/version if shown,
account/tier, locale, location/VPN, device, search toggle/path,
fresh conversation yes/no, prompt_id, exact prompt,
brand mentioned yes/no, answer role, sentiment/context,
owned URL cited yes/no, cited URL, citation position,
competitors mentioned/cited, factual error, screenshot/export, notes
【Proposal】 Run each prompt three times per platform in fresh sessions when
budget permits. Report run-level variance; do not select the best response.
Calculate with explicit denominators:
- A valid run is a planned run that returns an inspectable answer without an
account, tool, policy, network, or collection error.
- A citation-capable run is a valid run on a surface/configuration documented
or visibly configured to return web sources. Keep runs with zero citations in
this denominator.
- Generative appearance rate = runs mentioning the brand / all valid runs.
- AI citation rate = runs citing an owned URL / citation-capable runs.
- Citation-to-mention rate = brand-mention runs with owned citation /
brand-mention runs.
- Share of cited voice = owned citation occurrences / all tracked
category-entity citation occurrences.
- Source coverage = distinct owned URLs cited / priority owned URLs tested.
- Citation accuracy rate = correct supported owned citations / audited owned
citations.
Report every metric as numerator/denominator (rate). If the denominator is zero,
report N/A and the zero-denominator reason. Report each platform separately.
Combine only with predeclared weights tied to audience usage; show the unweighted
data beside the composite.
8.7 Evidence ladder and attribution
| Level |
Evidence |
Safe claim |
| E0 |
no direct observation |
hypothesis only |
| E1 |
crawler/config inspection |
eligible or blocked at tested layer |
| E2 |
one generated-answer observation |
appeared/did not appear in that run |
| E3 |
repeated fixed-corpus runs |
observed rate for platform/window/config |
| E4 |
repeated tests plus logs/referrals/conversions |
association across retrieval and business outcomes |
- Bot hits prove requests, not citation.
- AI referral traffic undercounts activity when referrers are absent or altered.
- A citation proves one answer used a source, not that every statement came from it.
- Manual tests are snapshots, not population estimates.
- Never claim causality from a before/after change without a design that rules out
query, model, index, competitor, seasonality, and personalization changes.
8.8 Experiment design
【Proposal】 Run controlled GEO experiments:
- Choose matched page/query groups and save a pre-change baseline.
- Change one class of lever: answer units, evidence, original data, entity facts,
technical access, or internal discovery.
- Preserve prompt corpus, platform settings, locale, and collection protocol.
- Record crawl/index lag; do not start the post window before the changed page is
observable.
- Compare run-level rates and uncertainty, plus organic, referral, and conversion
guardrails.
- Keep, revise, or revert based on evidence. Store counterexamples.
Treat local GEO-tool scores as readiness diagnostics, not outcome validation.
8.9 GEO vs SEO priority
| Condition |
Priority |
| Site cannot be crawled, rendered, indexed, or canonically understood |
SEO foundation first |
| Google AI Overviews/AI Mode is the target |
Standard Google SEO eligibility first; add answer/evidence improvements |
| Audience uses answer engines for category discovery/comparison |
GEO experiment high |
| Site has unique data/expertise but weak extractability |
GEO content high |
| Demand is navigational, local-map, or transaction-led with weak AI usage |
SEO/local/CRO high |
| Brand is mentioned but facts/citations are wrong |
GEO entity/source-of-truth high |
| No baseline or query corpus exists |
Measurement first |
Default to shared work—technical access, useful pages, verifiable facts, and clear
structure—then fund platform-specific work only when measured opportunity warrants it.
8.10 llms.txt
llms.txt is a community proposal, not a universal ranking or citation directive.
The mounted GEO sources support checking and experimenting with it, but this V6
found no official OpenAI, Google Search, Perplexity, Anthropic, or Doubao claim
that publishing it improves inclusion.
【Proposal】 Add it only when:
- maintenance ownership exists;
- links are canonical, public, and useful;
- it does not expose private or unpublished material;
- it is measured as an experiment.
Its absence is not an SEO/GEO defect. Never give it readiness points merely for
existing.
8.11 GEO readiness rubric
【Proposal】 Score each dimension 0=blocked/absent, 1=partial/unverified,
2=working with evidence:
- retrieval eligibility;
- index/render availability;
- answer-unit clarity;
- claim evidence and first-party value;
- entity consistency;
- query-to-asset coverage;
- repeated platform measurement;
- attribution/business-outcome linkage.
Report the eight values separately. A total is an internal triage aid, not a
prediction of citation probability.
9. Backlinks and external authority
Profile analysis
- Use GSC links plus an authorized provider when available; state coverage limits.
- Review referring domains/pages, relevance, editorial context, destination,
anchor distribution, acquisition trend, and lost links.
- Separate manipulative links from normal scraper/spam noise.
- Check unlinked brand mentions and incorrect citations that can be reclaimed.
Risk handling
Do not call links “toxic” from a provider score alone. Remove or disavow only with
documented evidence of manipulative activity and material risk, such as a manual
action or known paid-link scheme. Preserve a decision log.
Earn links and citations
- original datasets, tools, benchmarks, and transparent methods;
- expert contributions and primary-source commentary;
- genuinely useful comparison/reference pages;
- broken-link replacement where the asset is a true substitute;
- correction outreach for inaccurate facts or broken citations.
No PBNs, paid-link laundering, mass guest-post spam, or automated outreach sludge.
10. Analytics wiring (GSC + GA4)
GSC
- Distinguish verification-token presence, current verified ownership, API/UI
access, and data availability.
- A meta/DNS token is evidence of configuration, not proof of current access.
- Domain properties may cover subdomains; do not demand a subdomain meta token.
- Record property type, date range, search type, country/device filters, and
anonymization/row-limit caveats.
GA4
- Distinguish source-code presence, network request, DebugView/realtime event,
stream configuration, and report/API access.
- On Liz house properties, preserve
G-TXVLTJJ878 exactly. Outside section 13,
call it a measurement ID.
- Test consent state, duplicate tags, cross-domain rules, SPA navigation, and
conversion/key-event semantics when in scope.
- Never infer collected users or conversions from a tag string alone.
Reporting shape
Show clicks, impressions, CTR, average position, organic sessions/users, engaged
sessions, conversions/key events, landing page, query cluster, locale/device, and
comparison window only when the source is accessible. Otherwise mark UNVERIFIED.
11. SERP and competitor analysis
Capture exact query, locale, device, date/time, personalization state, and source.
- Record organic results, local/video/image/shopping/news features, featured
snippets, discussions, AI Overviews/AI Mode, and cited sources.
- Compare intent coverage, evidence, entity authority, information gain, format,
freshness, UX, links, and conversion path—not word count alone.
- Separate true business competitors, organic competitors, and AI-cited sources.
- Identify answer/source gaps the site can satisfy honestly.
- Re-check volatile SERPs; one observation is not a stable market fact.
Comparison/alternative parity floor
【Proposal】 Before publishing a comparison or alternative page, open the
competitor's same-type URL in Ahrefs/Semrush (or authorized equivalent). Require
parity or better on: traffic estimate, keyword coverage, content depth, evidence
density. Else BLOCK publish. Qualitative SERP notes without this floor check
are incomplete.
Output an opportunity table with query cluster, observed surface, winning source,
why it may satisfy the task, evidence gap, feasible asset, impact, effort, and
confidence.
12. Rank, AI-visibility, and outcome tracking
Tracking setup
- Freeze keyword/query corpus, landing-page mapping, locale, device, and platform.
- Track organic rank/landing URL beside generated mentions/citations.
- Annotate releases, migrations, campaigns, algorithm events, model/product
changes, and known outages.
- Segment branded/unbranded, intent, market, template, and conversion value.
- Keep raw observations so metric definitions can be recomputed.
Page-level result ownership
Assign one owner per priority URL. Success = that URL's GSC impressions + CTR
(and mapped conversions when available), not “we shipped SEO.” Unowned URLs stay
UNVERIFIED for outcome claims.
Decision rules
【Proposal】 Predeclare stop/continue rules per initiative. Example:
- continue when leading evidence improves without technical or conversion harm;
- investigate when the canonical URL changes, variance spikes, or citations become
inaccurate;
- stop when the site has no publish path, no audience fit, no measurable surface,
or maintenance cost exceeds documented value.
Do not call a SERP permanently occupied or an initiative causal from two points.
13. House rules (Liz) — verbatim-accurate
Follow exactly on house properties:
- GA4 shared property ID
G-TXVLTJJ878 — never change/remove.
- No brand/asset changes: logos, fonts, palette, icons.
- Copy pain-first, EN default; no AI-slop filler.
- Minimal change; never rewrite whole files; no new dependencies.
- Never fabricate GSC verification tokens; REPORT presence only.
- Don't touch
projects/_templates/**, docs/AUDIT.md, payment code.
- red-flowers (reading site) is private-by-design (noindex + WAF CN) — never SEO-refine it.
- Verification:
git status -s matches claimed files (leave pre-existing dirt unstaged)
grep GA4 ID count on touched HTML
- sitemap
<loc> count sane; robots head sane
- Stage selectively; never
git add -A
Extra house semantics (operational)
- Apply these rules only after confirming a Liz house property.
- Preserve pre-existing changes and report them separately.
- Treat production bindings as strong host-intent evidence; investigate conflicts
before migration.
- Do not stage or commit unless explicitly requested.
14. Execution algorithm
- Scope gate: objective, target, access, authorization, exclusions, sample.
- Snapshot: git/worktree state if local; date, host, redirects, robots, sitemaps.
- Technical gate: reach/render/crawl/index/canonical before content polishing.
- Analytics gate: identify observable and unverified layers.
- Page/content/schema/CWV audit on representative templates.
- Demand/SERP/competitor work when query evidence is in scope.
- GEO audit:
- platform and retrieval-path matrix;
- crawler/WAF eligibility;
- query-to-asset and answer-unit review;
- fixed-corpus baseline;
- citation, mention, accuracy, and source metrics.
- Backlink/external-authority review when data exists.
- Prioritize dependency first, then impact × confidence ÷ effort; keep the raw
dimensions visible rather than pretending the quotient is precise.
- Make only authorized minimal changes.
- Verify changed files, behavioral checks, analytics preservation, and
regressions.
- Report method, evidence, unknowns, rollback, owners, and next measurement.
Severity rubric
| Severity |
Definition |
Examples |
| P0 |
production, legal, privacy, or discovery catastrophe requiring immediate action |
public private data; sitewide accidental noindex; destructive redirect |
| P1 |
material discovery, trust, or conversion blocker |
broken canonical migration; key templates unavailable; false product facts |
| P2 |
bounded opportunity or quality defect |
weak answer units; missing contextual links; incomplete evidence |
| P3 |
optional experiment or polish |
llms.txt trial; low-value formatting refinement |
Severity depends on affected scope and business impact. Missing optional schema or
an AI crawler blocked by policy is not automatically P0/P1.
15. Commands and evidence cheatsheet
Set variables; do not paste angle-bracket placeholders into a shell:
host="example.com"
base="https://${host}"
# Final GET status, effective URL, redirects, timing.
curl -sS -L --max-redirs 10 --connect-timeout 10 --max-time 30 \
-o /dev/null -w 'status=%{http_code} url=%{url_effective} redirects=%{num_redirects} time=%{time_total}\n' \
"${base}/"
# Full response headers for one GET.
curl -sS --connect-timeout 10 --max-time 30 -D - -o /dev/null "${base}/"
# Robots and sitemap content. Inspect complete files or save outside protected paths.
curl -sS --connect-timeout 10 --max-time 30 "${base}/robots.txt"
curl -sS --connect-timeout 10 --max-time 30 "${base}/sitemap.xml"
# Compare responses to user-agent strings; this does not authenticate vendor bots.
curl -sS -A "OAI-SearchBot" -o /dev/null -w '%{http_code}\n' "${base}/"
curl -sS -A "PerplexityBot" -o /dev/null -w '%{http_code}\n' "${base}/"
# House verification.
git status -s
grep -oF "G-TXVLTJJ878" "path/to/touched.html" | wc -l
Parse XML/HTML with a real parser when correctness matters. Do not use line-count
grep as an XML element count, regex as a complete HTML parser, HEAD as proof of GET
behavior, or source presence as proof that analytics fires.
Never claim GSC/GA4 UI state, index inclusion, bot identity, or generated citation
without direct evidence.
16. Output template
# SEO/GEO report — [host] — [date]
## Verdict
SHIP | FIX | BLOCK | KILL
## Method and limits
- Scope/sample:
- Access:
- Locale/device/window:
- Unknowns:
## P0 / P1 / P2
| Priority | Finding | Evidence | Scope | Confidence | Impact | Effort | Owner |
## Technical / on-page / content / schema / CWV
- Observed:
- Expected:
- Evidence:
- Action:
## GEO
- Target platforms and retrieval paths:
- Crawler/WAF eligibility:
- Query corpus and run protocol:
- Valid runs and run-level variance:
- Appearance rate — numerator/denominator:
- AI citation rate — numerator/denominator:
- Citation-to-mention rate — numerator/denominator:
- Share of cited voice — numerator/denominator:
- Source coverage — numerator/denominator:
- Citation accuracy — numerator/denominator:
- Cited URLs and competitors:
- Answer/evidence/entity gaps:
## Analytics and outcomes
- GSC verification/access:
- GA4 G-TXVLTJJ878 on house properties: intact|missing|n/a
- Organic/referral/conversion evidence:
## Changes and verification
- Files/actions:
- Tests:
- Rollback:
## Next actions
1. [dependency-first action, owner, due/measurement]
17. Failure patterns (do not repeat)
| Pattern |
Instead |
| Optimize content on noindex URLs |
Fix indexability first |
| Sitemap of duplicate/wrong-host URLs |
Fix canonical host, then sitemap |
| DNS-based canonical migration |
Read bindings; then migrate |
| Fake GSC meta to clear a checklist |
Report only |
| price:0 JSON-LD |
Real price or omit Offer |
| AI-slop refresh |
Pain-first specifics |
git add -A |
Explicit paths |
| SEO-refine red-flowers |
Skip |
| Demand subdomain GSC tokens under domain property |
Mark n/a |
| Ignore CF managed AI bot blocks |
Flag to operator |
V6 additions:
| Pattern |
Instead |
Treat robots Disallow as noindex |
Test crawling and indexing controls separately |
| Treat allowed AI bot as citation proof |
Report eligibility only; run fixed-corpus tests |
| Treat bot user-agent as authenticated |
Verify current official IP ranges and logs |
Treat GPTBot or ClaudeBot as search crawlers |
Separate training, search, and user-fetch agents |
Treat Google-Extended as AI Overview control |
Use Googlebot/Search preview controls per official docs |
Require or score llms.txt as a standard |
Optional measured experiment only |
| Publish generic FAQ/schema for “GEO” |
Build useful answer units backed by evidence |
| Cherry-pick one favorable AI answer |
Report all valid fixed-protocol runs |
| Merge platform rates without denominators |
Report platform/surface metrics separately |
| Claim causality from before/after |
Control protocol and state confounders/limits |
| Ship decision page as listicle |
Pass format-vs-intent gate; rewrite as comparison/alternative |
| Publish below competitor same-type parity |
Meet Ahrefs/Semrush parity floor first |
| Report “did SEO” without URL owner/metrics |
Assign page owner; track GSC impressions + CTR per URL |
18. Source synthesis and claim policy
Use this precedence:
- current first-party platform documentation for product/crawler controls;
- primary research for measured effects;
- direct site, log, analytics, and fixed-protocol observations;
- mounted practitioner tools as hypotheses and executable heuristics;
【Proposal】 for unvalidated thresholds or internal decision rules.
Mounted source synthesis:
- GEO Optimizer: audit model, crawler/access diagnostics, Princeton GEO method
summary, measurement/drift concepts.
- geo-skills: answer-block, citability, platform-test, crawler, and com
…(truncated)
1---2name: seo-master3description: Audits, improves, and reports evidence-backed SEO and generative-engine visibility across technical, on-page, content, schema, performance, GEO/AEO, links, analytics, SERP, and tracking workflows. Use when a user requests an SEO/GEO audit, AI citation or brand-mention analysis, search remediation, content optimization, or a prioritized organic-discovery plan.4---56# seo-master78Evidence first. Separate observed facts, inferences, and proposals. Never present9technical eligibility as a promise of ranking, citation, recommendation, or10traffic.1112## 0. Quick start1314Use the smallest workflow that answers the request:1516```text17Audit [site] for technical SEO + GEO. Read-only. Sample [N] URLs.18Fix on-page + truthful schema for [URL]. Preserve brand and analytics.19Measure AI visibility for [brand] across [platforms] with [query set].20Compare SEO vs GEO opportunities for [market, locale, device].21```2223Before work, state:2425- target host, locale, audience, site type, and conversion goal;26- read-only audit or authorized edits;27- access available: repository, browser, logs, GSC, GA4, rank data, AI tools;28- sample or full crawl, URL cap, and time window;29- house property: yes/no. Apply section 13 only when yes.3031If required access is missing, continue with observable evidence and label the32unobservable layer `UNVERIFIED`; do not fill gaps with estimates.3334## 1. Scope & when-not-to3536### Do run when3738- The site has a public discovery or acquisition goal.39- The user asks about rankings, citations, AI mentions, crawlability, content,40 schema, performance, links, analytics, or search reporting.41- A deployable path exists and the requested access is authorized.4243### Do NOT run (anti-patterns)4445| Situation | Action |46|-----------|--------|47| Private-by-design sites (e.g. red-flowers: noindex + WAF CN) | Skip entirely — never SEO-refine |48| Dead / abandoned projects with no publish path | Report kill criteria; do not polish |49| Spam tactics (PBNs, cloaking, link farms, doorway pages) | Refuse; propose white-hat alternative |50| Keyword stuffing / hidden text | Strip; rewrite pain-first |51| Fabricating GSC verification tokens / fake metrics | REPORT presence only; never invent |52| Brand/asset redesign under SEO pretext | Out of scope (see House rules) |53| `projects/_templates/**`, `docs/AUDIT.md`, payment code | Do not touch |5455### Scope questions5657Ask only questions that materially alter execution:58591. Which canonical host, markets, languages, and priority conversions?602. Audit only, recommendations, or code/content changes?613. Which evidence sources are authorized and available?624. Which AI surfaces matter to the audience: ChatGPT, Gemini, Google AI63 Overviews/AI Mode, Perplexity, Doubao, Claude, or another engine?645. What must remain untouched?6566Record unresolved choices as assumptions. Never let a missing optional data67source block checks that can be performed safely.6869## 2. Technical SEO audit7071Audit in dependency order: reachability → renderability → crawlability →72indexability → canonicalization → discovery → internal architecture.7374### 2.1 Evidence sample7576Include at minimum, when in scope:7778- homepage;79- one URL per indexable template;80- priority conversion pages;81- one paginated, filtered, localized, and JavaScript-heavy URL where present;82- redirect, 404, soft-404 candidate, and canonical duplicate samples;83- URLs reported by GSC or logs, if access exists.8485For every finding record URL, timestamp, method, observed value, expected value,86evidence, confidence, impact, and affected scope. Do not extrapolate a sampled87finding to the whole site without a crawl or corroborating data.8889### 2.2 Reachability, rendering, and crawlability9091- Record DNS/TLS failures, final GET status, redirect chain, latency, and host.92- Compare raw HTML with rendered DOM for canonical, robots meta, headings,93 links, primary copy, and structured data.94- Parse `robots.txt` for the tested user agent. A `Disallow` controls crawling;95 it is not a page-level `noindex`.96- Check WAF/CDN responses separately from robots policy. A spoofed user-agent97 request does not authenticate a vendor crawler.98- Test robots and sitemap files on every host/subdomain that serves sampled URLs,99 redirects to or from the canonical host, or appears in hreflang/sitemaps.100- Flag infinite spaces: parameters, faceted paths, calendars, session IDs,101 internal search, and duplicate sort orders.102103### 2.3 Indexability104105For each sampled URL classify:106107`INDEXABLE`, `EXCLUDED-AS-INTENDED`, `BLOCKED-BY-TECHNICAL-CONFLICT`, or108`UNVERIFIED-IN-INDEX`.109110Check:111112- final status is indexable content, not redirect/error/soft error;113- `meta robots` and `X-Robots-Tag`;114- whether robots permits the tested crawler to fetch page-level directives; if115 blocked, classify the directive's crawler visibility as `UNVERIFIED`;116- canonical target status, indexability, locale, and content equivalence;117- GSC URL Inspection/indexing reports when authorized;118- orphan status and internal-link discoverability.119120Never claim “indexed” from a 200 response, sitemap entry, or `site:` query alone.121122### 2.4 Canonical, host, redirects, and hreflang123124- Treat owner intent, production bindings, redirects, canonicals, sitemaps,125 internal links, and analytics configuration as evidence. Resolve conflicts;126 do not assume one signal is universally authoritative.127- Require one-hop permanent redirects for retired canonical URLs where safe.128- Reject loops, chains, protocol/host oscillation, mass redirects to irrelevant129 pages, and query loss that changes meaning.130- Require self-referential canonical on canonical pages unless a documented131 cross-domain or variant strategy says otherwise.132- For localized equivalents, require reciprocal valid language-region codes,133 indexable canonical targets, and `x-default` only when a true fallback exists.134135### 2.5 Sitemaps and architecture136137- Parse XML; do not count matching lines.138- Validate sitemap-index recursion, XML syntax, HTTP status, canonical host,139 indexable URLs, `lastmod` truthfulness, and URL limits.140- Exclude redirects, errors, duplicates, blocked/noindex URLs, internal search,141 and noncanonical parameters.142- Confirm priority pages receive descriptive internal links from the topic,143 product, category, or locale hubs that contain them in the site architecture.144- Report crawl depth as observed distribution, not a universal pass/fail number.145146### 2.6 Technical scorecard147148Use statuses, not a pseudo-scientific score:149150| Layer | Status | Evidence | Scope | Confidence | Next action |151|---|---|---|---|---|---|152| Reach/render | PASS/FIX/BLOCK/UNVERIFIED | URL/log | sample/all | H/M/L | ... |153| Crawl/index | PASS/FIX/BLOCK/UNVERIFIED | directive/GSC | sample/all | H/M/L | ... |154| Canonical/redirect | PASS/FIX/BLOCK/UNVERIFIED | chain/HTML | sample/all | H/M/L | ... |155| Sitemap/links | PASS/FIX/BLOCK/UNVERIFIED | parsed XML/crawl | sample/all | H/M/L | ... |156157## 3. On-page SEO158159Optimize the page for one primary intent and the jobs needed to satisfy it.160161### Format-vs-intent gate (first filter)162163Classify primary intent, then reject format mismatch before copy polish.164165| Intent | Fit format | Reject |166|---|---|---|167| Explore / shortlist | Listicle, category roundup | Decision page forced into “N alternatives” |168| Decide / choose (X vs Y, alt to X) | Comparison or alternative page | Listicle that never answers “which should I pick” |169| Learn / how-to | Explainer, procedure | Comparison shell with no decision criteria |170171Decision-intent pages fail this gate if they are listicles. Fix format first.172173### Comparison/alternative page skeleton174175- [ ] H1 states the decision question (e.g. `How is X better than Y?`)176- [ ] Above-the-fold: direct answer + proof hooks177- [ ] Real-data band (benchmarks, pricing, limits — sourced)178- [ ] Evidence-dense differentiators; product video/screenshots where claims need demo179- [ ] Fair criteria table; disclose conflicts1801811. Match title, visible heading, opening, sections, media, and CTA to intent.1822. Make titles and descriptions specific and nonduplicative; judge truncation183 from rendered SERPs when available, not a fixed character quota.1843. Use a clear page topic and logical heading hierarchy. Do not enforce “exactly185 one H1” as a ranking rule; flag confusing document structure.1864. Put the useful answer before biography, history, or promotional filler.1875. Add internal links where they advance the task, using descriptive anchors.1886. Give images meaningful alternatives when informative; use empty `alt` for189 decorative images.1907. Check visible copy against canonical, schema, Open Graph, and product facts.1918. Remove doorway duplication, hidden text, keyword stuffing, and templated192 filler.193194Output: `KEEP`, `REWRITE`, `MERGE`, `REDIRECT`, `NOINDEX`, or `DELETE` for each195page, with evidence and destination where applicable.196197## 4. Content & E-E-A-T198199E-E-A-T is a qualitative quality lens, not a published Google score.200201### 4.1 Evidence gate202203For consequential claims, require:204205- a named source, first-party evidence, or clearly described methodology;206- author/reviewer identity and demonstrated competence for the claim where an207 incorrect answer could affect health, safety, rights, or material decisions;208- publication and material-update dates when freshness matters;209- ownership, editorial, contact, correction, privacy, and commercial disclosure210 required by the site's function, jurisdiction, risk, or commercial incentives;211- claims that match the cited source and do not exceed it.212213For YMYL topics, escalate unsupported medical, legal, financial, or safety advice.214Never create credentials, reviews, test results, customers, or outcomes.215216### 4.2 CORE-EEAT review217218`【Proposal】` Use this internal rubric for diagnosis only:219220| Dimension | Test |221|---|---|222| Coverage | Does the page complete the user's task and answer necessary follow-ups? |223| Originality | Is there first-hand evidence, analysis, data, or a useful synthesis? |224| Recency | Are time-sensitive claims reviewed and dated? |225| Evidence | Can a reviewer trace consequential claims to reliable sources? |226| Experience | Are first-hand methods, constraints, and outcomes demonstrated? |227| Expertise | Does the author/reviewer demonstrate credentials or first-hand competence for the claim? |228| Authority | Do independent sources in the same field recognize the entity or work? |229| Trust | Are identity, incentives, limitations, and corrections transparent? |230231Assign `PASS`, `WEAK`, `FAIL`, or `UNVERIFIED` per dimension. Do not average the232labels into a claimed Google score.233234### 4.3 Semantic gap and refresh235236- Build an entity/attribute/question map from the query set, SERP, customer237 language, and first-party data.238- Add missing concepts only when they help the task; avoid synonym padding.239- Preserve URLs with earned value unless evidence favors merge/redirect.240- Refresh changed facts and examples; do not change dates without material work.241- Prefer one useful new fact, example, test, or decision aid over generic length.242243## 5. Keyword and demand research244245Treat keywords as evidence of demand and language, not insertion targets.246247### Intent-first table248249| Query cluster | Intent | Audience/job | Evidence source | Business value | Existing URL | Action |250|---|---|---|---|---|---|---|251| ... | learn/compare/buy/navigate | ... | GSC/SERP/research | H/M/L | ... | keep/create/merge |252253### Workflow2542551. Seed from products, problems, customer interviews, site search, sales/support,256 GSC, paid-search terms, and competitor/category language.2572. Expand variants by task, entity, comparison, alternative, constraint, locale,258 and funnel stage.2593. Inspect current SERPs by locale/device/date; classify dominant intent and260 result type.2614. Cluster by shared intent and answer requirements, not an arbitrary overlap262 threshold.2635. Map one canonical page or page family per cluster; identify cannibalization.2646. Prioritize by evidenced demand, conversion fit, strategic value, feasibility,265 authority gap, and maintenance cost.2667. Label volume, difficulty, and trend data with provider, geography, and date.267268`【Proposal】` If comparable volume data is unavailable, use ordinal H/M/L inputs269with written reasons. Do not invent numeric opportunity scores.270271## 6. Structured data / JSON-LD272273Structured data describes visible truth; it does not guarantee rankings, rich274results, or AI citations.275276### Type selection277278- Use the most specific Schema.org type supported by page content.279- `Organization`/`LocalBusiness`: real identity, sameAs, contact, location.280- `WebSite`: site identity; add actions only when the action truly works.281- `Article`/`NewsArticle`: author, dates, headline, image, publisher.282- `Product`/`SoftwareApplication`: real product attributes; include `Offer` only283 with current price, currency, availability, and URL.284- `BreadcrumbList`: visible navigational hierarchy.285- `VideoObject`, `Event`, `JobPosting`, `Recipe`, or other types only when page286 content and current provider eligibility support them.287- FAQ or How-to markup may describe visible content, but never promise a Google288 rich result or AI extraction; check current provider documentation first.289290### Validation2912921. Extract every JSON-LD block from raw and rendered HTML.2932. Parse as JSON; resolve duplicate/conflicting entities and stable `@id` links.2943. Match every claim to visible content and current facts.2954. Validate vocabulary with Schema.org tooling.2965. Validate provider eligibility with the provider's current rich-result docs297 and test when that provider supports the page's structured-data type.2986. Save errors/warnings with URL and timestamp. Never “validate mentally.”2997. **No fake `price: 0`** on paid products; omit Offer if price unknown.300301## 7. Performance / Core Web Vitals302303Use field data at the 75th percentile when available; use lab data to diagnose.304Google's current “good” thresholds are LCP ≤2.5 s, INP ≤200 ms, and CLS ≤0.1305([official reference](https://web.dev/articles/vitals)). Re-check the reference306at runtime.307308Record:309310- source: CrUX, PageSpeed Insights, GSC CWV, RUM, or lab;311- URL-level or origin-level aggregation;312- device, locale, connection/CPU profile, sample window, and collection date;313- field status and lab reproduction separately.314315Prioritize the measured bottleneck:316317- LCP: server response, critical resource discovery, image/font priority.318- INP: long tasks, main-thread work, hydration, event handlers.319- CLS: unsized media/ads/embeds, injected content, font swaps.320321Do not claim a pass from one Lighthouse run. Re-test representative templates322under fixed conditions and verify field movement after the reporting window.323324## 8. GEO / AEO — generative recommendation and citation325326GEO here means increasing the probability that a brand, product, fact, or page327is retrieved, mentioned, recommended, or cited in a generated answer. It is not328“traditional SEO plus an `llms.txt` file.”329330### 8.1 Source-selection model331332Diagnose three distinct paths:333334| Path | What can happen | Observable levers | Measurement limit |335|---|---|---|---|336| Model memory | A model recalls learned associations without live retrieval | durable entity/fact consistency; broad legitimate recognition | training corpus and attribution usually unobservable |337| Search/retrieval | Engine retrieves an index or live sources for the prompt | crawl/index eligibility, relevance, answer fit, authority, freshness | retrieval/reranking systems are platform-specific black boxes |338| User-directed fetch | A user action causes the engine to fetch a page | fetcher access, page availability, renderability, answer clarity | one fetch does not imply future indexing or citation |339340This separation is supported by vendors publishing different training, search,341and user-fetch agents (OpenAI and Anthropic official crawler docs). Never use a342bot hit, training opt-in, or organic rank as proof of generated inclusion.343344For every observed answer, record whether search/retrieval was visibly enabled.345If the interface does not disclose it, mark the path `UNKNOWN`.346347### 8.2 Platform matrix348349Verify current behavior at execution time; interfaces and controls change.350351| Surface | Safe current statement | Primary control/evidence | Do not infer |352|---|---|---|---|353| ChatGPT search | `OAI-SearchBot` is used to surface sites in ChatGPT search; its control is independent of `GPTBot` training control | robots policy, official IP ranges, answer citations | allowed means ranked/cited |354| Google AI Overviews / AI Mode | Standard SEO eligibility applies; Google says there are no extra requirements or special optimizations | `Googlebot`; indexability; `nosnippet`, `data-nosnippet`, `max-snippet`, `noindex` | `Google-Extended` controls these Search features |355| Gemini outside Search | Product and grounding behavior must be tested in the named Gemini surface | visible citations, official product docs; `Google-Extended` for specified training/grounding uses | Google Search behavior equals every Gemini product |356| Perplexity | `PerplexityBot` surfaces and links sites in search results; `Perplexity-User` serves user-triggered requests | robots policy for bot, official IP lists, visible citations | crawler access guarantees recommendation |357| Claude search | `Claude-SearchBot` supports search quality; `Claude-User` supports user-directed retrieval; `ClaudeBot` is for potential training data | robots policy, official docs, visible citations | `ClaudeBot` is the citation crawler |358| Doubao | No first-party public crawler/citation control was verified for this V6 research pass | repeatable manual tests, visible citations, consented referral/log data | `Bytespider` purpose or access guarantees Doubao inclusion |359360Official references:361362- [OpenAI crawlers](https://platform.openai.com/docs/bots)363- [Google AI features in Search](https://developers.google.com/search/docs/appearance/ai-features)364- [Google common crawlers / Google-Extended](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers)365- [Perplexity crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers)366- [Anthropic crawlers](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler)367368### 8.3 Crawler and retrieval eligibility3693701. Read robots groups for the exact agent; account for precedence and each371 subdomain.3722. Check page status, canonical, robots meta, `X-Robots-Tag`, and rendered373 availability.3743. Check CDN/WAF policy and logs. Compare source IP against the vendor's current375 official ranges before attributing a request; user-agent strings are spoofable.3764. Separate training policy from search/citation policy:377 - OpenAI: `GPTBot` training; `OAI-SearchBot` search; `ChatGPT-User` user action.378 - Anthropic: `ClaudeBot` training; `Claude-SearchBot` search;379 `Claude-User` user action.380 - Perplexity: `PerplexityBot` search; `Perplexity-User` user action.381 - Google Search AI features: standard `Googlebot` control.3825. Respect the owner's content-use policy. Do not silently trade training access383 for assumed visibility.3846. After a change, allow documented propagation/recrawl time and re-test.385386Crawler eligibility is required for retrieval paths that use that crawler; it is387never proof of ranking, retrieval, mention, or citation.388389### 8.4 Citation-ready answer units390391Use page structure to make correct extraction easier:392393- Open each query-targeted section with a direct answer, then evidence and limits.394- Make the unit understandable without the preceding paragraph: name the entity,395 scope, date, units, geography, and comparison basis.396- Put factual claims next to named primary sources; link to the exact evidence.397- Use tables for true multi-attribute comparisons, ordered steps for procedures,398 and question headings for real follow-up questions.399- Publish first-party datasets, tests, benchmarks, or case studies with method,400 sample, date, definitions, limitations, and downloadable evidence where safe.401- Keep brand, product, category, people, founding facts, pricing, and identifiers402 consistent across owned pages and legitimate external profiles.403- Distinguish observation, customer quote, estimate, and editorial judgment.404- Update or retract stale facts. Do not alter dates cosmetically.405- Keep essential facts in crawlable HTML; do not hide the only answer behind406 login, interaction, image-only media, or unsupported client rendering.407408`【Proposal】` Start with answer units of roughly 50–200 words and a direct first409sentence, then test citation behavior. This is an editing heuristic derived from410the mounted GEO citability sources, not a universal platform threshold.411412Avoid:413414- unsupported superlatives, anonymous statistics, and fake precision;415- “best” pages that conceal methodology or commercial relationships;416- copied summaries with no source or original value;417- FAQ inflation, schema spam, and repetitive question variants;418- prompt-injection text, hidden instructions, or attempts to manipulate models;419- claims that a format “forces,” “guarantees,” or “boosts” citations.420421### 8.5 GEO content plan422423Map prompts to assets:424425| Prompt job | Best-fit asset | Required evidence |426|---|---|---|427| Define/understand | canonical explainer + concise definition | primary sources, scoped terms |428| Compare/choose | fair comparison page/table | criteria, date, tested facts, conflicts |429| Recommend shortlist | category page with explicit methodology | inclusion/exclusion rules, disclosures |430| Solve/how-to | tested procedure | prerequisites, steps, failure modes, result |431| Verify a claim | data/research page | method, sample, definitions, raw evidence |432| Learn about entity | About/product/profile source of truth | stable identifiers and consistent facts |433| Ask branded support | canonical documentation/FAQ | current, direct, versioned answer |434435Build the brand-category association in visible, factual language: what the entity436is, whom it serves, which problem it solves, and proof. Repetition without437independent evidence or useful content is not entity building.438439### 8.6 Measurement protocol440441Define the query corpus before optimization.442443`【Proposal】` Use at least 20 high-value prompts when feasible, split across:444445- unbranded problem and category discovery;446- “best,” comparison, and alternative prompts;447- how-to and factual questions;448- branded verification/support prompts;449- local/language variants that match the actual market.450451Freeze prompt text for trend measurement. For each run record:452453```text454run_id, timestamp, platform, product/surface, model/version if shown,455account/tier, locale, location/VPN, device, search toggle/path,456fresh conversation yes/no, prompt_id, exact prompt,457brand mentioned yes/no, answer role, sentiment/context,458owned URL cited yes/no, cited URL, citation position,459competitors mentioned/cited, factual error, screenshot/export, notes460```461462`【Proposal】` Run each prompt three times per platform in fresh sessions when463budget permits. Report run-level variance; do not select the best response.464465Calculate with explicit denominators:466467- A **valid run** is a planned run that returns an inspectable answer without an468 account, tool, policy, network, or collection error.469- A **citation-capable run** is a valid run on a surface/configuration documented470 or visibly configured to return web sources. Keep runs with zero citations in471 this denominator.472- **Generative appearance rate** = runs mentioning the brand / all valid runs.473- **AI citation rate** = runs citing an owned URL / citation-capable runs.474- **Citation-to-mention rate** = brand-mention runs with owned citation /475 brand-mention runs.476- **Share of cited voice** = owned citation occurrences / all tracked477 category-entity citation occurrences.478- **Source coverage** = distinct owned URLs cited / priority owned URLs tested.479- **Citation accuracy rate** = correct supported owned citations / audited owned480 citations.481482Report every metric as `numerator/denominator (rate)`. If the denominator is zero,483report `N/A` and the zero-denominator reason. Report each platform separately.484Combine only with predeclared weights tied to audience usage; show the unweighted485data beside the composite.486487### 8.7 Evidence ladder and attribution488489| Level | Evidence | Safe claim |490|---|---|---|491| E0 | no direct observation | hypothesis only |492| E1 | crawler/config inspection | eligible or blocked at tested layer |493| E2 | one generated-answer observation | appeared/did not appear in that run |494| E3 | repeated fixed-corpus runs | observed rate for platform/window/config |495| E4 | repeated tests plus logs/referrals/conversions | association across retrieval and business outcomes |496497- Bot hits prove requests, not citation.498- AI referral traffic undercounts activity when referrers are absent or altered.499- A citation proves one answer used a source, not that every statement came from it.500- Manual tests are snapshots, not population estimates.501- Never claim causality from a before/after change without a design that rules out502 query, model, index, competitor, seasonality, and personalization changes.503504### 8.8 Experiment design505506`【Proposal】` Run controlled GEO experiments:5075081. Choose matched page/query groups and save a pre-change baseline.5092. Change one class of lever: answer units, evidence, original data, entity facts,510 technical access, or internal discovery.5113. Preserve prompt corpus, platform settings, locale, and collection protocol.5124. Record crawl/index lag; do not start the post window before the changed page is513 observable.5145. Compare run-level rates and uncertainty, plus organic, referral, and conversion515 guardrails.5166. Keep, revise, or revert based on evidence. Store counterexamples.517518Treat local GEO-tool scores as readiness diagnostics, not outcome validation.519520### 8.9 GEO vs SEO priority521522| Condition | Priority |523|---|---|524| Site cannot be crawled, rendered, indexed, or canonically understood | SEO foundation first |525| Google AI Overviews/AI Mode is the target | Standard Google SEO eligibility first; add answer/evidence improvements |526| Audience uses answer engines for category discovery/comparison | GEO experiment high |527| Site has unique data/expertise but weak extractability | GEO content high |528| Demand is navigational, local-map, or transaction-led with weak AI usage | SEO/local/CRO high |529| Brand is mentioned but facts/citations are wrong | GEO entity/source-of-truth high |530| No baseline or query corpus exists | Measurement first |531532Default to shared work—technical access, useful pages, verifiable facts, and clear533structure—then fund platform-specific work only when measured opportunity warrants it.534535### 8.10 `llms.txt`536537`llms.txt` is a community proposal, not a universal ranking or citation directive.538The mounted GEO sources support checking and experimenting with it, but this V6539found no official OpenAI, Google Search, Perplexity, Anthropic, or Doubao claim540that publishing it improves inclusion.541542`【Proposal】` Add it only when:543544- maintenance ownership exists;545- links are canonical, public, and useful;546- it does not expose private or unpublished material;547- it is measured as an experiment.548549Its absence is not an SEO/GEO defect. Never give it readiness points merely for550existing.551552### 8.11 GEO readiness rubric553554`【Proposal】` Score each dimension `0=blocked/absent`, `1=partial/unverified`,555`2=working with evidence`:5565571. retrieval eligibility;5582. index/render availability;5593. answer-unit clarity;5604. claim evidence and first-party value;5615. entity consistency;5626. query-to-asset coverage;5637. repeated platform measurement;5648. attribution/business-outcome linkage.565566Report the eight values separately. A total is an internal triage aid, not a567prediction of citation probability.568569## 9. Backlinks and external authority570571### Profile analysis572573- Use GSC links plus an authorized provider when available; state coverage limits.574- Review referring domains/pages, relevance, editorial context, destination,575 anchor distribution, acquisition trend, and lost links.576- Separate manipulative links from normal scraper/spam noise.577- Check unlinked brand mentions and incorrect citations that can be reclaimed.578579### Risk handling580581Do not call links “toxic” from a provider score alone. Remove or disavow only with582documented evidence of manipulative activity and material risk, such as a manual583action or known paid-link scheme. Preserve a decision log.584585### Earn links and citations586587- original datasets, tools, benchmarks, and transparent methods;588- expert contributions and primary-source commentary;589- genuinely useful comparison/reference pages;590- broken-link replacement where the asset is a true substitute;591- correction outreach for inaccurate facts or broken citations.592593No PBNs, paid-link laundering, mass guest-post spam, or automated outreach sludge.594595## 10. Analytics wiring (GSC + GA4)596597### GSC598599- Distinguish verification-token presence, current verified ownership, API/UI600 access, and data availability.601- A meta/DNS token is evidence of configuration, not proof of current access.602- Domain properties may cover subdomains; do not demand a subdomain meta token.603- Record property type, date range, search type, country/device filters, and604 anonymization/row-limit caveats.605606### GA4607608- Distinguish source-code presence, network request, DebugView/realtime event,609 stream configuration, and report/API access.610- On Liz house properties, preserve `G-TXVLTJJ878` exactly. Outside section 13,611 call it a measurement ID.612- Test consent state, duplicate tags, cross-domain rules, SPA navigation, and613 conversion/key-event semantics when in scope.614- Never infer collected users or conversions from a tag string alone.615616### Reporting shape617618Show clicks, impressions, CTR, average position, organic sessions/users, engaged619sessions, conversions/key events, landing page, query cluster, locale/device, and620comparison window only when the source is accessible. Otherwise mark `UNVERIFIED`.621622## 11. SERP and competitor analysis623624Capture exact query, locale, device, date/time, personalization state, and source.625626- Record organic results, local/video/image/shopping/news features, featured627 snippets, discussions, AI Overviews/AI Mode, and cited sources.628- Compare intent coverage, evidence, entity authority, information gain, format,629 freshness, UX, links, and conversion path—not word count alone.630- Separate true business competitors, organic competitors, and AI-cited sources.631- Identify answer/source gaps the site can satisfy honestly.632- Re-check volatile SERPs; one observation is not a stable market fact.633634### Comparison/alternative parity floor635636`【Proposal】` Before publishing a comparison or alternative page, open the637competitor's same-type URL in Ahrefs/Semrush (or authorized equivalent). Require638parity or better on: traffic estimate, keyword coverage, content depth, evidence639density. Else `BLOCK` publish. Qualitative SERP notes without this floor check640are incomplete.641642Output an opportunity table with query cluster, observed surface, winning source,643why it may satisfy the task, evidence gap, feasible asset, impact, effort, and644confidence.645646## 12. Rank, AI-visibility, and outcome tracking647648### Tracking setup649650- Freeze keyword/query corpus, landing-page mapping, locale, device, and platform.651- Track organic rank/landing URL beside generated mentions/citations.652- Annotate releases, migrations, campaigns, algorithm events, model/product653 changes, and known outages.654- Segment branded/unbranded, intent, market, template, and conversion value.655- Keep raw observations so metric definitions can be recomputed.656657### Page-level result ownership658659Assign one owner per priority URL. Success = that URL's GSC impressions + CTR660(and mapped conversions when available), not “we shipped SEO.” Unowned URLs stay661`UNVERIFIED` for outcome claims.662663### Decision rules664665`【Proposal】` Predeclare stop/continue rules per initiative. Example:666667- continue when leading evidence improves without technical or conversion harm;668- investigate when the canonical URL changes, variance spikes, or citations become669 inaccurate;670- stop when the site has no publish path, no audience fit, no measurable surface,671 or maintenance cost exceeds documented value.672673Do not call a SERP permanently occupied or an initiative causal from two points.674675## 13. House rules (Liz) — verbatim-accurate676677Follow exactly on house properties:6786791. **GA4 shared property ID `G-TXVLTJJ878` — never change/remove.**6802. **No brand/asset changes:** logos, fonts, palette, icons.6813. **Copy pain-first, EN default; no AI-slop filler.**6824. **Minimal change;** never rewrite whole files; **no new dependencies.**6835. **Never fabricate GSC verification tokens;** REPORT presence only.6846. **Don't touch** `projects/_templates/**`, `docs/AUDIT.md`, payment code.6857. **red-flowers** (reading site) is **private-by-design** (noindex + WAF CN) — **never SEO-refine it.**6868. **Verification:**687 - `git status -s` matches claimed files (leave pre-existing dirt unstaged)688 - `grep` GA4 ID count on touched HTML689 - sitemap `<loc>` count sane; robots head sane690 - **Stage selectively; never `git add -A`**691692### Extra house semantics (operational)693694- Apply these rules only after confirming a Liz house property.695- Preserve pre-existing changes and report them separately.696- Treat production bindings as strong host-intent evidence; investigate conflicts697 before migration.698- Do not stage or commit unless explicitly requested.699700## 14. Execution algorithm7017021. Scope gate: objective, target, access, authorization, exclusions, sample.7032. Snapshot: git/worktree state if local; date, host, redirects, robots, sitemaps.7043. Technical gate: reach/render/crawl/index/canonical before content polishing.7054. Analytics gate: identify observable and unverified layers.7065. Page/content/schema/CWV audit on representative templates.7076. Demand/SERP/competitor work when query evidence is in scope.7087. GEO audit:709 - platform and retrieval-path matrix;710 - crawler/WAF eligibility;711 - query-to-asset and answer-unit review;712 - fixed-corpus baseline;713 - citation, mention, accuracy, and source metrics.7148. Backlink/external-authority review when data exists.7159. Prioritize dependency first, then impact × confidence ÷ effort; keep the raw716 dimensions visible rather than pretending the quotient is precise.71710. Make only authorized minimal changes.71811. Verify changed files, behavioral checks, analytics preservation, and719 regressions.72012. Report method, evidence, unknowns, rollback, owners, and next measurement.721722### Severity rubric723724| Severity | Definition | Examples |725|---|---|---|726| P0 | production, legal, privacy, or discovery catastrophe requiring immediate action | public private data; sitewide accidental noindex; destructive redirect |727| P1 | material discovery, trust, or conversion blocker | broken canonical migration; key templates unavailable; false product facts |728| P2 | bounded opportunity or quality defect | weak answer units; missing contextual links; incomplete evidence |729| P3 | optional experiment or polish | `llms.txt` trial; low-value formatting refinement |730731Severity depends on affected scope and business impact. Missing optional schema or732an AI crawler blocked by policy is not automatically P0/P1.733734## 15. Commands and evidence cheatsheet735736Set variables; do not paste angle-bracket placeholders into a shell:737738```bash739host="example.com"740base="https://${host}"741742# Final GET status, effective URL, redirects, timing.743curl -sS -L --max-redirs 10 --connect-timeout 10 --max-time 30 \744 -o /dev/null -w 'status=%{http_code} url=%{url_effective} redirects=%{num_redirects} time=%{time_total}\n' \745 "${base}/"746747# Full response headers for one GET.748curl -sS --connect-timeout 10 --max-time 30 -D - -o /dev/null "${base}/"749750# Robots and sitemap content. Inspect complete files or save outside protected paths.751curl -sS --connect-timeout 10 --max-time 30 "${base}/robots.txt"752curl -sS --connect-timeout 10 --max-time 30 "${base}/sitemap.xml"753754# Compare responses to user-agent strings; this does not authenticate vendor bots.755curl -sS -A "OAI-SearchBot" -o /dev/null -w '%{http_code}\n' "${base}/"756curl -sS -A "PerplexityBot" -o /dev/null -w '%{http_code}\n' "${base}/"757758# House verification.759git status -s760grep -oF "G-TXVLTJJ878" "path/to/touched.html" | wc -l761```762763Parse XML/HTML with a real parser when correctness matters. Do not use line-count764grep as an XML element count, regex as a complete HTML parser, HEAD as proof of GET765behavior, or source presence as proof that analytics fires.766767Never claim GSC/GA4 UI state, index inclusion, bot identity, or generated citation768without direct evidence.769770## 16. Output template771772```markdown773# SEO/GEO report — [host] — [date]774775## Verdict776SHIP | FIX | BLOCK | KILL777778## Method and limits779- Scope/sample:780- Access:781- Locale/device/window:782- Unknowns:783784## P0 / P1 / P2785| Priority | Finding | Evidence | Scope | Confidence | Impact | Effort | Owner |786787## Technical / on-page / content / schema / CWV788- Observed:789- Expected:790- Evidence:791- Action:792793## GEO794- Target platforms and retrieval paths:795- Crawler/WAF eligibility:796- Query corpus and run protocol:797- Valid runs and run-level variance:798- Appearance rate — numerator/denominator:799- AI citation rate — numerator/denominator:800- Citation-to-mention rate — numerator/denominator:801- Share of cited voice — numerator/denominator:802- Source coverage — numerator/denominator:803- Citation accuracy — numerator/denominator:804- Cited URLs and competitors:805- Answer/evidence/entity gaps:806807## Analytics and outcomes808- GSC verification/access:809- GA4 G-TXVLTJJ878 on house properties: intact|missing|n/a810- Organic/referral/conversion evidence:811812## Changes and verification813- Files/actions:814- Tests:815- Rollback:816817## Next actions8181. [dependency-first action, owner, due/measurement]819```820821## 17. Failure patterns (do not repeat)822823| Pattern | Instead |824|---------|---------|825| Optimize content on noindex URLs | Fix indexability first |826| Sitemap of duplicate/wrong-host URLs | Fix canonical host, then sitemap |827| DNS-based canonical migration | Read bindings; then migrate |828| Fake GSC meta to clear a checklist | Report only |829| price:0 JSON-LD | Real price or omit Offer |830| AI-slop refresh | Pain-first specifics |831| `git add -A` | Explicit paths |832| SEO-refine red-flowers | Skip |833| Demand subdomain GSC tokens under domain property | Mark n/a |834| Ignore CF managed AI bot blocks | Flag to operator |835836V6 additions:837838| Pattern | Instead |839|---|---|840| Treat robots `Disallow` as `noindex` | Test crawling and indexing controls separately |841| Treat allowed AI bot as citation proof | Report eligibility only; run fixed-corpus tests |842| Treat bot user-agent as authenticated | Verify current official IP ranges and logs |843| Treat `GPTBot` or `ClaudeBot` as search crawlers | Separate training, search, and user-fetch agents |844| Treat `Google-Extended` as AI Overview control | Use Googlebot/Search preview controls per official docs |845| Require or score `llms.txt` as a standard | Optional measured experiment only |846| Publish generic FAQ/schema for “GEO” | Build useful answer units backed by evidence |847| Cherry-pick one favorable AI answer | Report all valid fixed-protocol runs |848| Merge platform rates without denominators | Report platform/surface metrics separately |849| Claim causality from before/after | Control protocol and state confounders/limits |850| Ship decision page as listicle | Pass format-vs-intent gate; rewrite as comparison/alternative |851| Publish below competitor same-type parity | Meet Ahrefs/Semrush parity floor first |852| Report “did SEO” without URL owner/metrics | Assign page owner; track GSC impressions + CTR per URL |853854## 18. Source synthesis and claim policy855856Use this precedence:8578581. current first-party platform documentation for product/crawler controls;8592. primary research for measured effects;8603. direct site, log, analytics, and fixed-protocol observations;8614. mounted practitioner tools as hypotheses and executable heuristics;8625. `【Proposal】` for unvalidated thresholds or internal decision rules.863864Mounted source synthesis:865866- GEO Optimizer: audit model, crawler/access diagnostics, Princeton GEO method867 summary, measurement/drift concepts.868- geo-skills: answer-block, citability, platform-test, crawler, and com869870…(truncated)