Purpose
Produce a Research Synthesis Brief -- a structured, evidence-graded synthesis of user research from multiple sources that separates validated findings from hypotheses, grades every claim by evidence quality, and connects insights to product decisions. The core question this skill answers: "What does the evidence actually say?" This is not a research plan or a methodology guide -- it is a reasoning engine that takes raw research inputs and produces findings a PM can act on, with explicit confidence levels and evidence gaps that drive the next research cycle.
When to Use / When NOT to Use
Use this skill when:
- Synthesizing findings from multiple user interviews into a coherent picture
- Evaluating whether research evidence is strong enough to support a product decision
- Running a discovery sprint and need to structure findings as they accumulate
- Analyzing qualitative data (interview transcripts, support tickets, forum posts) for patterns
- Building research-backed feature hypotheses before committing engineering resources
- Resolving conflicting signals across different research sources (surveys say X, interviews say Y, behavioral data says Z)
- Assessing what you still do NOT know and where evidence is dangerously thin
Do NOT use this skill when:
- You need competitive market analysis (-> Competitive & Market Analysis skill -- that is market-side structural analysis, this is demand-side primary research)
- You need to design metrics or experiments (-> Metric Design & Experimentation skill -- that is measurement, this is evidence gathering and synthesis)
- You need to write a product specification (-> Spec Writing skill -- use this skill's output as INPUT to spec writing)
- You need a research plan template without existing data to synthesize (this skill processes research, it does not design research methodology from scratch)
- You need statistical analysis of quantitative experiment results (-> Metric Design & Experimentation skill)
Anti-inputs (what this skill does NOT handle):
- Research methodology design (sampling strategy, interview guide creation -> research ops)
- Statistical hypothesis testing (-> Metric Design & Experimentation skill)
- Competitive intelligence gathering (-> Competitive & Market Analysis skill)
- Product specification decisions (-> Spec Writing skill)
- Dashboard or visualization design (-> BI tooling)
Format Rules (Read First)
These rules govern every output produced by this codex. They are not style preferences -- they are quality enforcement mechanisms derived from the systematic failure modes of research synthesis.
Take positions. Never hedge with weasel words. "Likely," "may," "could," and "seems" are banned. Flag uncertainty with explicit confidence levels: H (>70% confident), M (40-70%), L (<40%). Example: "Users abandon the onboarding flow at the permissions step because the value proposition is unclear before data access is granted (H)" not "Users may be confused by the onboarding flow."
Per-cell evidence tier annotation is mandatory in all comparison and finding matrices. Every cell in a finding summary, source comparison, or insight classification table must carry an inline evidence tier tag:
(T1),(T2),(T3)etc. Where evidence is absent:(T6: inferred). A matrix with untagged cells is incomplete.The O->I->R->C->W cascade applies to ALL research-driven recommendations, not just final sections. Format: Observation [evidence tier] -> Implication [mechanism] -> Response [specific action] -> Confidence [H/M/L + assumption] -> Watch Indicator [observable signal].
Begin with Framework Selection (Step 0b) before applying any framework. Identify the research question type; select the 3-5 load-bearing frameworks; note which to skip. Applying all 8 frameworks to an interview analysis question inflates length and dilutes signal.
Contradictions between sources are signal, not noise. When two research sources reach different conclusions (e.g., surveys show high satisfaction but interviews reveal deep frustration), surface the contradiction explicitly and state which to weight more heavily and why, citing the evidence hierarchy.
Flag time-sensitive claims. Any claim based on research older than 6 months must carry
[POTENTIALLY STALE -- verify before presenting]. Evidence tier and recency are independent -- a T2 finding from 18 months ago may be less useful than a T4 report from last month if the market has shifted.Flag thin-evidence conclusions. If a key finding rests only on Tier 4-6 evidence, prepend it with
[EVIDENCE-LIMITED: validate with Tier 1-2 before acting].Every framework reference gets a one-line contextual explanation the first time it appears. Not "Evidence Triangulation" but "Triangulation -- converging multiple independent sources to validate a finding. When three weak signals from different methods agree, confidence increases multiplicatively, not additively." The reader needs to know why this lens matters for their decision.
The document must be navigable by someone who didn't create it. Include a reading guide (by time and by role), a notation key, and layered depth. A VP should be able to read only the Executive Summary and act. A PM should be able to read through Key Findings and skip the Deep Analysis. The full document is for the researcher.
Output Template (Mandatory Document Skeleton)
Every Research Synthesis Brief MUST follow this exact structure. Copy this skeleton and fill it in. Do not reorder sections, skip sections, or invent new top-level sections. If a framework was skipped in Step 0b, note "Skipped -- not load-bearing for this research question" in that section.
# Research Synthesis Brief: [Subject -- e.g., "Enterprise User Onboarding Pain Points"]
> **Date:** [YYYY-MM-DD] | **Confidence band:** [Overall H/M/L] | **Staleness window:** [Date after which key findings need revalidation] | **Sources synthesized:** [count and types]
---
## Executive Summary
[5 sentences max. A VP reads only this and makes a decision. No framework names, no jargon, no evidence tier tags. Plain language a non-PM exec can act on. Final sentence = the recommended action in bold.]
---
## How to Read This Document
**What this is:** A research evidence synthesis -- not a research report. It grades evidence quality, separates validated findings from hypotheses, and connects insights to product decisions.
**Reading by time available:**
| Time | Read | You'll get |
|---|---|---|
| **5 min** | Executive Summary only | The key finding + recommended action |
| **15 min** | Executive Summary + Key Findings (sections 1-3) | Core evidence-graded conclusions with source quality |
| **30 min** | Full document through Recommendations | Complete synthesis with framework analysis and evidence gaps |
| **Deep dive** | Everything including Appendix | Full source analysis, methodology notes, adversarial critique |
**Reading by role:**
| Role | Start with | Then read | Skip unless curious |
|---|---|---|---|
| VP / Exec | Executive Summary | Recommendations, Assumption Registry | Everything in between |
| PM Lead | Executive Summary | Key Findings (sections 2-5), Research Gap Map, Recommendations | Deep framework sections |
| Researcher / Analyst | Full document in order | Source Quality Assessment, Cross-Source Contradictions, Adversarial Self-Critique | Nothing -- this is your primary artifact |
| Designer | Executive Summary | User Need Patterns (section 3), Behavioral vs. Stated Findings | Evidence methodology sections |
| Engineer | Executive Summary | Validated Requirements (section 6), Signal vs. Noise summary | Research methodology |
---
## Notation Key
**Confidence levels** -- applied to every finding:
- **H (>70% confident)** -- Multiple independent sources converge. Act on it.
- **M (40-70%)** -- Direction is probable but evidence is mixed or thin. Validate before committing resources.
- **L (<40%)** -- This is a hypothesis, not a finding. Do not act without further evidence.
**Evidence tiers** -- how we know what we claim to know (tagged inline as T1-T6):
- **T1** -- Direct behavioral data: usage analytics, A/B results, usage logs (strongest -- what users DO)
- **T2** -- Primary qualitative research: your own interviews (n>=5 with consistent patterns), usability tests
- **T3** -- Secondary research with disclosed methodology: academic studies, well-sampled surveys
- **T4** -- Industry reports and analyst commentary: Gartner, Forrester, market surveys
- **T5** -- Anecdotal evidence: individual interviews (<5), forum posts, support tickets, app reviews
- **T6** -- Inference / expert opinion without data (weakest -- treat as starting hypothesis)
**Insight classification:**
- **VALIDATED** -- Multiple independent sources converge (>=2 tiers, >=3 sources)
- **EMERGING** -- Single strong signal or consistent pattern from one source type
- **HYPOTHESIS** -- Inference from patterns; plausible but not directly observed
- **CONTRADICTED** -- Conflicting evidence across sources; resolution required
**Recommendation format** (O->I->R->C->W):
- **O**bservation -- What the evidence shows (with evidence tier)
- **I**mplication -- Why it matters (the mechanism)
- **R**esponse -- What to do (specific action + owner + timeline)
- **C**onfidence -- How sure we are (H/M/L + key assumption)
- **W**atch -- How to know if we're wrong (observable signal)
**Flags:**
- `[POTENTIALLY STALE]` -- Source data is >6 months old; verify before presenting
- `[EVIDENCE-LIMITED]` -- Conclusion rests on T4-T6 evidence only; validate with stronger data before acting
---
## Step 0: Context Fitness Check
Before selecting frameworks, verify that a Research Synthesis Brief is the right artifact for this question.
| Question | If Yes | If No |
|---|---|---|
| **Is the core question about what users need, want, or do?** | Proceed to Framework Selection | A Research Synthesis Brief is the wrong artifact. Consider: Competitive War Map (market structure question), Measurement Framework (instrumentation question), or Problem Framing (problem definition question). |
| **Do you have multiple research sources to synthesize?** | Full synthesis -- apply Evidence Triangulation and Source Quality Assessment | Single-source analysis. Flag prominently: "This brief synthesizes a single source type. All findings are tentative until triangulated with independent sources." Cap findings at M confidence unless behavioral data (T1). |
| **Is the research about current or prospective users?** | Standard evidence hierarchy applies | Flag: "Research subjects are not current users. Stated preferences carry even less weight than usual. Behavioral proxies (competitor usage, adjacent product data) should be weighted above interview responses." |
| **Has the product shipped, or is this pre-product research?** | T1 behavioral data may be available -- prioritize it | No T1 data exists. Highest available tier is T2 (interviews). Flag: "Pre-product research -- no behavioral validation available. All findings are directional until validated post-launch." |
**If any answer triggers a "No" path:** State this prominently at the top of the Executive Summary.
---
## Step 0b: Framework Selection
| Research question type | Primary frameworks (apply in full) | Supporting frameworks (scan only) | Skipped (why) |
|---|---|---|---|
| [e.g., "Interview synthesis"] | [e.g., Interview Analysis, Evidence Triangulation, Insight Classification] | [e.g., Signal vs. Noise, Research Gap Mapping] | [e.g., "Competitive Discovery -- no competitor data available"] |
---
## 1. Source Inventory & Quality Assessment
| Source | Type | Sample | Methodology | Tier | Recency | Key Bias Risk |
|---|---|---|---|---|---|---|
| [e.g., "User interviews (n=12)"] | Qualitative | n=12, enterprise segment | Semi-structured, 45 min | T2 | [date] | Self-selection, social desirability |
| [e.g., "Product analytics"] | Behavioral | N=22K DAU | Full population | T1 | [date] | Survivorship (churned users absent) |
| [e.g., "Gartner report"] | Industry | Unknown sample | Undisclosed | T4 | [date] | Vendor influence, category bias |
**Source coverage assessment:** [Which user segments, use cases, or questions have NO evidence? Flag these as research gaps.]
---
## 2. Key Findings (Evidence-Graded)
**Finding 1: [Insight Header -- conclusion, not topic]**
- **Classification:** VALIDATED / EMERGING / HYPOTHESIS / CONTRADICTED
- **Confidence:** H / M / L
- **Evidence:** [Source A (TX), Source B (TX), Source C (TX)]
- **Detail:** [2-3 sentences explaining the finding with specifics]
- **Implication:** [What this means for product decisions]
**Finding 2: [Insight Header]**
[Same structure]
**Finding 3: [Insight Header]**
[Same structure]
[Continue for all key findings. Minimum 5, maximum 12.]
---
## 3. User Need Patterns
| Need | Frequency (how many sources) | Intensity (how strongly expressed) | Behavioral Evidence? | Classification | Confidence |
|---|---|---|---|---|---|
| [need] | X/Y sources | High/Med/Low | Yes (T1) / No | VALIDATED/EMERGING/HYPOTHESIS | H/M/L |
| [need] | | | | | |
**Frequency vs. Intensity Matrix:**
| | Low Frequency | High Frequency |
|---|---|---|
| **High Intensity** | Niche but critical -- potential power-user feature or underserved segment | Core unmet need -- high-priority opportunity |
| **Low Intensity** | Noise -- deprioritize | Table stakes -- expected but not differentiating |
---
## 4. Stated vs. Revealed Preferences
| What Users SAY | What Users DO | Gap | Implication |
|---|---|---|---|
| [stated preference (TX)] | [observed behavior (TX)] | [description of gap] | [what the gap means for product decisions] |
**Evidence hierarchy reminder:** Behavioral data (T1) > Usability tests (T2) > Surveys (T3) > Interviews (T2-T5 depending on n) > Focus groups (T5). When stated and revealed preferences conflict, behavioral data wins unless you can identify a specific reason the behavioral data is misleading (e.g., the current product forces a behavior that doesn't reflect preference).
---
## 5. Signal vs. Noise Assessment
| Signal Candidate | Source | Frequency | Convergence | Structural? | Verdict |
|---|---|---|---|---|---|
| [signal] | [source (TX)] | Recurring/One-off | Converges with X other signals / Isolated | Structure change / Weather | SIGNAL / NOISE / INVESTIGATE |
---
## 6. Evidence Triangulation Summary
| Finding | Source 1 | Source 2 | Source 3 | Convergence? | Confidence After Triangulation |
|---|---|---|---|---|---|
| [finding] | [source (TX): supports/contradicts/silent] | [source (TX): ...] | [source (TX): ...] | Strong/Partial/Weak/Contradicted | H/M/L |
---
## 7. Research Gap Map
| Question We Cannot Answer | Why It Matters | What Evidence Would Answer It | Recommended Method | Priority |
|---|---|---|---|---|
| [unanswered question] | [product decision it blocks] | [what T1/T2 evidence would look like] | [interview/survey/analytics/prototype test] | Critical/High/Medium |
---
## 8. Cross-Source Contradictions
| Contradiction | Source A says | Source B says | Resolution / Which to weight |
|---|---|---|---|
| [e.g., "Satisfaction vs. churn"] | [Surveys: 85% satisfied (T3)] | [Analytics: 40% churn in 90 days (T1)] | [Behavioral data (T1) wins. Survey satisfaction likely reflects switching cost tolerance, not genuine satisfaction. Investigate further.] |
---
## 9. Competitive Discovery Findings (if applicable)
| Competitor | Source | Unmet Need Identified | Evidence Quality | Implication for Our Product |
|---|---|---|---|---|
| [competitor] | [reviews/forums/job postings (TX)] | [what their users want but don't get] | T4/T5 | [how we could address this] |
---
## 10. Strategic Recommendations (O->I->R->C->W Cascade)
**Recommendation 1: [Title]**
- **Observation** [TX]: [What the evidence shows]
- **Implication**: [Why it matters -- the mechanism]
- **Response**: [Specific action + owner + timeline]
- **Confidence**: [H/M/L] -- assumes [key assumption]
- **Watch**: [Observable signal]; if [threshold], re-assess
**Recommendation 2: [Title]**
- **Observation** [TX]: ...
- **Implication**: ...
- **Response**: ...
- **Confidence**: ...
- **Watch**: ...
**Recommendation 3: [Title]**
[Same structure]
---
## Assumption Registry
| # | Assumption | Finding it underpins | Confidence | Evidence | What would invalidate this |
|---|---|---|---|---|---|
| 1 | | | H/M/L | (TX) | |
| 2 | | | H/M/L | (TX) | |
| 3 | | | H/M/L | (TX) | |
---
## Adversarial Self-Critique
**Weakness 1: [Title]**
[Steelmanned argument against the synthesis. What assumption is being made? What evidence would disprove it? Scenario where this recommendation is catastrophically wrong.]
**Weakness 2: [Title]**
[Same depth]
**Weakness 3: [Title]**
[Same depth]
---
## Revision Triggers
| Trigger | What to re-assess | Timeline |
|---|---|---|
| [Observable event] | [Which findings break] | [When to check] |
---
## Sources
[All sources cited in the synthesis, with evidence tier and date.]
Rules for using this template:
- Do not skip sections. If a section isn't applicable, write "Skipped -- [reason]" and move on.
- Every table cell with a finding or claim must have an evidence tier tag --
(T1)through(T6). - Section headers are conclusions, not labels. Replace generic headers (e.g., "User Need Analysis") with insight headers (e.g., "Users Need Real-Time Collaboration, Not Better Dashboards") after completing the section.
- The Executive Summary is written last but appears first. Do not write it until all sections are complete.
- Findings are ordered by confidence, not by source. Lead with VALIDATED findings, then EMERGING, then HYPOTHESIS. CONTRADICTED findings get their own section.
Domain Frameworks
This section IS the knowledge weapon. Each framework is encoded with its scoring rubrics, decision tables, and application methodology -- not merely referenced. A PM using this skill produces research synthesis that requires these frameworks; without them, the output degrades to a summary of interview notes.
Framework 1: Evidence Triangulation
The core analytical engine for validating findings across multiple sources. Triangulation is not "did three people say it?" -- it is "do three independent methods converge on the same conclusion?"
The Triangulation Principle: A finding supported by a single source type is an observation. A finding supported by two independent source types is probable. A finding supported by three independent source types with different bias profiles is as close to validated as qualitative research gets.
Why independence matters: Five interviews from the same customer segment, conducted by the same researcher, with the same interview guide, are one source -- not five. Independence requires: different methods (interview vs. analytics vs. survey), different populations, or different time periods.
Triangulation Types:
| Type | Method | When to Use | Strength |
|---|---|---|---|
| Method triangulation | Same question, different research methods (interview + survey + analytics) | Default -- always try this first | Different methods have different biases; convergence across methods is the strongest validation |
| Source triangulation | Same method, different populations (enterprise users + SMB users + churned users) | When you need to test whether a finding generalizes | Reveals segment-specific vs. universal patterns |
| Temporal triangulation | Same question, different time periods (Q1 interviews vs. Q3 interviews) | When recency or trend direction matters | Distinguishes stable needs from transient reactions |
| Investigator triangulation | Same data, different analysts | When the finding is high-stakes and interpretation matters | Reduces single-analyst bias in qualitative coding |
Convergence Scoring:
| Convergence Level | Definition | Confidence Impact | Action |
|---|---|---|---|
| Strong | 3+ independent source types agree, no contradictions | Upgrade finding to H confidence | Act on this finding |
| Partial | 2 source types agree, 1 silent or ambiguous | M confidence | Usable for planning; validate before major commitment |
| Weak | 1 source type only, or 2 sources but same method | L confidence | Hypothesis only; design research to validate |
| Contradicted | Sources actively disagree | Flag explicitly | Do NOT average. Investigate why sources disagree -- the disagreement is often the most valuable insight |
The "3 Weak Signals" Question: When do 3 weak signals (T5) equal 1 strong signal? Only when they are independent. Three forum posts from the same community are one signal repeated three times. Three signals from different channels (forum + support ticket + sales call note), about the same problem, from users in different segments -- that is genuine convergence.
When do 3 weak signals NOT equal 1 strong signal? When they share a common cause. If a tech blogger wrote about a problem, and then three users cited the same blog post in support tickets, the "three signals" are actually one signal amplified. Always trace weak signals to their origin to check for shared roots.
Output format -- Triangulation Matrix:
| Finding | Interviews (T2) | Analytics (T1) | Survey (T3) | Support (T5) | Convergence | Confidence |
|---------|:---:|:---:|:---:|:---:|---|---|
| [Finding 1] | Supports (n=8/12) | Supports (funnel data) | Silent | Supports (47 tickets) | Strong (3 types) | H |
| [Finding 2] | Supports (n=4/12) | Contradicts | Silent | Silent | Contradicted | Investigate |
Framework 2: Interview Analysis Protocol
The systematic methodology for extracting reliable insights from qualitative interviews. The challenge: interviews are rich in detail and poor in reliability. This protocol converts unstructured qualitative data into structured, gradable evidence.
The Core Problem with Interviews: Interviews are the most commonly used and most commonly misinterpreted research method. Users are unreliable narrators of their own behavior, preferences, and needs. They will tell you what they think you want to hear (social desirability bias), what sounds rational rather than what they actually did (post-hoc rationalization), and what happened most recently rather than what matters most (recency bias). Despite this, interviews are irreplaceable for discovering why behind behavioral patterns.
Coding Protocol (Pattern Extraction from Transcripts):
| Step | Action | Purpose |
|---|---|---|
| 1. Open coding | Read transcripts without a framework. Mark every statement that could be an insight, a need, a frustration, or a behavior. | Avoid confirmation bias -- let the data speak before imposing categories |
| 2. Axial coding | Group open codes into categories. Merge similar codes. Name the categories by the underlying need, not the surface statement. | "I want a dark mode" and "the screen hurts my eyes" are the same code: visual comfort need |
| 3. Frequency count | Count how many independent interviewees raised each coded category. | Frequency = breadth of need. Minimum 3/n to flag as a pattern (not a one-off) |
| 4. Intensity scoring | Rate each mention by emotional intensity: casual mention vs. unprompted frustration vs. "I would switch products for this." | Intensity = depth of need. A need mentioned by 2 users with extreme intensity may matter more than a need mentioned by 8 users in passing |
| 5. Verbatim extraction | Pull 2-3 direct quotes per coded category. Preserve exact language, not your paraphrase. | Verbatims are T2 evidence. Your interpretation is T6. Keep both, label both. |
| 6. Interpretation layer | Now overlay your interpretation. For each category: what is the underlying job-to-be-done? What is the user actually trying to accomplish? | Separate the data layer (verbatims, frequencies) from the interpretation layer (your inference about what it means). Readers should be able to disagree with your interpretation while accepting your data. |
Frequency vs. Intensity Decision Table:
| Low Intensity (casual mention) | High Intensity (unprompted, emotional, "would switch for this") | |
|---|---|---|
| High Frequency (>=50% of n) | Table stakes need -- users expect it but won't pay extra for it. Build it, but it won't differentiate. | Core unmet need -- highest priority. This is where product value lives. |
| Low Frequency (<30% of n) | Background noise -- deprioritize unless segment-specific analysis reveals a cluster. | Power-user or niche signal -- investigate whether this represents an underserved segment. Small n + high intensity = potential beachhead market. |
Verbatim vs. Interpreted Claims:
| Type | Example | Evidence Tier | How to Use |
|---|---|---|---|
| Verbatim | "I spend 45 minutes every Monday manually copying data from Salesforce into our dashboard." | T2 (if n>=5 with consistent pattern) or T5 (if isolated) | Present as primary evidence. Let the reader draw their own conclusion. |
| Interpreted | "Users experience significant friction in the data integration workflow." | T6 (your inference) | Present as your interpretation, explicitly labeled. Tie back to the verbatims that support it. |
| Hybrid (best practice) | "Users report spending 30-60 minutes per week on manual data integration (n=7/12 interviewees, unprompted). This suggests the integration workflow is a primary source of friction, not a minor inconvenience." | T2 + T6 (separated) | Data sentence is T2. Interpretation sentence is T6. Both are useful; conflating them is not. |
Sample Size Guidance for Interview Research:
| n | What You Can Claim | What You Cannot Claim |
|---|---|---|
| 1-4 | Individual perspectives; hypothesis generation | Patterns, prevalence, segment-level conclusions |
| 5-8 | Emerging patterns (if 60%+ consistency); directional themes | Definitive prevalence; confidence above M |
| 9-15 | Validated qualitative patterns (if 60%+ consistency); segment-level themes with careful scoping | Statistical prevalence; quantitative generalization |
| 16-30 | Robust qualitative saturation (when new interviews stop producing new codes); strong segment comparisons | Population-level statistics (that still requires quantitative methods) |
| 30+ | Thematic saturation is almost certain; quasi-quantitative prevalence claims become defensible | Nothing changes about the method's core limitation: you still cannot claim "X% of users feel Y" from interviews alone |
Critical rule: Never cite "12 out of 15 interviewees said X" as if it means "80% of users feel X." Interview samples are never representative in the statistical sense. The correct framing: "A clear pattern emerged across 12 of 15 interviewees, suggesting this is a widespread need in the [segment] population. Quantitative validation recommended before sizing the opportunity."
Framework 3: Research Quality Assessment
The systematic methodology for evaluating whether a research source is trustworthy enough to inform product decisions. Not all evidence is created equal, and mixing T1 behavioral data with T5 anecdotes without distinguishing them is the most common research synthesis failure.
Source Evaluation Rubric:
| Dimension | Questions to Ask | Scoring |
|---|---|---|
| Sample | How many subjects? How selected? Representative of target population? | n>=30 quantitative = adequate; n>=8 qualitative with saturation = adequate; convenience sample = flag bias risk |
| Methodology | Disclosed? Replicable? Standard method or novel? Controls for known biases? | Disclosed + standard = trustworthy. Undisclosed = T4 at best. Novel without validation = treat as exploratory. |
| Bias detection | Who funded it? Who conducted it? What incentives exist to reach a specific conclusion? | Vendor-funded research about their own product = T5 regardless of methodology. Academic with disclosed funding = T3. |
| Recency | When was the data collected? Has the market changed since? | <6 months = current. 6-12 months = usable with [POTENTIALLY STALE]. >12 months = historical context only unless nothing has changed. |
| Relevance | Does this research study the same user population, use case, and context as your product? | Same population + same context = directly applicable. Adjacent population or different context = inference required (downgrade one tier). |
The Relevance Trap: A beautifully designed, large-sample study about enterprise SaaS onboarding is T3 evidence for your enterprise SaaS onboarding question. The same study is T4-T5 evidence for your consumer mobile onboarding question, because the populations and contexts differ enough that findings may not transfer. Always evaluate relevance to YOUR specific context, not just research quality in the abstract.
Source Quality Decision Table:
| Quality Dimension | Strong | Adequate | Weak | Disqualifying |
|---|---|---|---|---|
| Sample size (quantitative) | n>=1000 | n>=100 | n=30-99 | n<30 |
| Sample size (qualitative) | n>=15 with saturation | n=8-14 | n=5-7 | n<5 (anecdotal) |
| Methodology disclosure | Full methodology + data available | Methodology described | "Survey" or "interviews" with no detail | No methodology mentioned |
| Bias risk | Independent researcher, no conflicts | Funded research with disclosed conflict | Vendor-funded | Vendor self-report as evidence |
| Recency | <6 months | 6-12 months | 12-24 months | >24 months without [STALE] flag |
| Relevance | Same population, same context | Same population, different context | Different population, similar context | Different population AND context |
Tier Assignment Protocol:
- Start with the source type's default tier (T1-T6 from Evidence Standards below)
- Apply quality assessment: strong on all dimensions = keep tier. Weak on any critical dimension (sample, methodology, relevance) = downgrade one tier.
- Apply recency: >12 months = downgrade one tier or add
[POTENTIALLY STALE]. - Apply relevance: different context = downgrade one tier.
- Floor: No source drops below T6 regardless of downgrades. But T6 evidence should never be the sole basis for a finding.
Framework 4: Signal vs. Noise Filter
The systematic methodology for distinguishing genuine user needs from noise: stated preferences that don't predict behavior, outlier opinions amplified by recency, feature requests that mask deeper problems, and the ever-present recency bias that makes the latest interview dominate all previous data.
The Five-Filter Test:
Apply these five filters to every candidate insight before elevating it to a finding:
| Filter | Question | If Yes -> Signal | If No -> Noise (or Investigate) |
|---|---|---|---|
| 1. Frequency | Has this appeared in >=3 independent sources? | Pattern, not one-off | One-off unless high-intensity (check Filter 2) |
| 2. Intensity | Did users express this unprompted, with emotion, or as a dealbreaker? | Even at low frequency, investigate as niche signal | Low-intensity + low-frequency = deprioritize |
| 3. Behavioral corroboration | Is there behavioral evidence (T1) that supports the stated need? | Validated signal | Stated-only. May be aspiration, not need. Flag for behavioral validation. |
| 4. Structural significance | Does this represent a change in market structure (durable) or market weather (temporary)? | Structural = high strategic weight | Weather = monitor but don't over-invest |
| 5. Independence | Are the sources genuinely independent, or could they share a common origin? | Independent convergence | Amplified signal -- trace to origin before counting |
Common Noise Patterns:
| Noise Type | What It Looks Like | Why It Fools You | How to Detect |
|---|---|---|---|
| Feature request as proxy | "I need a Gantt chart view" | The user is expressing a need for project visibility, not literally requesting Gantt charts. The request masks the job-to-be-done. | Ask "What would you do with that?" until you reach the underlying need. The fifth "why" reveals the real job. |
| Recency bias | The last 3 interviews all mentioned X; it must be critical | The last interviews are freshest in memory. The 20 previous interviews that didn't mention X are forgotten. | Always count frequency across the FULL dataset, not just recent additions. Weight by temporal triangulation, not by memory. |
| Social desirability | "I would definitely pay more for that feature" | In interviews, users overstate willingness to pay, understate price sensitivity, and agree with the interviewer's implied preferences. | Stated willingness to pay is T5 evidence. Only revealed behavior (T1) validates pricing. Never set price based on interview responses alone. |
| Expert user bias | "The API should support GraphQL" | Power users who agree to interviews are not representative of the full user base. Their needs are real but may be niche. | Segment findings by user sophistication. If a need only appears in expert-user interviews, it's a segment signal, not a market signal. |
| Survivorship bias | "Our users love the product!" | You're only interviewing users who stayed. The 40% who churned are absent from your research. | Include churned users and non-adopters in the research plan. If your sample is only active users, flag: "Survivorship bias: findings reflect retained users only." |
The Noise Escalation Ladder: Not all noise should be discarded. Some noise deserves monitoring:
| Action | When |
|---|---|
| Discard | Failed all 5 filters. One person, one time, no corroboration, no structural significance. |
| Monitor | Failed 3-4 filters but passed one with high intensity or structural significance. Add to watch list for the next research cycle. |
| Investigate | Passed 2-3 filters. Not enough evidence to act, but too much to ignore. Design a targeted research probe. |
| Elevate to Finding | Passed 4-5 filters. This is signal. Classify and grade per the Insight Classification framework. |
Framework 5: Insight Classification
The taxonomy for categorizing research findings by their evidence strength and actionability. Not all insights are created equal, and treating a hypothesis with the same weight as a validated finding is the core failure mode of research synthesis.
Classification Taxonomy:
| Classification | Definition | Evidence Requirement | Action Permission | Visual Marker |
|---|---|---|---|---|
| VALIDATED | Multiple independent sources converge; finding has been triangulated across methods | >=2 evidence tiers, >=3 independent sources, no unresolved contradictions | Act on this finding. Use for resource allocation, roadmap decisions, spec requirements. | Solid -- build on it |
| EMERGING | Single strong signal or consistent pattern from one source type; promising but not yet triangulated | 1 evidence tier with strong signal (e.g., 8/12 interviewees, or clear behavioral data) | Plan for this finding. Include in roadmap discussions but validate before committing major resources. | Directional -- plan with it |
| HYPOTHESIS | Inference from patterns; plausible interpretation of indirect evidence but not directly observed | Logical inference from validated or emerging findings; no direct evidence | Test this finding. Design a specific research probe or experiment. Do not build features based on hypotheses alone. | Unvalidated -- test it |
| CONTRADICTED | Conflicting evidence across sources; the finding status is genuinely uncertain | Two or more sources actively disagree; resolution requires additional research | Investigate, do not act. The contradiction itself is the most important finding -- it reveals something unexpected about user behavior or market structure. | Conflicted -- investigate it |
Classification Decision Tree:
Is there direct behavioral evidence (T1)?
YES -> Is it corroborated by another source type?
YES -> VALIDATED (H confidence)
NO -> EMERGING (M confidence) -- strong but single-method
NO -> Is there qualitative evidence from n>=5?
YES -> Are there contradicting sources?
YES -> CONTRADICTED -- investigate the gap
NO -> Is there a second independent source type that agrees?
YES -> VALIDATED (M confidence, upgrade to H when behavioral data confirms)
NO -> EMERGING (M confidence)
NO -> Is the inference logically derived from validated/emerging findings?
YES -> HYPOTHESIS (L confidence)
NO -> Not a finding. Discard or file as speculation.
Classification Upgrade/Downgrade Triggers:
| Trigger | Classification Change |
|---|---|
| New behavioral data (T1) confirms an EMERGING finding | EMERGING -> VALIDATED |
| A contradicting source appears for a VALIDATED finding | VALIDATED -> CONTRADICTED (investigate immediately) |
| Time passes without revalidation (>6 months) | Any classification gets [POTENTIALLY STALE] |
| The user population changes (new segment, new market) | All classifications from old population get relevance downgrade |
| An EMERGING finding fails a targeted validation test | EMERGING -> discard or HYPOTHESIS |
Framework 6: Demand-Side Analysis
The framework for understanding what users actually do versus what they say -- the hierarchy of evidence for user needs, and how to resolve the inevitable conflicts between different evidence types.
The Evidence Hierarchy for User Needs:
| Rank | Evidence Type | What It Tells You | Reliability | Common Failure |
|---|---|---|---|---|
| 1 | Behavioral data (analytics) | What users actually do at scale | Highest -- behavior doesn't lie (but can be misinterpreted) | Observing behavior without understanding context. Users click X -- but why? |
| 2 | Usability tests | What users do in controlled settings with a specific task | High -- observed behavior, but in artificial conditions | Artificial task framing changes behavior. "Find the settings page" != natural navigation. |
| 3 | Surveys (well-sampled) | What users say they do and prefer, at scale | Moderate -- scale offsets self-report bias somewhat | Leading questions, response bias, survivorship in sample |
| 4 | Interviews (n>=5) | Why users do what they do; underlying motivations and frustrations | Moderate -- depth offsets small sample, but self-report bias remains | Social desirability, post-hoc rationalization, interviewer influence |
| 5 | Focus groups | How users talk about needs in social settings | Low-moderate -- group dynamics distort individual preferences | Dominant personalities drive consensus; social desirability amplified |
| 6 | Expert opinion | What experienced practitioners believe users need | Low -- expert prediction of user behavior is notoriously unreliable | Curse of knowledge: experts cannot accurately simulate non-expert experience |
The Stated vs. Revealed Preference Gap: The most important and most underappreciated dynamic in user research. Users will tell you one thing and do another -- not because they're lying, but because:
- They don't have accurate introspective access to their own decision-making
- They describe their idealized self, not their actual self
- They anchor on recent experiences rather than stable preferences
- They want to be helpful and will agree with your implied hypothesis
Resolution Protocol When Stated != Revealed:
| Situation | Resolution | Rationale |
|---|---|---|
| Users SAY they want feature X; behavioral data shows they never use similar features | Weight behavioral data. Investigate why the gap exists (awareness? discoverability? mismatch between stated and actual job?) but do NOT build X based on stated preference alone. | Wh |
…(truncated)