1---2name: ux-research-and-usability-testing3description: Use when planning moderated usability sessions, participant tasks, observation evidence, interviews, surveys, card/tree tests, diary studies, or findings synthesis. Do not use for expert-only heuristic inspection or unsupported personas; route those to heuristic evaluation or journey mapping.4---56# UX Research and Usability Testing78<!-- dual-compat-start -->9## Use When10- You need to choose a research method and cannot tell which one answers the question (generative vs. evaluative, attitudinal vs. behavioural, qual vs. quant). Use `references/research-method-selector.md`.11- You are running user interviews, contextual inquiry, surveys, card sorts, tree tests, or diary studies and need a defensible plan, screener, and guide.12- You are running a usability test — moderated or unmoderated — and need a real protocol: tasks, success criteria, think-aloud, severity rating. Use `references/usability-test-protocol.md`.13- You have raw research notes and must synthesize them into findings and then into **decisions** (not a report that gets shelved).14- A stakeholder is asserting a user "fact" with no evidence, or a design debate has stalled on opinion and needs data to break the tie.1516## Do Not Use When17- The work is applying cognitive/behavioural theory or heuristics to a design (Nielsen, Norman, Gestalt, biases, cognitive load) → use the sibling `ux-psychology`.18- The work is the end-to-end product/discovery process, RACI, and rituals → use the sibling `enterprise-ux-process`.19- The work is purely writing interface copy, microcopy, or error messages → use group 10 `ux-writing-and-microcopy` / `error-empty-and-system-messaging`.20- The work is building the prototype that the test will run on → use group 05 `wireframing-and-prototyping`.2122## Required Inputs2324| Input | Source | Evidence |25|---|---|---|26| Research question and decision | Sponsor/research lead | Named decision owner and action the study can change |27| Target population and recruitment limits | Product/research operations | Inclusion, exclusion, sample, and consent requirements |28| Stimulus, tasks, data, and risk controls | Product, legal, privacy | Stable artefact, protocol, retention, and sensitive-data rules |29- **The decision the research must inform.** Every study starts from a decision that is currently being made on a guess. No decision → no study.30- The research question(s), stated so they can be answered with evidence, and what you currently believe (the assumption being tested).31- The target user / segment, and access to them (recruiting source, screener constraints, incentive budget).32- The artifact under test where the method is evaluative: live product, prototype, wireframe, IA, or competitor.33- Timeline and constraints (how many participants, moderated vs. unmoderated, remote vs. in-person, regulatory/consent constraints).3435## Workflow361. **Anchor to a decision.** Write the decision, the question, and the current assumption in one line each. If you cannot name the decision the result will change, stop — you are doing research theatre. This is the research-to-decision spine; everything traces back to it.372. **Select the method.** Use `references/research-method-selector.md` to pick on two axes — *generative vs. evaluative* and *attitudinal (what they say) vs. behavioural (what they do)*. Behaviour beats opinion when they disagree. Prefer the cheapest method that actually answers the question; do not run a survey to answer a "why".383. **Plan the study.** Write a one-page plan: decision, question, method, participants (number + screener), tasks/topics, success metrics, schedule, and what result would change the decision in each direction. Pre-committing the decision rule prevents post-hoc rationalisation.394. **Recruit honestly.** Screen for real target users; screen *out* friends, colleagues, and people who can guess the hypothesis. State the incentive. Get informed consent and recording permission in writing — non-negotiable for any session.405. **Run the session.** For usability tests follow `references/usability-test-protocol.md` exactly: neutral intro, realistic task scenarios (never instructions that name the UI element), think-aloud, no leading or rescuing, observe behaviour first and ask "why" after. For interviews: open questions, past-behaviour over hypotheticals ("tell me about the last time…"), embrace silence, never pitch.416. **Capture observations, not conclusions.** Record what happened (quote, action, where they got stuck) separately from your interpretation. Tag each with the participant ID so a claim can be traced to evidence later.427. **Synthesize.** Cluster observations into themes (affinity mapping). For usability issues, rate severity (frequency × impact × persistence) so the team fixes the right things first. Quantify where honest to (task success rate, time, error count, SUS); never fabricate precision from a 5-person sample.438. **Convert findings to decisions.** For each theme write: the finding → the evidence (participant count + quotes) → the recommended decision → confidence. Close the loop back to step 1. A finding with no recommended action is incomplete.449. **Apply the doctrine lens.** Research is also how you defend that the product looks *authored*, not templated (`doctrine/design-doctrine.md` §0). Usability findings that surface "this feels generic / I didn't trust it" are slop signals (`doctrine/references/ai-slop-taxonomy.md`), not just nuisance comments — escalate them.4510. **Check accessibility coverage.** If the test never included a keyboard-only, screen-reader, low-vision, or motor-impaired participant, state that as a known gap; pair findings with `doctrine/references/wcag-2.2-criteria.md` before claiming the experience "works for users".4647## Decision Rules4849| Condition | Choice | Wrong-choice failure |50|---|---|---|51| Need to understand why/mental model | Interview or moderated study | Survey percentages cannot explain mechanism |52| Need to observe task usability | Behavioural usability test | Preference questions substitute opinion for performance |53| Navigation findability is the question | Tree test before visual prototype | Visual styling confounds information-architecture evidence |5455## Capability Contract5657- Must protect consent, privacy, recruitment fairness, and raw evidence; planning/review is read-only unless study execution is authorised.58- Do not contact participants, record sessions, spend incentives, or publish identifiable data without explicit authority and approved handling.5960## Degraded Mode6162- If decision, population, consent, or data-handling rules are missing, stop recruitment/collection and return the blockers.63- Without participant access, deliver a validated protocol or secondary-evidence synthesis, not fabricated findings. Recover weak sessions by documenting deviation, excluding compromised evidence where necessary, and revising the protocol.6465## Quality Standards6667- Findings separate observation, interpretation, prevalence, and limitation; claims do not exceed sample or method.68- Every recommendation traces to evidence and names the decision, owner, confidence limits, and next validation.6970## Anti-Patterns71- **Research theatre** — running a study whose result cannot change any decision. The most common and most expensive mistake.72- **Leading the witness** — "Don't you find this easy?", naming the button in the task, nodding at the answer you want. Invalidates the data.73- **Rescuing the participant** — jumping in the moment they struggle. The struggle *is* the finding.74- **Asking people to predict their behaviour** — "Would you use this?" / "How much would you pay?" Self-reported future behaviour is near-worthless; ask about the last real instance instead.75- **Confirmation harvesting** — recruiting fans, cherry-picking the quote that fits the roadmap, ignoring the disconfirming participant.76- **Sample-size theatre** — reporting "80% of users" from 5 people, or running 30 sessions when 5 would have surfaced the same top issues (Nielsen's diminishing-returns curve).77- **The shelved report** — a 40-slide deck with no recommended decision. Findings without "so we should…" are inert.78- **Opinion laundering** — presenting a designer's preference as a "research finding" with no traceable evidence.7980## Outputs8182| Output | Consumer | Evidence and acceptance |83|---|---|---|84| Research plan/protocol | Sponsor and research operations | Question, method, sample, tasks, consent, analysis, and stop rules are complete |85| Findings and decision trace | Product/design | Evidence references, themes, severity, limitations, decisions, and owners are explicit |86- A one-page research plan (decision → question → method → participants → success criteria → decision rule).87- A screener and a moderation guide / interview guide.88- A usability-test protocol with task scenarios and pre-stated success criteria.89- A synthesis: themes, severity-rated issues, quantified metrics where honest, traced to participant evidence.90- A research-to-decision summary: each finding paired with a recommended decision and a confidence level.9192## Examples93- See `examples/research-plan-and-synthesis.md` — a real, worked study for an onboarding flow: the plan, the moderated unmoderated mix, the raw observations, the affinity synthesis with severity ratings, and the research-to-decision table that changed the roadmap. Never lorem.9495## References96- `references/research-method-selector.md` — choose the method by generative/evaluative × attitudinal/behavioural; cost vs. answer-fit.97- `references/usability-test-protocol.md` — a real moderated + unmoderated protocol: tasks, think-aloud script, success criteria, severity rating, SUS.98- `doctrine/design-doctrine.md` — §0 Mission (the authored-not-templated moat) and the Anti-Slop Charter; usability findings are evidence the product reads as human-made.99- `doctrine/references/ai-slop-taxonomy.md` — the interface/product slop tells; "feels generic / didn't trust it" findings map here.100- `doctrine/references/wcag-2.2-criteria.md` — accessibility coverage gate; a study without disabled participants is incomplete, not done.101- Sibling skills: `05-ux-process-research-and-psychology/ux-psychology` (theory to apply), `enterprise-ux-process` (the surrounding process), `wireframing-and-prototyping` (what you put under test).102<!-- dual-compat-end -->