UX Usability Study Kit
This skill helps you run a moderated usability study end to end: figure out what you're trying to learn, recruit the right people, prepare everything you say and hand to a participant, take notes you can actually use, and turn raw observations into findings a team will act on.
A usability test is simple in spirit — you give real, representative people realistic tasks and watch where the product helps or fails them — but it goes wrong in predictable ways: fuzzy goals, the wrong participants, leading questions, tasks that hand users the answer, and notes that can't be synthesized afterward. The guidance below exists to avoid those failure modes, not to impose ceremony. Scale it to the study: a scrappy five-person test of one flow needs a fraction of what a formal benchmark study needs.
How to work through a study
Usability work is a pipeline, and each artifact feeds the next — success metrics shape the screener, the screener shapes recruiting, the tasks come from the research questions, and the note structure comes from the tasks. So don't jump straight to writing a script. Establish the foundation first, then generate the downstream materials from it.
Ask the user which pieces they need. Most people want either the full kit or one specific artifact ("just write me a screener"). If they want one piece, still ground it in the foundation questions below — a screener written without knowing who the target user is will be generic and useless.
1. Establish the foundation (always do this first)
Before writing anything, pin down four things. If the user hasn't supplied them, ask — briefly, in one message, not as an interrogation.
- What decision does this study inform? A usability test is worth running when the results will change what someone builds. "Should the new nav ship as is?" is a decision. "Let's get some feedback" is not. Anchoring on the decision keeps the whole study focused.
- What are the research questions? The 3–6 specific things you need to learn, phrased as questions ("Can a new user find where to add a payment method without help?"). Everything downstream traces back to these.
- Who is a representative user? The characteristics that make someone a valid test participant — role, relevant experience, tools they use, context. This becomes the screener.
- What does success look like, and is it measurable? See the metrics section below.
2. Define success metrics
Vague goals produce vague findings. Where it makes sense, attach a measurable target to a task so the team agrees in advance what "good enough" means. The common, well-worn usability metrics are:
- Success rate — did the participant complete the task, partially complete it, or fail? (The most important and most universal metric.)
- Time on task — how long completion took.
- Errors — wrong turns, mistakes, and recoveries along the way.
- Satisfaction — the participant's subjective ease rating after a task or the whole session.
Write targets at the right altitude. A task-level target reads like: "At least 4 of 5 participants complete 'add a new payment method' unaided in under 90 seconds." Don't over-quantify a small qualitative study — with five participants, patterns matter more than precise percentages, and a false air of statistical precision misleads. Use metrics to sharpen focus, not to dress up five data points as significance.
3. Recruiting screener
The screener is a short questionnaire (usually run by phone, form, or a recruiting panel) that qualifies people against your representative-user definition. A good screener:
- Opens with the logistics (what the session is, how long, incentive, that it's a one-on-one session and not a sales call, that you're testing the product and not them).
- Disguises the criteria you're screening for. If you ask "Do you manage a team's budget? (yes qualifies)" people say what gets them the incentive. Ask behavioral questions with plausible distractor answers instead, so the qualifying answer isn't obvious.
- Screens out the wrong fit as much as it screens in the right one — e.g. exclude people who work in the industry you're studying if they'd behave unlike real users, or competitors' employees.
- Aims for a mix, not clones — a spread of ages, experience levels, and backgrounds within your target, unless the study specifically targets one segment.
- Collects the practical stuff: availability, device/browser if remote, whether they'll need any assistive technology, consent to be recorded.
For a standard qualitative study, five participants per distinct user group is a sensible default — it surfaces the large majority of the serious issues while keeping recruiting and moderation manageable. Say so if the user hasn't picked a number, but treat it as a rule of thumb, not a law.
4. Consent form
A short, plain-language form the participant signs (or agrees to) before the session. It should cover: what they'll be asked to do, that participation is voluntary and they can stop at any time, what's being recorded (screen, voice, face), how the recording will be stored and for how long, that their identity will be kept confidential and not tied to the recording, and who to contact with questions. Keep it human and readable — a wall of legalese makes people nervous right before a session. Flag clearly that the specifics (retention period, organization name, whether video is redacted) are placeholders the user's legal/privacy owner should confirm; don't invent binding legal terms.
5. Facilitator script
This is what the moderator says, start to finish. Its job is to make the participant comfortable and honest while keeping the moderator from accidentally biasing them. Structure:
- Intro / framing. Thank them, set expectations, and say the two things that most improve data quality: we're testing the product, not you — there are no wrong answers, and if something is confusing, that's useful information about the design, not a failing on your part.
- Think-aloud protocol. Ask them to narrate their thoughts, expectations, and reactions as they work — "tell me what you're looking at, what you expect to happen, what's going through your head." This is the single richest source of insight in a usability test. Note that people go quiet when concentrating; the moderator's main job is to gently prompt ("what are you thinking now?") without leading.
- Warm-up questions. A few easy background questions to relax them and gather context about how they'd normally approach this kind of task.
- Tasks (see below).
- Wrap-up. Broad reflection questions, a satisfaction rating, space for anything they want to add, and thanks.
The moderator's discipline is worth calling out in the script as reminders: don't answer "where would you click?" questions (turn them back: "what would you try?"), don't rescue a struggling participant too early (struggle is data), don't explain how the product works, and avoid leading or loaded phrasing.
6. Task scenarios
Tasks are the heart of the test, and writing them well is a craft. Each task should:
- Be a realistic scenario, not an instruction. Give the participant a goal and a reason, framed in their world — "You've just started a new job and need to expense your first client lunch. Show me how you'd do that." — not "Click the Expenses button, then New."
- Never contain the interface's own words. If the button says "Reimbursement" and your task says "get reimbursed," you've handed them the answer and tested nothing. Describe the goal in the user's language.
- Have a clear completion state so the moderator and notetaker agree on whether it succeeded, and map back to a research question.
- Be ordered sensibly — usually simplest/most-common first to build confidence, with independent tasks so a failure on one doesn't block the next.
For each task, prepare the moderator-facing detail: the exact wording read to the participant, the expected success path(s), what counts as success, and which research question it answers.
7. Note-taking structure
Notes taken as freeform prose can't be synthesized. Give the team a consistent structure so observations from different sessions and notetakers line up. A practical row-per-observation format captures, for each noteworthy moment:
- Timestamp (or elapsed time) so you can find the clip later.
- Task it happened in.
- The observation — what the participant did or said (behavior and verbatim quotes, kept separate from interpretation).
- Severity / flag — a quick marker for "this is a serious problem" vs. a minor nit, and a flag for direct quotes worth reusing.
Encourage recording what happened, not just conclusions — "hesitated 20s on the totals screen, said 'I don't know if this includes tax'" is reusable; "totals screen is confusing" has already thrown away the evidence.
8. Findings report / synthesis
After the sessions, turn observations into a prioritized, actionable report. The synthesis process:
- Cluster observations across participants into distinct issues. An issue that hit 4 of 5 participants is very different from a one-off.
- Rate severity by combining how badly it hurt the user (blocked the task? annoyed them? cosmetic?) with how many participants hit it and how frequent the path is. A simple high/medium/low is usually enough; explain the reasoning.
- Tie each finding to evidence — the participants and moments that show it, ideally with a quote — so the team trusts it and can watch the clip.
- Recommend, don't just diagnose. Pair each significant issue with a concrete design direction, even a tentative one.
- Report against the metrics you set in step 2 (success rates, times) so the original decision can actually be made.
Suggested report shape:
# Usability Study: [product / flow] — Findings
## Study at a glance
What we tested, who with, when, method
## What we wanted to learn
The research questions and the decision this informs
## Results against our success metrics
Task-by-task success rates / times vs. targets
## Findings (most to least severe)
For each: severity, what happened, who it affected, evidence/quote, recommendation
## What to do next
Prioritized recommendations
Producing the artifacts
Generate whatever the user asks for as clean, ready-to-use documents. Default to
Markdown so the content is easy to review and edit. If the user wants
hand-out-ready files — a screener to send a recruiter, a consent form to print,
a note-taking sheet — offer to produce them as .docx or .xlsx (a spreadsheet
is the natural home for the note-taking and issue-tracking structures). Use
clearly marked [placeholders] for anything project-specific (dates, incentive
amounts, product names, legal specifics) rather than inventing details, and keep
the wording original and plain.
Bundled tools
scripts/sus.py— score the System Usability Scale from participants' 10-item responses (the fixed +/− and ×2.5 math) with an interpretation against the ~68 benchmark. E.g.python scripts/sus.py 4 2 4 2 5 1 4 2 5 1.scripts/discovery_rate.py— the problem-discovery curve1−(1−L)^n; use it to justify a sample size and show diminishing returns instead of just asserting "five users is enough." E.g.python scripts/discovery_rate.py --users 5.
Prefer running these over doing the arithmetic by hand — both formulas are easy to get subtly wrong.
Pairs well with
- persona-builder — define the representative users you recruit and write tasks for.
- heuristic-evaluation / accessibility-review — expert audits that complement testing with real users.
- inclusive-design — check the flow and its copy for exclusion before you test.
Sources
This methodology reflects widely-established usability practice (Nielsen Norman Group, usability.gov, the broader UX field) and is written to be shared freely. The only caution is about other people's material: don't reproduce a specific organization's copyrighted templates verbatim, and don't present another company's confidential research as an example.