Sampling conversations
Every manual review of support conversations is a sample, including the ones nobody
calls a sample. The question is only whether it was drawn deliberately.
This is about how to draw it and how to weight the results. How many you need for a
given precision is a different question — a sample can be perfectly sized and still
answer about the wrong population.
The frame comes first
The sampling frame is the list of conversations you could have drawn from. Almost
every generalisation failure is a frame problem, not a size problem.
Write down explicitly:
- What is in the frame — window, channels, queues, statuses, languages.
- What is silently missing. Conversations not synced from a channel; deleted,
merged or redacted records; anything the API does not return; a channel connected
halfway through the window. Each of these removes a non-random slice.
- What you excluded on purpose, and the rule.
A frame that omits a channel cannot support any organisation-wide claim, however
large the sample. Say which population your result actually describes — that sentence
is the deliverable as much as the number is.
Choose the design to match the question
Simple random. The default, and correct when you want one overall estimate. Boring
and defensible. Use it unless you have a specific reason not to.
Stratified. Split the frame into strata — channel, queue, team, language, risk band
— and sample within each. Use it when:
- you need per-stratum estimates, not just an overall one (almost always true for QA);
- strata differ a lot in the thing you are measuring, which improves precision for the
same total sample;
- some stratum is small but important and simple random would return three of them.
Sample small strata at a higher rate, then weight the results back by the
stratum's true share when computing an overall figure. This is the step that gets
skipped, and skipping it means the overall number silently over-weights the segments
you oversampled. If you cannot weight, report per-stratum results only and no overall
figure.
Risk-weighted / purposive. Deliberately over-sample high-risk conversations —
complaints, low scores, regulated topics, vulnerable customers. Correct for finding
problems, and it cannot produce a rate that describes the population. A defect rate
from a risk-weighted sample describes risky conversations, and reporting it as the
overall defect rate is a serious and common error.
Run both when both questions matter: a random core for estimating, plus a risk-weighted
supplement for finding. Keep them separate in the analysis and never pool them into one
rate.
Census. For rare, high-severity categories — regulatory complaints, safeguarding —
review everything. Sampling a population of forty to save effort is a false economy.
Randomise properly
- Use a real random draw over the frame, with a recorded seed so the sample is
reproducible and auditable.
- "The first 100 returned" is not random. API results carry an implicit order —
usually recency or id — and both correlate with the things you are measuring.
- Neither is "conversations from last Tuesday". Day of week and time of day
correlate strongly with staffing, contact mix and quality.
- Sample the whole window, not a convenient slice of it.
- Record the seed, the frame definition and the draw date. Someone will ask how the
sample was picked, and "randomly" is not an answer that survives an audit.
Sampling for QA coverage
Ongoing QA sampling is the same problem with an extra constraint: it must also be
fair to the people being measured.
- Every agent needs enough evaluations for the use. If scores drive coaching only,
a handful is fine. If they drive ranking or pay, they need the volume that supports
it — and if the programme cannot afford that volume, the honest conclusion is that
the scores may not be used for ranking.
- Equalise by agent, not by ticket, when the purpose is per-agent measurement.
Ticket-proportional sampling gives high-volume agents narrow intervals and low-volume
agents useless ones, and then compares them.
- Do not let reviewers self-select which conversations to review. Reviewer choice
is the largest single source of bias in QA programmes, and it is invisible in the
output.
- Check what the sampler actually did. Compare the composition of what was reviewed
against the composition of what was eligible, periodically. Samplers drift, rules get
edited, and a channel or team silently drops out. This check is cheap and it
invalidates comparisons when it fails.
Traps
- Post-hoc filtering breaks the sample. Drawing 200 and then analysing "the ones
that were interesting" produces a purposive sample with a random sample's confidence
intervals attached. If you filter after drawing, report the filter and treat the
result as purposive.
- Replacing unavailable items non-randomly. If a drawn conversation cannot be
reviewed — unavailable transcript, wrong language — record it as a non-response and
either replace it with another random draw or report the non-response rate. Quietly
picking the next one down the list re-introduces the ordering bias.
- Reviewing until you find something. That is search, not sampling, and it produces
no rate at all.
- Stale frames. A frame built a week ago no longer matches the population.
- Language. A sample drawn without regard to language, reviewed only by reviewers
who speak one of them, silently becomes a single-language sample.
Present results to the user
- The frame — what was in it, what is silently missing, what you excluded.
- The design — simple, stratified, risk-weighted or census — and why it matches the
question.
- The draw — seed, date, sizes per stratum, and the sampling rate each implies.
- Weighting, if strata were sampled at different rates, and the weighted overall
estimate. Or a plain statement that no overall estimate is available.
- Non-response — drawn items that could not be reviewed, and how they were handled.
- The population your result actually describes, in one sentence. This is the
sentence that stops a risk-weighted finding being quoted as an overall rate.
1---2name: cx-conversation-sampling3description: Use to draw a defensible sample of support conversations for QA review, an audit or a manual analysis, so the results generalise to the population rather than to whatever was easy to pull. Trigger for "which tickets should we review", "how do we pick a sample", "is our QA sampling representative", stratified or risk-based sampling, review coverage design, or a finding based on a handful of hand-picked tickets.4---56# Sampling conversations78Every manual review of support conversations is a sample, including the ones nobody9calls a sample. The question is only whether it was drawn deliberately.1011This is about **how to draw it and how to weight the results**. How many you need for a12given precision is a different question — a sample can be perfectly sized and still13answer about the wrong population.1415## The frame comes first1617The **sampling frame** is the list of conversations you could have drawn from. Almost18every generalisation failure is a frame problem, not a size problem.1920Write down explicitly:2122- **What is in the frame** — window, channels, queues, statuses, languages.23- **What is silently missing.** Conversations not synced from a channel; deleted,24 merged or redacted records; anything the API does not return; a channel connected25 halfway through the window. Each of these removes a non-random slice.26- **What you excluded on purpose**, and the rule.2728**A frame that omits a channel cannot support any organisation-wide claim**, however29large the sample. Say which population your result actually describes — that sentence30is the deliverable as much as the number is.3132## Choose the design to match the question3334**Simple random.** The default, and correct when you want one overall estimate. Boring35and defensible. Use it unless you have a specific reason not to.3637**Stratified.** Split the frame into strata — channel, queue, team, language, risk band38— and sample within each. Use it when:3940- you need per-stratum estimates, not just an overall one (almost always true for QA);41- strata differ a lot in the thing you are measuring, which improves precision for the42 same total sample;43- some stratum is small but important and simple random would return three of them.4445**Sample small strata at a higher rate**, then **weight the results back by the46stratum's true share** when computing an overall figure. This is the step that gets47skipped, and skipping it means the overall number silently over-weights the segments48you oversampled. If you cannot weight, report per-stratum results only and no overall49figure.5051**Risk-weighted / purposive.** Deliberately over-sample high-risk conversations —52complaints, low scores, regulated topics, vulnerable customers. Correct for finding53problems, and **it cannot produce a rate that describes the population**. A defect rate54from a risk-weighted sample describes risky conversations, and reporting it as the55overall defect rate is a serious and common error.5657Run both when both questions matter: a random core for estimating, plus a risk-weighted58supplement for finding. Keep them separate in the analysis and never pool them into one59rate.6061**Census.** For rare, high-severity categories — regulatory complaints, safeguarding —62review everything. Sampling a population of forty to save effort is a false economy.6364## Randomise properly6566- **Use a real random draw over the frame**, with a recorded seed so the sample is67 reproducible and auditable.68- **"The first 100 returned" is not random.** API results carry an implicit order —69 usually recency or id — and both correlate with the things you are measuring.70- **Neither is "conversations from last Tuesday".** Day of week and time of day71 correlate strongly with staffing, contact mix and quality.72- **Sample the whole window, not a convenient slice of it.**73- **Record the seed, the frame definition and the draw date.** Someone will ask how the74 sample was picked, and "randomly" is not an answer that survives an audit.7576## Sampling for QA coverage7778Ongoing QA sampling is the same problem with an extra constraint: it must also be79*fair to the people being measured*.8081- **Every agent needs enough evaluations for the use.** If scores drive coaching only,82 a handful is fine. If they drive ranking or pay, they need the volume that supports83 it — and if the programme cannot afford that volume, the honest conclusion is that84 the scores may not be used for ranking.85- **Equalise by agent, not by ticket**, when the purpose is per-agent measurement.86 Ticket-proportional sampling gives high-volume agents narrow intervals and low-volume87 agents useless ones, and then compares them.88- **Do not let reviewers self-select** which conversations to review. Reviewer choice89 is the largest single source of bias in QA programmes, and it is invisible in the90 output.91- **Check what the sampler actually did.** Compare the composition of what was reviewed92 against the composition of what was eligible, periodically. Samplers drift, rules get93 edited, and a channel or team silently drops out. This check is cheap and it94 invalidates comparisons when it fails.9596## Traps9798- **Post-hoc filtering breaks the sample.** Drawing 200 and then analysing "the ones99 that were interesting" produces a purposive sample with a random sample's confidence100 intervals attached. If you filter after drawing, report the filter and treat the101 result as purposive.102- **Replacing unavailable items non-randomly.** If a drawn conversation cannot be103 reviewed — unavailable transcript, wrong language — record it as a non-response and104 either replace it with another random draw or report the non-response rate. Quietly105 picking the next one down the list re-introduces the ordering bias.106- **Reviewing until you find something.** That is search, not sampling, and it produces107 no rate at all.108- **Stale frames.** A frame built a week ago no longer matches the population.109- **Language.** A sample drawn without regard to language, reviewed only by reviewers110 who speak one of them, silently becomes a single-language sample.111112## Present results to the user1131141. **The frame** — what was in it, what is silently missing, what you excluded.1152. **The design** — simple, stratified, risk-weighted or census — and why it matches the116 question.1173. **The draw** — seed, date, sizes per stratum, and the sampling rate each implies.1184. **Weighting**, if strata were sampled at different rates, and the weighted overall119 estimate. Or a plain statement that no overall estimate is available.1205. **Non-response** — drawn items that could not be reviewed, and how they were handled.1216. **The population your result actually describes**, in one sentence. This is the122 sentence that stops a risk-weighted finding being quoted as an overall rate.