Want help applying this guide?Ask the DripTell team
A support manager opens ten conversations before calibration. Nine have low satisfaction scores and the tenth came from a new agent. The review may find real problems, but it cannot describe normal service quality because the selection was designed to find trouble.
A useful customer service QA sampling plan needs two separate lanes. Use a random baseline to understand everyday quality. Use a targeted risk lane to find complaints, policy exceptions, new automation, repeated transfers, or other cases that deserve attention. Review both with the same standards, but never blend their results into one score.
Start with the question the sample must answer
Sampling begins before anyone opens a conversation. Write down the question first.
Are you estimating overall quality, checking a refund rule, looking for unsafe AI replies, or coaching a new teammate? Each question needs a different selection rule.
Define the review period, included channels, eligible teams, languages, issue types, and closed or open status. Record exclusions too. If voice calls are missing because they are stored elsewhere, say so. A clean shared inbox can make the population easier to reconstruct, but the reviewer still has to define what belongs in it.
Keep a random baseline and a risk lane
The baseline should be boring. Draw it randomly from the full eligible population after exclusions. NIST explains that in a simple random sample, every response in the population has an equal chance of selection. Random selection does not guarantee that one small sample perfectly represents the operation, but it avoids choosing only familiar, dramatic, or easy cases on average (NIST).

- 1Define the questionState what the review must estimate or investigate.
- 2Freeze the populationRecord the period, channels, teams, statuses, and exclusions.
- 3Draw the random baselineGive every eligible conversation the same chance of selection.
- 4Add the risk laneSelect complaints, exceptions, changes, and other named signals.
- 5Report separatelyUse one scorecard but preserve each lane's denominator and result.
The risk lane is deliberately biased. That is useful when the bias is explicit. Pull conversations involving complaints, vulnerable customers, policy overrides, a new workflow, failed handoffs, unusual refunds, low confidence AI answers, or a sudden change in customer sentiment. Microsoft's current evaluation-plan documentation likewise separates the records identified by conditions from the records selected by sampling, which makes the selection rule visible (Microsoft Learn).
Do not average the two lanes together. If half the review set was selected because something already looked wrong, the combined failure rate is not an estimate of overall quality.
| Sampling lane | What it answers | Selection rule | Honest reporting |
|---|---|---|---|
| Random baseline | What normal service looks like | Random draw from the defined population | Trend as its own rate with the population and period |
| Targeted risk | Where serious failure may hide | Explicit risk or learning signals | Count findings by selection reason without claiming prevalence |
| Change cohort | Whether a new rule or automation behaves as intended | Cases exposed to the named change | Compare the bounded cohort before and after the change |
| Recovery check | Whether a known defect stayed fixed | Cases matching the previous failure pattern | Report recurrence inside that pattern only |
Build the sampling frame before choosing cases
The sampling frame is the list of conversations that could be selected. Weak QA programs often skip this step and let a dashboard's default view become the frame. That can silently exclude archived chats, overnight work, certain channels, unresolved cases, or conversations owned by another team.
Export the eligible IDs first and freeze the list for the review period. Use a reproducible random method for the baseline. Store the date, filter, selection record, and reviewer. The aim is to explain why these cases were reviewed and draw a comparable sample next week.
Stratify only when the overall population would hide an important group. For example, a small Arabic support queue or a newly launched messaging channel may disappear from a simple draw. In that case, draw randomly inside each declared group and report the groups separately or weight them correctly. Never quietly oversample a group and then present its raw score as the company average.
Review both lanes with one scorecard
Selection and scoring are different controls. Once a conversation enters review, apply the same observable criteria regardless of agent, channel, or why it was selected. The existing QA scorecard method is useful here because it focuses attention on resolution, accuracy, ownership, policy, and customer effort rather than style preferences.
Calibrate reviewers with a few shared cases before the cycle. If two reviewers disagree, discuss the rule and evidence before treating the difference as a performance issue. A rule that repeatedly needs private interpretation is not stable enough for coaching or trend reporting.
Automated screening can help sort a large population, especially when workflow rules or AI assistance are involved. It should not erase the random lane. A model trained to find known risks can miss quiet, ordinary failures that were never labelled.
Report the lanes separately
A QA report should show the population, sample counts, method, exclusions, scorecard version, reviewer agreement, and findings for each lane. The random baseline describes normal work cautiously. The targeted lane lists risks and learning opportunities.
Break patterns down by issue type, channel, workflow version, handoff, or policy when the groups are large enough to be useful. Avoid agent league tables from tiny samples. Three reviewed conversations can start a coaching conversation, but they cannot prove that one person is worse than the team.
Use inbox reporting to locate operational patterns, then return to the conversations for evidence. Protect access to recordings, transcripts, and reviewer notes through appropriate security controls. QA needs enough context to judge the interaction, not unlimited access to every customer detail.
Change the plan when the operation changes
Keep the core baseline stable long enough to show a trend. Add temporary targeted cohorts when a new channel, policy, team, or automation changes the work. Close those cohorts when the question is answered, or they will quietly become permanent bias.
The most practical review meeting ends with one owned action. It might clarify a policy, change a routing rule, repair knowledge, or coach a judgment call. Then sample again. DripTell's support workspace can keep the conversation history and owner beside that follow-up, but the sampling discipline belongs to the team.
Good QA does not try to prove that service is good. It creates a trustworthy baseline and a sharper place to look when risk appears.
Frequently Asked Questions
How many conversations should customer service QA review
There is no universal number. Choose enough random cases to support the decision you need to make, considering population size, variation, and desired precision. Treat small samples as learning evidence, not precise companywide estimates.
Should managers choose their own QA samples
Managers can nominate targeted cases for coaching or risk review, but those cases should stay outside the random baseline. Otherwise personal visibility and recent incidents will distort the quality trend.
Can AI replace random QA sampling
AI can screen more conversations and surface known risk signals. Keep a random human-reviewed baseline to catch unlabelled failure, check the screening system, and avoid confusing model attention with representative evidence.
DripTell Editorial
Practical guidance reviewed by the DripTell product and customer workflow team.
See how DripTell checks product claims, uses primary sources and handles corrections.
Editorial and source policy



