At 10:12, an agent labels a conversation “billing,” “refund,” “urgent,” and “angry customer.” At 10:18, another agent handles the same situation but chooses only “payment problem.” Both cases are resolved correctly. At the end of the month, the report says refund demand fell while payment problems rose.
Nothing operational changed. The vocabulary changed inside the data.
That is why customer conversation tagging is not clerical decoration. A label can drive a queue, trigger an automation, define a quality sample, or become the denominator in a management report. If people apply overlapping labels with different meanings, every downstream decision inherits the ambiguity.
The remedy is not a longer tag list. It is a small governed model that gives each kind of information a proper home. Keep durable customer facts in fields, changing work conditions in statuses, final results in outcomes, and reserve tags for flexible cross-cutting observations.
A tag should answer one reusable question
Before adding a tag, write the question it will answer. “Which conversations concerned a delivery address error?” is useful because a team can investigate causes and change a form or process. “Which conversations felt difficult?” is not yet useful because two agents may interpret difficult differently and no decision is attached.
A good label has a named user, a repeatable definition, and an action. The user might be a support lead planning training, a product manager finding defects, or an operations analyst checking demand. The action might be reviewing ten cases, changing a routing rule, or comparing an outcome before and after a release.
Reject labels that merely repeat the channel, current owner, or ticket status when those values already exist elsewhere. Reject labels created for one presentation unless someone will maintain them. If a question matters only once, answer it with a temporary analysis rather than permanently expanding the operating vocabulary.
This test makes tag creation slightly slower and every later report faster to trust.
Separate facts states outcomes and tags
Most tag sprawl comes from putting four different data types into one flat list.
Durable customer facts describe the customer or account beyond one conversation. Language, market, plan, consent state, account tier, and acquisition source usually belong in customer fields. A fact can change, but it should have one current authoritative value rather than several contradictory tags.
Changing workflow states describe where the work is now. Waiting for customer, assigned, escalated, pending review, and closed are states. Store them in status, owner, queue, or workflow fields so a transition replaces the previous value and carries a timestamp.
Final outcomes describe what happened. Refunded, address corrected, duplicate request, no response, and policy exception are examples. An outcome should be recorded at completion and governed independently from the issue that started the conversation.
Flexible tags capture a cross-cutting observation that may apply alongside other information. Examples include release related, accessibility feedback, competitor mention, or suspected documentation gap. Multiple tags can coexist because they are not trying to represent one exclusive state.
The distinction matters. If “escalated” remains a tag after the escalation ends, a later report cannot tell whether the case is currently escalated, was once escalated, or was tagged by mistake. If “enterprise” is both a CRM field and a conversation tag, the two can disagree. If “refund” means both a request and a completed outcome, the report cannot measure conversion from request to refund.
Write a small vocabulary before adding rules
Start with the decisions the team makes, not with every phrase customers use. Group the decisions into a few families such as contact reason, operational cause, risk signal, product feedback, and research cohort. Then decide which families should actually be fields or outcomes rather than tags.
For every surviving tag, document five things:
- the exact label and formatting;
- one sentence defining when to apply it;
- two positive examples;
- one close example that must not receive it;
- the owner and next review date.
Formatting is not cosmetic. Zendesk documents that tags add context and can be used in views, search, and business rules. It also warns that underscores, dashes, and slashes produce different tags. Pick one connector, one singular or plural convention, and one language for system identifiers.
Keep the visible vocabulary plain enough for an agent to choose under time pressure. Hierarchy can help administrators without forcing people to scan a giant list. Microsoft explains that hierarchical categories can support grouping, search, reporting, sorting, and segmentation. Use that structure to narrow choices, not to create decorative depth.
Do not set a universal maximum number of tags. A specialist operation may need more than a small retail team. The useful limit is cognitive: an agent should see a short relevant set for the current type of work, and every active label should still have an owner and a decision attached.
Control who can create and change tags
Letting every user invent labels feels flexible until spelling variants, private shorthand, and near duplicates reach reports. A better arrangement separates suggestion from publication. Agents can propose a missing label with an example. A small owner group checks whether an existing field, state, outcome, or tag already answers the need.
Approve a new tag only with a definition, owner, affected reports, automation impact, and review date. Changes to tags used in routing or automation need the same caution as changes to workflow rules. A rename may split history. A deletion may stop a trigger. A broader definition may make this month incomparable with last month.
Keep a simple change log with the old value, new value, reason, effective time, and affected dashboards. The log is part of the metric definition, not administrative paperwork.
Test the labels on real conversations
A taxonomy can look clear in a meeting and fail during actual work. Test it before connecting it to important automation.
Select a varied sample of recent conversations with personal details removed from the review material. Ask two reviewers to label each case independently using only the written guide. Compare their exact choices and discuss every disagreement. Do not begin by arguing about a universal agreement target. First find why reasonable people differed.
Common causes are overlapping definitions, missing exclusions, tags that require information the agent cannot see, and labels that mix reason with outcome. Rewrite the guide, merge indistinguishable labels, and test a fresh sample. Also record “not enough information” as a valid finding. Forcing a guess produces tidy completion rates and unreliable data.
After launch, sample cases by tag and by agent. Look for labels that appear only when one person is working, tags that suddenly vanish, and combinations that should be impossible. These patterns are often definition or interface problems before they are performance problems.
Clean the taxonomy without rewriting history
Begin cleanup by freezing uncontrolled creation. Export active tags and their usage counts, but do not assume the most common labels are the most useful. Mark each as keep, merge, move to field, move to state, move to outcome, or retire.
Create a mapping from every retired value to its successor and give the change an effective date. Preserve the original value in historical data. If management needs a continuous trend, build a reporting layer that groups old and new values under a stable concept, and label the definition version used. Do not silently rewrite old records and then compare the result with a report produced under the former rules.
Retirement also needs operational checks. Search views, macros, routing rules, automations, quality forms, exports, and dashboards for the old label. Zendesk’s tag management guidance notes that rules can add, remove, or set tags, which is precisely why deleting a visible label is not enough. Remove or update each dependency deliberately.
Review the vocabulary on a schedule and after meaningful product or policy changes. A tag without recent use may be obsolete, too obscure, or hidden from the correct team. Investigate before removing it.
Use tags as evidence not as decoration
A useful report shows the definition and denominator beside the number. Instead of “refund tags increased 18 percent,” state whether the tag means refund requested, refund approved, or refund completed, and whether the denominator is conversations, customers, orders, or resolved cases. Compare only periods using compatible definitions.
Read tagged totals with outcomes. A rise in documentation gap tags followed by fewer repeat contacts after an article update is a stronger operational story than the tag count alone. A fall might mean the problem improved, the tag moved, or agents stopped applying it. Sample real conversations before choosing the explanation.
In DripTell, the omnichannel inbox keeps assignment, status, notes, channel, and customer context with the conversation, while the customer CRM provides custom fields, tags, source, lead stage, and ownership. That gives each part of the model an appropriate place. The design work still belongs to the operating team: decide which values are authoritative, who may change them, and what each report will trigger.
Start with the twenty labels used most often in the last month. Move customer facts and work states out of the tag list, define the remaining labels, and test them on real cases. If you want to map the model to your routing and reporting workflow, talk with the DripTell team.
Common questions about conversation tagging
Should every conversation have a tag
No. Required fields, status, owner, and outcome may already capture everything needed. Apply a tag only when it answers a defined cross-cutting question. A compulsory meaningless tag encourages guesses and makes completeness look better than accuracy.
Can automation apply tags
Yes, when the rule uses observable evidence and its errors are reviewed. Record whether a person or rule applied the label, test precision on a sample, and give uncertain cases a review path. Do not let an inferred tag silently become a durable customer fact.
How often should tags be reviewed
Review usage and disagreements monthly during a cleanup period, then choose a cadence that matches how quickly products and policies change. Always review a tag when its definition, automation dependency, or reporting role changes.
DripTell Editorial
Practical guidance reviewed by the DripTell product and customer workflow team.
See how DripTell checks product claims, uses primary sources and handles corrections.
Editorial and source policy



