Taxonomy Drift. Why the Same Customer Problem Becomes Forty Different Insights

Chris Lloyd • 2026-09-21
#product #customer-success #marketing #executive #sales

A support lead pastes a batch of tickets into Claude and asks it to pull out the recurring pain points.

It does cleanly, in seconds, with sensible-sounding labels. She does the same thing again next week with a fresh batch. The labels are different. Not wrong, exactly, "onboarding friction" instead of "setup confusion," "pricing objection" instead of "cost concern", just different enough that the two batches can't be compared. Nobody notices for a while, because each individual output still looks great. That's the trap, taxonomy drift doesn't announce itself with an error message. It shows up months later, when someone tries to answer a simple question, "is this getting better or worse?" and discovers there's no consistent record of what "this" was ever called.

We named this problem in our article on Dabbling vs. Doing as the thing that "breaks first" in DIY AI analysis. It's worth a piece of its own, because it's also the hardest failure mode to notice while it's happening, and the most expensive one to unwind once your leadership team has already made a decision based on numbers that were never actually comparable.

The market is finally naming this out loud

Enterprise AI adoption research is starting to catch up with what taxonomy-drift victims already know from experience. NVIDIA's 2026 State of AI Report found that having sufficient, usable data, not model capability, not compute cost, is now the single most commonly cited top challenge to scaling AI, named by 48% of respondents (NVIDIA, 2026).

IDC's FutureScape 2026 predictions put a number on what that failure actually costs, companies that don't prioritise high-quality, AI-ready data will face a 15% productivity loss by 2027 as they try to scale GenAI and agentic solutions (IDC FutureScape 2026).

Read those two numbers together and the message is blunt, the model was never the bottleneck. The mess underneath it was. Taxonomy drift is what that mess looks like in practice, one prompt at a time.

What it looks like from the inside

We didn't have to speculate about this. It shows up constantly in Four/Four's own first-party conversation data, across customers who are actively trying to make DIY AI tagging work.

AI tagging without shared context invents its own taxonomy. One customer's team described exactly this failure mode in their own automated pipeline: "Current AI-based automation for tagging and categorizing customer feedback lacks the context needed for accurate grouping, resulting in the creation of duplicate or incorrect insights," flagging a direct need for "improvement in AI context recognition". This is the mechanical root of taxonomy drift, that without a shared, enforced structure, every classification pass is a fresh guess, and "fresh guesses" don't stay consistent with each other.

Teams know they need structure, and can't get it without deliberate design. Another account told us directly that "there is confusion regarding the use of labels and other organisational attributes for structuring insights, slowing down knowledge management," and asked plainly for help "understanding how to use labels to provide more structure and context to insights... to group and analyse more efficiently". This is a sophisticated, engaged customer, and the taxonomy problem is still hard enough that they came looking for a structured system to solve it, rather than trying to hand-roll one.

Ungoverned AI adoption doesn't stay contained to one team. One customer flagged a strategic-level version of the same risk, "adoption of in-house AI tools without governance and systems integration leads to massive inconsistencies, which poses a huge risk for organisations that invest heavily in consistent brand perception" raised in the same conversation where they described "inconsistent narrative and lack of structured prompts in proposal generation... leads to poor-quality, inconsistent sales proposals that fail to effectively communicate value to clients". Taxonomy drift in customer insight is the quiet version of this. Inconsistent proposals are the loud version. Same root cause though, no shared structure governing how AI-assisted output gets produced.

Fragmentation compounds across channels, not just across prompts. A fourth account described the aggregate effect: "the challenge of gathering sentiment from a variety of internal and external channels, where information is siloed and inconsistent, makes it difficult to establish a comprehensive view of customer and team feedback, impeding decision-making". Every additional channel, Slack, support tickets, calls and email is another place for the same underlying pain point to get relabelled from scratch.

And this is the sharpest version of the pattern, already cited in our anchor article, one customer described the moment two commercial, purpose-built analysis tools disagreed with each other on account health, producing "confusion and discrepancies between analysis... leading to conflicting insights for their team about major account statuses and sentiment". If two structured platforms, each with their own internal taxonomy discipline, can still diverge enough to confuse a customer-success team about whether an account is healthy, a hand-written Claude prompt, changing slightly every time someone uses it, doesn't stand a chance of holding a line over a full quarter.

Why "just be more consistent" doesn't fix it

The instinctive fix, write a really good prompt, save it, tell everyone to use the same one, buys you a few weeks, maybe a quarter. Then someone tweaks it for their use case. Someone else joins the team and writes their own version because they didn't know the shared one existed. The underlying model updates and starts interpreting the same prompt slightly differently. None of this is a discipline failure on your team's part. It's what happens when consistency depends on a document nobody's incentivised to maintain, rather than a system that enforces the taxonomy structurally, validating new labels against existing ones, flagging near-duplicates before they're created, and treating "what do we call this" as infrastructure rather than as a Slack message pinned to a channel.

Three questions that are specific to this failure mode

  1. Merge testing — If you fed last month's tagged insights back through your current process, would the labels come out the same, or would half of them land somewhere new?
  2. Cross-tool agreement — If two people, or two tools, independently analysed the same batch of customer conversations, would they agree on what the top three issues were?
  3. Silent renaming — Can anyone on your team tell you, right now, every label that's ever been used for "the same" underlying pain point, or would that require someone's memory rather than a system?

If any of those made you pause, that's not a failure of effort. It's the default state of taxonomy under DIY AI analysis, and it's exactly why "dabbling" and "doing" diverge fastest here. Doing means the taxonomy is the product, held constant by design, not by whoever remembered the house style this week.

That's the layer we built Four/Four to own, so the AI agent doing the reasoning, whether it's Claude, GPT, or whatever comes next, is always reasoning over a taxonomy that's actually been held steady.

We use cookies as specified in our Privacy Policy. You agree to consent to the use of these technologies by clicking Allow Cookies.