What Happens After Your Team's Claude or OpenAI Prototype Works?

Four/Four 2026-08-17
#sales #executive #customer-success #product #marketing

Somewhere in your organisation right now, someone has built something impressive with Claude or ChatGPT. They pasted a customer call transcript into a chat window, wrote a decent prompt, and got back something that looked like a product requirements document. It took thirty minutes. Everyone in the room was rightly excited.

Here's the uncomfortable statistic that excitement tends to skip past, across enterprise AI pilots broadly, roughly 95% never reach measurable P&L impact. Not "were disappointing", they never got far enough to be measured against the business at all. Most AI initiatives right now aren't failing. They're stalling, quietly, somewhere between the demo and the deliverable, and nobody schedules a retro for a project that just... stops getting mentioned in stand-up.

We wrote this because we think the reason for that stall is specific, mechanical, and almost entirely avoidable, provided you're honest about what actually happened during that thirty-minute demo, and what it would take to do it again, correctly, ten thousand times.

What the demo actually proved

Building one good PRD from one good customer call is genuinely achievable with a frontier model, and it's worth being precise about what it took, because the honest version is more instructive than the impressive version. Turning a single messy transcript into a clean set of voice-of-customer statements took somewhere north of eight distinct rounds of prompt refinement. Getting the model to reliably call out to other systems, a CRM lookup here, a web search there, took deliberate tool-calling design, not a lucky prompt. And the output was only trustworthy because a person who understood the account sat in the loop the whole time, correcting drift as it happened.

That's not a criticism of the model. It's a description of the actual unit of work: raw, noisy conversation on one end, a clean, structured, comparable signal on the other, and a substantial amount of disciplined engineering in between, lets call it codification. Skip that middle step, and what you get instead is exactly what most DIY builds produce, inconsistent, one-off outputs that can't be compared across time, across accounts, or across the person who happened to write today's prompt.

Thirty minutes of expert, hands-on attention, for one call. Now multiply by the number of customer conversations your business actually has.

Where the DIY build genuinely wins and where it quietly stops

To be fair to the DIY instinct, building on Claude directly has real strengths, and we'd be dishonest to pretend otherwise. It's fast, days, not months, to get something in front of a stakeholder. Your team keeps full control over prompts and outputs, tuned to exactly your business's language. And there's no new vendor relationship to negotiate, no procurement cycle, no seat licences to justify.

Those are real advantages, and they're precisely why the pattern is so common. The trouble is where it goes next, and we don't have to speculate about this part, it shows up constantly in our own customers' words, including from teams already running Four/Four alongside their own Claude experiments.

Taxonomy breaks first. The same pain point, mentioned by forty different customers, becomes forty different insights, because forty different people (or forty different prompt sessions) each described it slightly differently, with no shared structure holding it together. This isn't hypothetical for teams comparing tools either, one Four/Four customer described exactly this failure mode when two analysis tools disagreed: "confusion and discrepancies between analysis done by Four/Four and another tool... leading to conflicting insights for their team about major account statuses and sentiment," undermining their ability to determine true customer health. If that can happen between two structured platforms, a hand-written prompt has no real chance of holding a consistent taxonomy across a quarter, let alone a year.

Governance arrives as an emergency, not a design decision. Across the market, roughly 53% of organisations that deployed AI agents without governance built in from day one experienced a permission breach. That statistic has a face inside our own customer base: one account told us about "a recent acute challenge related to the access and management of sensitive data, specifically around PII and GDPR requirements, which led to disruption in team workflows until access was restored", and another flagged that, as configured, "anyone can access and listen verbatim to any call, regardless of department or involvement," a live exposure risk nobody had signed off on. Nobody designs a system to fail this way. It fails this way because PII redaction, audit trails, and access control are unglamorous, and unglamorous work is exactly what gets skipped when the goal is "get a demo working by Friday."

Sync breaks quietly. Scripts that pull from your CRM work perfectly in testing and then degrade the moment an account gets renamed, merged, or re-owned, with no error recovery, because nobody built error recovery into a prototype. One customer described the practical version of this: a "bottleneck in resolving CRM sync errors due to inconsistent response rates from the team managing CRM data, resulting in delays that impact workflow and system reliability." Multiply that friction across every account change your CRM sees in a quarter.

Ownership becomes the real cost. The engineer who built the internal tool over a sprint two quarters ago has since moved teams, or is now fully consumed by a roadmap item that actually has their name on a scorecard. The tool doesn't get maintained because nobody owns "maintaining the internal AI thing" as a job. It becomes what every DIY build eventually becomes: an orphaned side project, competing for attention against work that someone is actually accountable for.

None of this is a failure of the model. Claude, GPT and their peers are extraordinarily capable at the actual analysis step. What they don't do, what no model does on its own, is remember your taxonomy, redact your PII by policy rather than by luck, recover from a broken CRM sync, or show up to a monthly review to explain why nobody's touched the pipeline since March.

It gets sharper, not softer, in voice

If this sounds like a text-and-transcripts problem, the same pattern shows up with a shorter fuse the moment voice enters the picture, and for teams doing customer research over calls rather than chat, that's most of you. Enterprise build-vs-buy analysis puts roughly 80% of the total cost of an in-house voice AI build after the first call goes live, not before. Latency orchestration alone, which is keeping speech-to-text, the model, and text-to-speech coordinated within a few hundred milliseconds, takes specialised expertise most product teams simply don't carry. "Model drift," as underlying models update, forces continual code changes just to keep pace. Add it up over three years, and in-house voice builds tend to run 40–60% more expensive than buying, once engineering time, retraining, and infrastructure are honestly counted. As one analysis puts it bluntly: "a functioning prototype is a dangerous metric for success." The prototype tells you the demo works. It tells you almost nothing about whether the system survives its hundredth call, let alone its ten-thousandth.

Even sophisticated teams want someone else's curation layer

Here's the detail that should reframe this whole conversation, because it didn't come from a market report, it came directly from our own customers. The most sophisticated teams in our book of business aren't choosing between Four/Four and Claude. They're asking for both, in a specific order, Four/Four doing the curation, Claude (or another agent) working on top of it. One customer put in a direct request "to use Four/Four as a headless data curation layer, enabling integration with external AI agents such as Claude for downstream processing," so that other applications could "plug in and work through curated data efficiently."

Read that carefully, because it's the whole argument in one sentence. The teams closest to this problem, running real budgets and real headcount against it, have independently concluded that the agent is the easy, replaceable part, and the governed, deduplicated, consistently-taxonomised signal underneath it is the part worth protecting, standing up once, and not rebuilding every time a new model ships.

Dabbling vs. doing: a five-question test

You don't need our diagnosis to know which side of this line you're on. Ask your team five questions about the internal AI tool that's currently analysing customer conversations:

  1. Taxonomy — If the same pain point appears in fifty different calls, does it land as one comparable insight, or fifty slightly different ones?
  2. Governance — Is PII redacted by policy, enforced automatically, or by whoever remembered to check before hitting send?
  3. Ownership — Is there a named person whose job includes maintaining this pipeline, with time allocated to it, or did it inherit an owner by accident?
  4. Resilience — When your CRM renames an account or merges two records, does the sync recover on its own, or does someone notice three weeks later that the numbers look wrong?
  5. Trend — Can you show, with confidence, how sentiment on a specific topic has moved over the last two quarters, or does every analysis start from zero?

If the honest answer to more than one of those is "we don't really know," you're dabbling. That's not a judgement, it's the default state for almost every team experimenting with frontier models right now, which is exactly why it's worth naming plainly rather than discovering it during a board question about AI ROI.

Doing looks different. A system that turns messy conversations into clean, comparable signal automatically, holds that taxonomy steady as volume grows, treats governance as infrastructure rather than an incident response, and has someone whose job is to keep it that way leaving your team, and your Claude or ChatGPT agents, to do what they're actually good at, reasoning over signal that's already trustworthy.

That's the system we've built. If you want to see what your own conversations look like once they've been through it, including the ones currently living in a Claude chat window we'll show you, using your actual data.

We use cookies as specified in our Privacy Policy. You agree to consent to the use of these technologies by clicking Allow Cookies.