8 questions to ask before you call your team's Claude project "done"

Chris Lloyd 2026-09-03
#marketing #sales #executive #customer-success #product

Somewhere in your organisation, a prototype built on Claude or ChatGPT is quietly being treated as a finished product. It works. Everyone who's seen the demo is impressed. Nobody has asked the eight questions below, which is exactly how a 30-minute win turns into an unowned, unscalable liability six months from now.

Across enterprise AI pilots broadly, roughly 95% never reach measurable P&L impact, not because they failed outright, but because nobody stress-tested them against questions like these before the volume, the compliance requirements, and the org-chart reality arrived.

Score your qualitative research project honestly against each item below. This isn't a judgement of the model, Claude and its peers are exceptionally capable at the reasoning step. It's a judgement of everything you built, or didn't build, around it.

How to score: for each item, give your project 0 (not addressed), 1 (partially addressed / manual workaround), or 2 (fully addressed, systematised). Add up your total out of 16 at the end.

1. Taxonomy consistency

Question: If the same pain point or theme shows up in fifty different conversations, does it land as one comparable signal or fifty slightly different ones, phrased however that day's prompt happened to phrase them?

Dabbling: Every analyst, and every prompt session, produces its own labels. Nobody can query "how often has X come up this quarter" and trust the answer. Doing: A structured taxonomy is enforced at the point of extraction, not cleaned up afterward in a spreadsheet.

Score: _ / 2

2. PII and access governance

Question: Is sensitive data redacted and access-controlled by policy, automatically, on every conversation, or does it depend on whoever happened to remember to check before hitting send?

Dabbling: Governance shows up as an incident, not a design decision. One Four/Four customer described a "recent acute challenge related to the access and management of sensitive data, specifically around PII and GDPR requirements, which led to disruption in team workflows." Doing: Redaction and role-based access are infrastructure, applied before anyone can query the data, not a policy document nobody enforces.

Score: _ / 2

3. Named ownership

Question: Is there a specific person whose job description includes maintaining this pipeline, with time formally allocated to it, or did it inherit an owner by accident when the person who built it moved on?

Dabbling: The build lives in one engineer's side-project time. When priorities shift, so does the tool's future. Doing: Ownership, escalation path, and maintenance time are named and reviewed, the same way you'd staff any other production system.

Score: _ / 2

4. CRM and system-of-record resilience

Question: When an account gets renamed, merged, or re-owned in your CRM, does the sync recover automatically, or does someone notice three weeks later that the numbers look wrong?

Dabbling: Scripts work perfectly in testing and degrade the moment real-world account churn hits, with no error recovery built in. Doing: Sync failures are monitored, alerted, and recoverable without a manual rebuild.

Score: _ / 2

5. Historical trend, not just snapshots

Question: Can you show, with confidence, how sentiment or volume on a specific topic has moved over the last two quarters, or does every analysis start from zero because last quarter's taxonomy doesn't map to this quarter's?

Dabbling: Every report is a one-off. Comparing quarter to quarter means re-deriving the labels by hand. Doing: A stable taxonomy means trend lines are a query, not a project.

Score: _ / 2

6. Volume math, done honestly

Question: Have you actually multiplied the effort of your best single output by the number of conversations your business generates in a month, or does the plan quietly assume the 30-minute demo scales linearly?

Dabbling: The unit economics were never run. One good PRD from one good call took eight rounds of prompt refinement and a subject-matter expert in the loop the whole time. Nobody has asked what that costs at ten thousand calls. Doing: The cost of codification at scale, engineering time, review time, drift correction, is modelled and owned before it's promised to a stakeholder.

Score: _ / 2

7. Tool-calling and integration reliability

Question: When the model needs to reach out to another system, a CRM lookup, a web search, a downstream agent, is that call-out designed and tested for failure modes, or is it a lucky prompt that happened to work in the demo?

Dabbling: Tool-calling was never load-tested against malformed inputs, timeouts, or partial failures. Doing: Integrations are engineered with the same rigour as any other production API dependency.

Score: _ / 2

8. Executive visibility

Question: If a board member or exec asked "what's the ROI of the AI work happening in this team," could you answer in a sentence with a number, or would you need three days to reconstruct what's actually running and why?

Dabbling: Nobody schedules a retro for a project that just stops getting mentioned in stand-up. It doesn't fail loudly; it fades. Doing: The system reports on itself (usage, accuracy, coverage) as a matter of course, not a fire drill before a board meeting.

Score: _ / 2

Your total: ___ / 16

0–6 — You're dabbling. Capably, even impressively, but dabbling. This is the default state for almost every team experimenting with frontier models right now. The risk isn't the experiment; it's mistaking it for a finished system.

7–11 — You're transitioning. Some of the hard, unglamorous work has started. This is the highest-risk zone: far enough along to feel done, not far enough along to survive its ten-thousandth conversation without someone finding out the hard way.

12–16 — You're doing. Codification, governance, and ownership are treated as infrastructure, not afterthoughts, which means your Claude or ChatGPT agents are finally free to do what they're actually good at: reasoning over signal that's already trustworthy

We use cookies as specified in our Privacy Policy. You agree to consent to the use of these technologies by clicking Allow Cookies.