Your data mess is the bottleneck for AI greatness, not the model itself.

Rob Dumbleton 2026-08-26
#product #executive #sales #customer-success #marketing

One of the joys and pains of being a commercial founder is that you have to do a lot of the grunt work when it comes to speaking to the market and identifying potential opportunities.

It’s a joy because you get to speak to real people about real problems and really get a grip on what’s important. This gives me a buzz, the thrill of the connect and the possibility of booking a meeting with someone we can help.

It’s also a pain, as when you’re shattered after a long day’s work and you’ve got the kids down to bed, you have to summon the enthusiasm to do US cold calls at 8pm. Smile and dial baby!

One of the recurring themes I’ve identified on these calls is the need to link CRM (structured) data to conversations (unstructured). This is a delight for me to hear, as when someone mentions this, I know they know their underlying technology when it comes to AI, ML and LLMs.

What irks me though, is when I speak to someone who’s clearly in the tech field, who works in product and who has ‘AI’ slapped all over their LinkedIn profile and they say, “we just bung everything into Claude and it seems to work just fine”.

Yea it might be a fob, ‘get this bloke off the phone’, but the more and more I think about it and see what people are still NOT doing with their AI capability, the more and more I’m disappointed with that answer.

Now, I don’t know if anyone reads these newsletters I write, I pore over them for days, even weeks in terms of what to talk about. Then I think write about what you know and spread what you’ve learned, so here we go.

Having been in the tech space for 20 years now, the one thing I’ve seen that has never left, and in fact is even more important today with AI, is data.

Back in 2010 when cloud got huge, ‘Big Data’ was all the rage. Loads of companies glossed over it as it’s hard to do and so kept their siloed systems and data sources. Then companies like Databricks and Snowflake came out, that predate the LLM era. They built their businesses on data warehousing and analytics for BI and traditional ML, but in doing so have become central to LLM workflows because of the problem that turns out to be harder than it looks: getting an organization's data into a state where a model can actually use it well.

Loads of companies have and continue to invest in their data architecture, with vendors like Snowflake showing no sign of slowing down.

Most enterprise data isn't sitting around in a clean, unified format. It's scattered across transactional databases, event logs, CRM exports, PDFs, support tickets, Slack threads, and dozens of SaaS tools, often duplicated, inconsistently labelled, and full of stale or conflicting versions of the same fact.

So for an LLM application, whether that's fine-tuning a model, running retrieval-augmented generation (RAG), or just letting an agent query internal data, that mess is the actual bottleneck, not the model itself. You need something to ingest all of it, deduplicate and clean it, track lineage and provenance, enforce access controls (so the model doesn't surface data a given user shouldn't see), and turn unstructured text into something query-able.

That data never has to leave its governed environment to be used by a model. In the case of systems like Snowflake, Databricks and (enter) Four/Four the pitch is the same, whoever controls the trusted, governed copy of an organisation's data controls how well any LLM built on top of it performs. If that data has to be exported to some separate AI stack, the platform loses relevance. Structuring data for LLMs" isn't really a new discipline, it's the same governance, quality, and access-control problem these platforms already solved, now with embeddings and retrieval bolted on top.

Having Claude or OpenAI or Copilot or another of the frontier models in your business is just table stakes. What you connect it to and how that data is organised is where you can really set yourselves apart from the competition. Turning “I think it works okay” to “I can prove it’s right”.

I’ll keep fighting the good fight on this one and hopefully not too much of a rant. Maybe I am a bit of a broken record, but I just really want people to understand what it takes to get ahead, stay there and use the superpower that they already have but are not using properly: first-party data, that in a lot of cases has probably cost millions to obtain in the first place.

We use cookies as specified in our Privacy Policy. You agree to consent to the use of these technologies by clicking Allow Cookies.