What a product audit actually finds
After two dozen audits, the finding is almost never the tech. It is that nobody agreed what done means, or the data was never there.
Every audit I run starts the same way. Someone reaches out, usually a founder or a VP, and says their AI feature isn’t landing the way they hoped, and they want someone from outside to tell them why. I’ve done maybe two dozen of these now, structured teardowns of someone else’s build. I used to think most of them would turn up bad models or wrong architecture choices. They mostly don’t. After enough audits, a pattern shows up that surprised me the first three or four times and doesn’t surprise me anymore: the tech is almost never the finding.
I’m not saying this is universal, I don’t have a hundred audits to prove it, but it’s consistent enough across two dozen of them that I’ve stopped expecting anything else. Sometimes the model choice is genuinely bad, sure. But the actual diagnosis, the thing that’s been quietly stalling a team for months, is usually something nobody wanted to say out loud in a standup.
Nobody agreed what done means
The most common finding, by a wide margin, is that three or four people on the same team are building toward three or four different finish lines. The engineer thinks done means the API returns a response without erroring. The PM has a completely different bar, usually whether the demo survives the exec review next Tuesday. And somewhere in support, done quietly means customers stop filing tickets about it, a bar nobody upstream even knows exists. None of these people are wrong exactly, they’re just aiming at different things, and nobody in the room forced a single answer. So the team ships, everyone privately grades it against their own version of finished, and half of them walk away disappointed while the other half thinks it went fine. That gap never shows up in a sprint retro. It shows up three months later as “why does this still feel unfinished.”
The data isn’t where anyone thinks it is
Second most common finding, and this one is a bit more painful because it usually means real rework: the product needs data the company assumed it already had, and it doesn’t. Not in the shape the model needs, not labeled correctly, sometimes not collected at all. I’ve watched a team spend six weeks tuning prompts to fix a quality problem that was actually a missing dataset, because tuning a prompt feels like progress and admitting the data doesn’t exist feels like admitting the last two quarters were built on sand. Basically nobody wants to be the one who says it out loud in the room. So the prompt gets tuned again, and the real problem sits there untouched.
Three teams, one hard part, no owner
The third pattern took me the longest to name properly. A genuinely hard piece of the problem, evaluation, or edge-case handling, or what the product does when the model is wrong, sits exactly in the overlap between two or three teams. Each one assumes it’s someone else’s job, because on paper it kind of is everyone’s job, which in practice means it belongs to no one. I found this exact gap once on a project where the AI layer, the backend, and the product team had each quietly built their half assuming the other half existed somewhere else. It didn’t exist anywhere. Nobody had lied about it, everyone had just assumed.
What the audit actually hands back
A fixed-price audit isn’t a deck full of hedged language and options to consider. It’s a written diagnosis of what’s actually broken, in plain sentences, followed by a ranked list of what to fix first and why that one goes first. No “it depends,” no “let’s align further.” I’ve had clients tell me it was the most direct feedback they’d gotten on their own product in over a year. That says something about how rarely anyone tells founders the truth to their face.
The tech being fine is, honestly, the less useful thing to find. It means the fix is people agreeing on something, which takes longer and feels far less satisfying than shipping a patch, and nobody puts it on a roadmap because it doesn’t look like a task. It just looks like a conversation somebody should have had eight months ago and kept putting off.
Have a product like this to ship?
Book a call →