← Field notes

How to scope an AI project so it does not balloon

Scope from the one metric that has to move and the smallest slice that proves it, then do the data reality check before anything else.

Nearly every AI project I get pulled into opens with the wrong thing. Someone is excited about the model, not the metric. They want to show me the demo, the retrieval setup, the chain of prompts that does something clever. Nobody opens with the number they’re actually trying to move. That’s backwards, and it’s the first thing I fix before a single line of code gets written.

Scope from one metric that has to change, not a vision doc full of adjectives. Signups that convert, or support tickets that close without a human touching them, something you can point at and say yes or no to in a few weeks. Then find the smallest slice that proves the number can move at all. Give it two weeks, not two quarters, and see what actually happens.

I learned this the expensive way, long before AI was even a thing people scoped projects around. Years ago I built a lamp that charged your phone wirelessly with a whole ambient lighting setup built into it. I also built a smart mirror that showed you the weather and your calendar while you brushed your teeth. Both were genuinely well built. I couldn’t sell a single unit of either one. I had scoped both from what was fun to build, not from what anyone would actually pay for. Nobody was walking into a shop asking for a mirror that read out their meetings. That gap between cool to build and someone will pay for this is where most first versions die, AI or not.

The data reality check nobody wants to do first

Before anyone touches a model, I ask what data actually exists to support the thing they’re imagining. Most AI projects don’t stall because the model got something wrong. Actually, that’s not quite true, they do stall on the model sometimes, but nine times out of ten the real problem is data that was supposed to be clean and ready and turns out to be scattered across three systems, half of it stale, with nobody owning the pipeline.

This is the least exciting part of scoping and also the part that decides whether a project ships in six weeks or six months. I run it before writing a single line of spec. Pull ten real examples of what the model will actually see in production, not the clean ones from the pitch deck. If those examples are messy or incomplete, sometimes they don’t even exist yet, that’s the actual project. The model comes after.

What a scoped-down version actually looks like

For Astrika this came down to one question. Would the reading even be believable if the chart underneath it was wrong. So the deterministic engine, the part that computes the actual planetary positions, got built and tested before a single word of the AI’s writing style got touched. Get the ground truth right first, then let the model do the part it’s actually good at, which is language, not arithmetic.

A properly scoped AI project looks a bit disappointing at the start. Just one metric to move and the smallest possible slice to test it, nothing more finished than it needs to be. Founders want the full vision in week one, I get it, I used to want that too when I was building lamps nobody asked for. But the version that ships in three weeks and tells you something true beats the version that looks impressive in a deck and takes four months to discover it was scoped wrong from day one.

I still think about that smart mirror sometimes. Really solid engineering. Nobody wanted it.

Have a product like this to ship?

Book a call