AI-assisted teams move fastest in their first few weeks and get progressively more expensive after that — not because the AI writes bad code, but because nobody's comparing the hundred small, independent decisions it's making in parallel. Here's the checkpoint system we use to catch that early, before it costs a quarter of a sprint instead of a five-minute conversation.
We've watched this pattern play out on more than one client engagement now, and it always has the same shape. The first month of building with AI generation — describing what you want in plain language and letting AI produce the prototype or the code, sometimes called "vibe coding" — looks like nothing but upside. By month three, a real chunk of that speed is gone, spent reconciling decisions that were each perfectly reasonable on their own and never got compared against each other.
This isn't an argument against AI-assisted building. It's an argument for knowing exactly where the bill comes due, and putting a cheap checkpoint in place before it does.
Why the first month looks like pure upside
The early advantage of AI-assisted building is real, not hype. When one person can describe a screen and get a working version back in minutes instead of days, a few things happen at once: people stop waiting on each other, ideas get tested before anyone has to defend them in a meeting, and the sheer volume of what a small team can attempt goes up.
That's genuinely valuable, and none of what follows is an argument against using AI generation early and often. The problem isn't the speed. It's what the speed hides.
Every AI-assisted session — one designer exploring a checkout flow on Monday, an engineer generating an admin panel on Tuesday, a PM sketching a settings page on Wednesday — makes a series of small, locally reasonable decisions. How to format a date. Which component to reach for when two similar ones exist. How to handle an empty state. Each decision is defensible in isolation. None of them get compared against each other, because there's no natural moment where that comparison happens. The team is moving too fast, and everything demos fine.
What actually accumulates
This is the part we see constantly on data-heavy products specifically — dashboards, internal tools, anything with more than one way to display the same underlying data.
Once a product has enough surface area, the small, independent decisions from different sessions start to collide. Three flows handle the same data type three different ways. Two components that look similar behave differently under edge cases nobody tested for, because each was generated against a different prompt with a different implicit assumption. A state that got handled cleanly in one AI session never got asked about in the other four sessions touching the same object.
None of this shows up as a bug report at first. It shows up as a growing sense that the product feels slightly inconsistent, followed — usually a couple of months in — by a genuine integration crisis: QA starts finding contradictions, engineers start tracing weird behavior back to three different code paths doing the same job differently, and the team ends up spending a real chunk of its sprint capacity untangling decisions that were each five minutes of AI generation time when they were made.
A recent body of qualitative HCI research — interviews with 22 people across product teams in enterprises, startups, and academic labs — put a number on this: by week twelve, the teams studied were spending 20–30% of their sprint capacity fixing bugs that traced back to AI-generated code, after early weeks that looked like nothing but a speed unlock. It's a qualitative study, 22 interviews rather than a large-scale survey, so we're treating the exact percentage as directional rather than gospel — but the shape of the curve it describes matches almost exactly what we've watched happen on client work.
That's the actual cost. Not that AI wrote bad code. That nobody reconciled a hundred small, independent, individually reasonable decisions before they compounded.
We've seen this before AI, too — it just moved slower
This isn't a new failure mode. AI just runs it faster, across more surface area, with fewer natural pauses to catch it.
On Cohere Commerce, a B2B retail intelligence platform where we've been the embedded design team for two years, we inherited a product with two competing component libraries — not because anyone decided to build two, but because enough individually reasonable decisions, made by different contributors at different times, never got compared against each other. By the time we started, "which button style is correct" didn't have an obvious answer, and every new feature was adding its own version of the same inconsistency instead of resolving it. It took a full design-system rebuild to unwind.
On an equipment-inspection app we redesigned for Arnlea, used by inspectors on oil rigs and power plants, we found the same piece of equipment being logged five different ways — purely from unconstrained free-text entry, no AI involved at all. Structured inputs replaced the free text, and the inconsistency mostly disappeared.
On Mintra's exam and training platform for maritime and energy crews, we took the opposite approach from the start: before drawing a single screen for the question bank, we mapped every state, every rule, and every transition a question could go through. That mapping is exactly the discipline this article is arguing for. It's just usually skipped, because it doesn't feel like "real" design work — until the alternative makes the cost obvious a few months later.
The pattern is the same in all three cases: unconstrained, parallel, individually-reasonable decisions accumulate into inconsistency when nobody's job is to compare them. AI didn't invent this. It just compresses a two-year drift into a ten-week one.
A checkpoint system that catches it in week three, not week twelve
The fix isn't slowing down generation, and it isn't adding a heavyweight process that kills the speed advantage that made AI-assisted building worth adopting in the first place. It's a small number of specific checkpoints, placed early:
None of these require slowing the team down. They require someone whose job is specifically to look across sessions instead of inside one session at a time — which, by definition, isn't a role AI can fill for itself. It doesn't know what the other four sessions decided.
The actual cost calculus
Here's the case we'd make to any founder or product lead deciding whether this is worth setting up now versus dealing with it later: catching a divergent decision in week three costs a five-minute conversation and a one-line update to a shared doc. Catching the same divergence in week twelve costs a multi-day untangling effort, a chunk of a sprint that was supposed to go toward new work, and — on a data-heavy product with a lot of surface area — a genuine risk that the inconsistency reached users before anyone caught it.
Teams don't choose the expensive version on purpose. They choose it by not having anyone whose job is to look for it early, because for the first month, there's nothing to look for.
If your team has been building with AI generation for a few months now, this is worth an honest hour of stocktaking: has anyone actually measured how much time is going into fixing divergence between AI-assisted sessions, or does it just feel like normal sprint work by now? That question is usually the fastest way to find out which side of the curve you're already on. It's also the kind of audit we run for teams building data-heavy products who'd rather catch this in week three than week twelve.
Gytis Markevičius
Founder, GytisMark Studio — gytismark.com
Other articles
Every engagement includes a senior designer thinking about the product strategically, catching issues before they reach you, and making sure the work holds up under scrutiny.
Let's talk about your product and your business goals.
Request your FREE trial →✦ Start with a conversation ✦
20 minutes. No pitch. We'll tell you honestly if we're the right fit.