Leet Force
Back to Blog
StrategyUpdated

Your AI-Built Prototype Works. That's the Expensive Part.

AI gives 35-40% on greenfield work and about 10% on complex legacy code. Shipping your prototype moves it from the first category to the second, on launch day. Here's what that costs and how to decide what to keep.

Abdullah Khan
By Abdullah KhanCEO, Leet Force
Your AI-Built Prototype Works. That's the Expensive Part.

The conversation happens most weeks now, and it always opens the same way: someone built a working version of the product in a fortnight, the board has seen it, and now there's a launch date. The question we get asked is how long it takes to "finish" it.

The honest answer usually surprises people, so it's worth explaining where it comes from rather than just quoting a number. Because the prototype working is not evidence that the expensive part is behind you. In most cases it's the reason nobody can see the expensive part yet.

This isn't an argument against building with AI. We use these tools daily and they've genuinely changed what a small team can do. It's an argument about what you actually have when the demo runs, and what it costs to turn that into something you can operate.

Why the prototype was cheap

Start with the finding that explains nearly everything else. Research from Stanford's Software Engineering Productivity programme, cited in DORA's 2026 ROI of AI-assisted Software Development report, found the productivity gain from AI varies enormously by the kind of work:

  • 35–40% gains on simple, greenfield tasks
  • Roughly 10% on complex legacy code

A prototype is the purest greenfield task there is. No existing architecture to respect, no established patterns, no users whose data you can't break, no edge cases discovered in production, no tests to keep passing. It is precisely the shape of work where these tools are at their best — which is why it took a fortnight, and why it genuinely was a bargain.

Now consider what happens the moment you put it in front of customers. It acquires users, data you can't lose, uptime expectations, a support burden, and a growing set of behaviours people depend on. Your greenfield project has become complex legacy code. Not in three years — on launch day.

The productivity that built it does not transfer to maintaining it. That's the whole problem in one sentence, and it's why "we built the hard part in two weeks, the rest should be quick" is almost exactly backwards.

What's structurally different about the code

The next question is fair: if it works, does it matter how it was written?

It matters for one specific thing — what it costs to change. And there's now measurement rather than opinion. GitClear has been analysing code changes across a large corpus (211 million changed lines for the 2020–2024 window) and tracking structural signals as AI authorship scaled:

  • Refactoring collapsed. Moved code — their proxy for consolidation and reuse — fell from about 25% of changed lines in 2021 to under 10% in 2024.
  • Copy-paste overtook it. Cloned lines rose from 8.3% to 12.3% over the same period. 2024 was the first year on record where copy-paste exceeded moved code.
  • Duplication reached a record. Block duplication per million changed lines went from 40.3 in 2023 to 73.0 year-to-date in 2026 — an 81% increase, and the highest level in their dataset.

The mechanism is unremarkable once stated. It's easier for a model to generate a fresh block that works than to find the existing function that nearly does the job and generalise it — partly a context limitation, partly that nothing in the workflow rewards consolidation. GitClear's framing is that today's default AI workflow is incentivised to deliver "a happy-path, a passing test, a closed ticket" while taxing "the reuse, consolidation, and error-surfacing that determine how expensive a codebase is to own in year three."

That is the cost you're buying. Duplication is not an aesthetic complaint. It's a multiplier on every future change: the pricing rule that exists in six places gets updated in five, and the sixth becomes a bug that surfaces at quarter end. Nothing about that is visible in a demo. It's visible in your third change request.

Why your team's estimate is probably wrong, in a known direction

This is the part worth understanding before you set the launch date, because it affects the number you're being given.

METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real tasks in their own mature repositories — projects averaging over 22,000 stars and a million lines of code. Tasks were randomly assigned to allow or forbid AI tooling.

The developers expected AI to speed them up by 24%. Afterwards, they believed it had sped them up by 20%. Measured, they were 19% slower with AI than without it.

Note carefully what that is and isn't. It's a 39-point gap between perceived and actual performance, in experienced engineers working on code they knew well. It is not evidence that AI slows developers down generally — METR say so explicitly, and it's worth repeating because the study gets misused in both directions. They state their results do not show that AI fails to speed up most developers, and flag real limitations: a small sample, likely participant selection effects, and developers with only around 50 hours of experience with the tool.

What survives the caveats is the calibration problem, and it's corroborated elsewhere. Stack Overflow's 2025 Developer Survey found the top frustration, cited by 66% of developers, is AI output that is "almost right, but not quite" — and the second, at 45%, is that debugging AI-generated code takes longer. Trust in AI accuracy fell to 29%, down from 40%, even as adoption climbed.

So when your team says the prototype is 80% done, they are being sincere. The research says that estimate is unusually likely to be optimistic, because the mode of work that produced the first 80% is not the mode of work that produces the rest. "Almost right" code demos beautifully and finishes slowly.

The J-curve, and the dip you should budget for

DORA's 2026 report gives this a shape worth knowing: a J-curve of value realisation. Organisations adopting AI typically see productivity dip before it rises. They call it the tuition cost of transformation, and attribute the dip to three things:

  1. The learning curve — teams adapting workflows to the tools.
  2. The verification tax — the cost of reviewing AI-generated code, which does not go away and scales with volume.
  3. Process adaptation — testing, review and change approval straining under more code arriving faster.

Their illustrative model also prices something they call the instability tax: a change failure rate rising from 5% to 6% after adoption, costing $344,000 in downtime for the organisation they model. Small percentage, real money.

This echoes DORA's 2025 finding that AI functions as an amplifier: strong teams with good tests, clean deployment and small batches get faster; weak ones ship low-quality work more quickly. If your prototype has no tests and no deployment pipeline, adding AI velocity to it does not produce a product. It produces more of what you have, sooner.

What to check before you decide anything

You can assess this in about two days, and you should, because the keep-or-rewrite decision is worth far more than the assessment costs. Six things, roughly in order of how often they're the problem:

  1. Are there tests, and do they test behaviour? AI writes tests readily, and they frequently assert that the code does what it does rather than what the business requires. A suite that passes against a bug is worse than no suite, because it manufactures confidence.
  2. How many times does each business rule appear? Pick your most important rule — pricing, permissions, tax, eligibility — and count the implementations. One is healthy. Three or more and you have GitClear's finding in your own repository, and you should price changes accordingly.
  3. What happens on the unhappy path? Prototypes are built along the happy path because that's what gets demoed. Turn off the network mid-transaction. Submit the form twice. Feed it a duplicate order. The gap between demo and product is mostly this.
  4. Where does authentication and authorisation actually live? The most common serious defect we find is authorisation checked in the interface but not on the server — the button is hidden, the endpoint is open. It is invisible in every demo and it is a breach.
  5. Can anyone explain the data model? Not the code — the model. If nobody on your team can describe why the schema is shaped the way it is, you don't own this system yet, whatever the repository says.
  6. Can you deploy it twice in a day without fear? If deployment is a person doing steps from memory, you don't have a production system. You have a demo with customers on it.

Keep, harden, or rewrite

Three outcomes. In our experience the middle one is most common and the last one is rarer than engineers instinctively want it to be.

Keep and extend

When: tests exist and are meaningful, the data model holds up, business rules mostly live in one place, and the unhappy paths merely need filling in.

This happens more often than the discourse suggests, particularly when an experienced engineer directed the AI rather than accepting whatever it produced. Add the missing hardening and carry on. Typically weeks.

Harden the core, replace the risky parts

When: the shape is right but specific areas are unsafe — usually authorisation, payments, anything touching money or personal data, and wherever duplication has concentrated.

The usual answer. Keep the parts that encode what you learned about your users, which is the genuinely valuable output of the prototype, and rebuild the parts where being wrong is expensive. Structuring this well is mostly about sequencing, and it's the same risk-first logic we apply to deciding how to contract a build.

Rewrite

When: the data model can't support what the business actually needs, or the same logic is scattered so widely that every change breaks something unrelated.

Genuinely necessary sometimes, and the trigger is almost always the data model rather than the code. Code is cheap to replace now. A schema with a year of customer data in it is not.

Even here the prototype earned its money: it told you what to build. That's a real return. It just wasn't the artefact — it was the learning.

How to think about the cost

We won't invent a range, because the number depends almost entirely on the audit above. But the shape is consistent, and it's the opposite of what people expect.

Prototype cost scales with features. Production cost scales with consequences. What drives the second number isn't how much the software does — it's what happens when each part of it is wrong. A prototype that displays the wrong number shows the wrong number. A production system that displays the wrong number sends the wrong invoice, and now you have a finance problem, a support problem, and possibly a regulatory one.

That's why "it's already 80% built" doesn't reduce the estimate the way people expect. The remaining work isn't the last 20% of the features. It's most of the engineering that makes any of the features safe to rely on — and, per the Stanford finding, it's being done in the mode where AI assistance contributes about 10% rather than 35–40%.

What we'd actually tell you to do

  • Keep building prototypes this way. Two weeks to a working demo is a genuine advantage, and validating an idea cheaply is exactly what it's for.
  • Stop calling the result an MVP. It's a prototype. The naming decides whether the next conversation is about hardening it or about shipping it, and those are different projects with different budgets.
  • Audit before you commit to a date. Two days of assessment ahead of a board commitment is the cheapest insurance available to you.
  • Budget the J-curve. The dip is normal and documented. Teams that plan for it get through it; teams that assumed a straight line conclude the project is failing and start making bad decisions in month two.
  • Fix the foundations before adding velocity. Tests and a deployment pipeline first. AI amplifies what's already there — including the absence of both.

The short version

  • The prototype was cheap because greenfield is where AI is strongest — 35–40% gains, versus roughly 10% on complex legacy code.
  • Shipping converts greenfield into legacy immediately. The productivity that built it doesn't carry over to running it.
  • The code is structurally different in a way that costs money later — record duplication, refactoring at a decade low. That's a tax on every future change, invisible in a demo.
  • Expect the estimate to be optimistic. Measured work has run behind perceived work by a wide margin, and "almost right" is the single most reported frustration in the field.
  • Audit first, then decide. Keep, harden, or rewrite — the data model usually decides which.

Sitting on a prototype with a launch date attached? Give us read access and a description of what it's meant to do, and we'll tell you which of the three it is and what stands between it and production — including when the answer is that it's in better shape than you feared. Have a look at how we approach product engineering, read what custom builds actually cost, or send it to an engineer.

Frequently asked questions

Why is it so expensive to turn an AI-built prototype into a real product?

Because the prototype was cheap for a reason that stops applying the moment you ship. Research from Stanford's Software Engineering Productivity programme, cited in DORA's 2026 ROI report, found AI delivers 35 to 40 percent productivity gains on simple greenfield tasks but only around 10 percent on complex legacy code. A prototype is pure greenfield. Once it has users, data you cannot lose and uptime expectations, it is complex legacy code, so the productivity that built it does not transfer to finishing it.

Is AI-generated code actually lower quality?

It is structurally different in ways that cost money later. GitClear's analysis of hundreds of millions of changed lines found refactoring fell from about 25 percent of changed lines in 2021 to under 10 percent in 2024, copy-pasted code rose from 8.3 to 12.3 percent, and block duplication reached a record 73.0 per million changed lines in 2026, up 81 percent on 2023. Duplication is not an aesthetic complaint, it is a multiplier on every future change, and none of it is visible in a demo.

Why do estimates for finishing an AI prototype run over?

There is a measured calibration problem. A METR randomised controlled trial with experienced developers on mature repositories found they were 19 percent slower with AI tools while believing they had been 20 percent faster, a 39 point gap. METR are explicit that this does not show AI fails to speed up developers generally. What survives their caveats is corroborated by Stack Overflow's 2025 survey, where the top frustration, cited by 66 percent, is output that is almost right but not quite, and 45 percent say debugging AI code takes longer.

Should I keep, harden, or rewrite an AI-generated prototype?

Hardening is most common. Keep and extend when tests are meaningful, the data model holds up and business rules mostly live in one place. Harden the core and replace the risky parts when the shape is right but authorisation, payments or anything touching money and personal data is unsafe. Rewrite when the data model cannot support what the business needs or logic is scattered so widely that every change breaks something unrelated. The trigger for a rewrite is almost always the data model, not the code.

How do I audit an AI-built prototype before committing to a launch date?

Six checks, about two days of work. Do tests assert business behaviour or just that the code does what it does? How many times does your most important business rule appear in the codebase? What happens on the unhappy path when you kill the network mid-transaction or submit twice? Is authorisation enforced on the server or only hidden in the interface? Can anyone explain the data model? And can you deploy twice in a day without fear?

What is the J-curve of AI adoption?

DORA's 2026 report describes value realisation from AI as a J-curve: productivity dips before it rises, which they call the tuition cost of transformation. They attribute the dip to three things, namely the learning curve as teams adapt workflows, the verification tax of reviewing AI-generated code which scales with volume, and process adaptation as testing and change approval strain under more code arriving faster. Teams that budget for the dip get through it; teams that assumed a straight line conclude the project is failing.

Should we stop building prototypes with AI?

No. Two weeks to a working demo is a real advantage and validating an idea cheaply is exactly what a prototype is for. The change worth making is in the naming and the next decision: call the result a prototype rather than an MVP, audit it before committing to a date, and fix tests and deployment before adding more velocity. DORA's 2025 finding is that AI amplifies what is already there, so adding speed to a codebase with no tests produces more of the same, sooner.

#ai#prototype#technical-debt#product#engineering