The Oracle Gap: Why AI Makes Some Engineering Teams Faster and Others Slower
One randomised trial of 4,867 developers found a 26% speedup. Another, of 16 expert maintainers, found them 19% slower while they felt faster. Both are right, and the variable that separates them is not the model. It is whether the work has an oracle.

Two of the most careful studies ever run on AI and software productivity disagree completely, and both of them are right.
The first is a set of three randomised controlled trials at Microsoft, Accenture and an anonymous Fortune 100 company, covering 4,867 developers. Giving them an AI coding assistant produced a 26.08% increase in completed tasks, with the largest gains going to the most junior engineers.
The second is METR's trial with 16 experienced open-source maintainers working 246 real issues in repositories they knew well — projects averaging over 22,000 stars and a million lines of code. With AI assistance they were 19% slower. They had forecast a 24% speedup beforehand. Afterwards, having just been measurably slowed down, they reported feeling 20% faster.
The industry mostly resolved this by picking a side. The honest resolution is that these two experiments were run in different regimes, and the variable that separates them is not experience, or model quality, or prompt skill. It is whether the work had an oracle: a cheap, automatic way to decide whether a given output is correct.
That distinction turns out to explain a great deal — including why your team's AI results probably look nothing like the case study you read, and why the industry's current answer to the problem cannot work.
The thing everyone agrees on, and the thing nobody does about it
There is now a consensus that verification is the constraint. DORA's 2025 research on AI-assisted development found that higher AI adoption is associated with an increase in software delivery throughput and an increase in delivery instability at the same time, and named the mechanism directly: "the time saved during initial code or content generation is often re-allocated to verification overhead and prompting overhead." They call it the verification tax. Industry surveys through 2026 have converged on the same finding from the other direction: reviewing AI-generated code is now the top-named bottleneck, and a large majority of developers say they do not fully trust the output they are shipping.
So far, so agreed. What follows is where the consensus goes wrong.
Almost every response treats verification cost as a fixed tax — something to be absorbed, governed, staffed, or routed through a smarter review queue. Buy a reviewing layer. Add context to the agents. Write better standards documents. Hire more reviewers.
Verification cost is not fixed. It is a property of how the system is built, and it varies across roughly four orders of magnitude between two pieces of code that do the same thing. Deciding whether a pure function that parses a date is correct can be done by a machine, exhaustively, in milliseconds. Deciding whether a 400-line change to a billing reconciliation path is correct can take a senior engineer an afternoon and still be wrong.
Computer science has a name for this and has studied it since the 1970s. It just isn't the name anyone is using right now.
The oracle problem
In software testing, an oracle is the mechanism that decides whether an observed behaviour is correct. Not the test input — the judgment. The survey that defines the field, Barr, Harman, McMinn, Shahbaz and Yoo in IEEE Transactions on Software Engineering (2015), puts the problem plainly: generating test inputs can be automated; deciding whether the resulting behaviour is right generally cannot. They identify oracle automation as the bottleneck that limits test automation overall. Without it, a human has to look.
Read that again with 2026 eyes. The field has known for decades that producing candidate behaviour is the cheap half and judging it is the expensive half. Large language models did not create this asymmetry. They removed the one thing that was hiding it.
Here is the part that matters, and it is the central claim of this article:
Writing code by hand was never purely production work. It was also the process by which an engineer built the mental model that let them judge whether the result was correct. The verification was bundled into the typing, invisibly and for free. Remove the typing and you do not just remove the production cost — you remove a verification subsidy nobody was accounting for.
This is not speculation. In January 2026 Anthropic published a randomised trial of 52 engineers learning an unfamiliar Python async library. The AI-assisted group finished about two minutes faster, a difference that did not reach statistical significance. On a quiz covering material they had used minutes earlier, they scored roughly 50% against 67% for the group that coded by hand — a 17-point gap, concentrated in debugging.
The detail inside that result is the useful one. Participants who used the assistant to ask conceptual questions scored 65% or higher. Those who delegated code generation scored below 40%. Same tool, same time budget, opposite outcomes, separated entirely by whether the human stayed in the comprehension loop. It is worth noting this is a vendor publishing an unflattering result about its own product category, which is a point in its favour rather than against it.
Lisanne Bainbridge described the general shape of this in 1983, in a paper about industrial control rooms called Ironies of Automation. Automate the parts a human does easily, and you leave them with only the hard part — monitoring a system that is usually right, and catching it on the rare occasion it isn't. Their skill at that task decays precisely because the automation is good. Forty years later we have rebuilt the same irony in the pull request queue.
Oracle strength is a spectrum, and it is the variable that matters
Most engineering work sits somewhere on a scale from "a machine can decide this instantly and exhaustively" to "only a senior human who understands the business can decide this, and they may be wrong."
Now put the two studies back on the same page. METR's developers were working deep in the bottom two bands: subtle changes to mature systems they personally maintained, where the oracle was their own accumulated understanding of a million-line codebase. Generation was nearly free. Judgment was the entire job, and it got harder, because reading unfamiliar code someone else produced is slower than reading code you just reasoned your way through.
Cui's 4,867 developers, across three large organisations and measured on tasks completed, were working much further up — more scoped units of work, more existing test coverage, more junior engineers whose personal oracle was weaker than the machinery around them. Handing them a generator was a straightforward win.
A necessary caution, because this is the point where a good argument usually overreaches. These two studies also differ in experience level, task type, measurement method and sample size, and METR's authors explicitly warn against generalising their result beyond expert developers in familiar, mature repositories. Oracle strength is the most useful lens I know of for reconciling them, and the evidence is consistent with it. It is not a controlled demonstration that oracle strength is the cause. Treat it as a model to test against your own data, not a law.
Why an AI reviewer cannot close the gap
The dominant commercial answer to the verification problem is to point a second model at the first model's output. This is intuitive and it is structurally broken, for a reason worth stating precisely.
An oracle is only useful to the extent its errors are independent of the generator's errors. That is the whole function of the thing. A compiler catches what you failed to notice because a compiler's notion of correctness was derived separately from your reasoning. A reviewer who shares your exact blind spots adds confidence without adding information.
Two bodies of evidence say LLM reviewers are substantially correlated with LLM generators. The first is the self-preference literature: models systematically score their own outputs higher, and the research finds the bias is most harmful on exactly the instances where the model performed poorly as a generator — stronger models produce fewer errors but show a greater tendency toward this bias when they are wrong. The failure is concentrated where you needed the check most.
The second is more direct. If a model does not understand a subtlety well enough to avoid it while writing the code, it does not understand it well enough to write the assertion that would catch it. The blind spot is a property of the model, not of the task it was pointed at.
This is not an argument against AI code review. It catches real defects and it is cheap. It is an argument against counting it as verification. In the ASE 2025 study Do LLMs Generate Useful Test Oracles?, Molinelli and colleagues found model-generated assertions scored comparably to developer-written ones on mutation detection — comparable, in a domain where developer-written oracles themselves catch under half of injected faults. Useful. Not a substitute for a check that is structurally independent.
Manufacturing oracles
If verification cost is the constraint and verification cost is a design property, then the highest-leverage engineering work available right now is building oracles. Not better prompts. Not more reviewers. Machinery that decides correctness without a human in the path.
These are ordered by strength, which is roughly inverse to how often they get used.
Make illegal states unrepresentable
The cheapest oracle is one that runs before the code does. Every invariant you push into the type system is a class of error a generator cannot produce at all, checked exhaustively, in milliseconds, for free, forever.
// Weak oracle: nothing here prevents the wrong order of arguments,
// and no test will catch it if both values happen to be plausible.
function transfer(from: string, to: string, amount: number) { /* ... */ }
// Strong oracle: the compiler rejects the mistake. A generated
// call site that confuses the two arguments does not build.
type AccountId = string & { readonly __brand: "AccountId" };
type Minor = number & { readonly __brand: "MinorUnits" };
function transfer(from: AccountId, to: AccountId, amount: Minor) { /* ... */ }
This mattered before AI. It matters more now, because the volume of code arriving at your boundaries went up by an order of magnitude and the compiler is the only reviewer that scales linearly with it.
Test properties, not examples
Example-based tests encode what the author thought to check. That is exactly the set of cases a model with similar training is also likely to think of. Property-based testing inverts it: state the invariant, let the machine hunt for a counterexample.
import fc from "fast-check";
// Not "parse('2026-01-31') === ..." but a law that must always hold.
fc.assert(
fc.property(fc.date({ min: new Date("1970-01-01") }), (d) => {
return parseDate(formatDate(d)).getTime() === d.getTime();
})
);
Round-trip laws, idempotence, commutativity, ordering and conservation are the workhorses. The generator searches the space you did not think about, and shrinks failures to a minimal reproducer — which is a second, underrated benefit, because a minimal reproducer is also the cheapest possible input to a debugging session.
Metamorphic testing: the technique for when you don't know the right answer
This is the most useful idea in this article and the one least present in current writing about AI code.
Often you cannot say what the correct output is. You can still say how outputs must relate to each other. A metamorphic relation is a property that must hold across multiple executions, and it gives you a partial oracle for systems that have no obvious one — search ranking, pricing, routing, recommendation, simulation, anything where the right answer is a matter of degree.
// You cannot assert the "correct" search results for a query.
// You can assert relations that any correct implementation obeys.
// 1. Adding a term that matches nothing must not add results.
expect(search(q + " zzqx").length).toBeLessThanOrEqual(search(q).length);
// 2. Reordering independent filters must not change the result set.
expect(ids(search(q, [filterA, filterB])))
.toEqual(ids(search(q, [filterB, filterA])));
// 3. A document that is a strict superset of a match must still match.
expect(search(q, { corpus: [doc + " extra prose"] })).not.toHaveLength(0);
Each relation is cheap to write, hard to satisfy accidentally, and — critically — derived from the specification rather than from the implementation. That independence is what makes it an oracle rather than a restatement.
Differential and shadow execution
When a trusted implementation exists, correctness is a diff. Run both against real traffic, compare, alert on divergence. This is the strongest practical oracle available for rewrites, migrations and optimisation work — which happens to be exactly the category people most want to hand to AI.
// Shadow the new path. Serve the old one. Compare asynchronously
// so a divergence is a dashboard entry, never a user-visible failure.
const result = legacyPricing(order);
queueMicrotask(() => {
try {
const candidate = generatedPricing(order);
if (!deepEqual(candidate, result)) {
recordDivergence({ order, expected: result, got: candidate });
}
} catch (err) {
recordDivergence({ order, expected: result, threw: String(err) });
}
});
return result;
The failure mode to design against is divergence volume. If the comparison is too strict — floating point, map ordering, timestamps — you will drown in noise and stop reading the dashboard within a week, which is worse than having no oracle because it feels like having one.
Assertions that survive into production
Tests check the inputs you imagined. Runtime invariants check the inputs your users actually send. For anything with a conservation law — money, inventory, seat allocation, double-entry anything — an assertion at the boundary is the highest-value oracle in the system, because it holds against traffic no test suite will ever generate.
When you cannot verify cheaply, make being wrong cheap
Some work has no affordable oracle. The correct response is not to verify harder but to change the cost of being wrong: narrow blast radius, feature flags, progressive rollout, fast and boring rollback, and migrations that are reversible by construction. Reversibility is not a substitute for correctness, but it converts an unbounded risk into a bounded one, and bounded risk is something an organisation can actually reason about.
A framework for deciding what to hand over
Put the two variables together — how strong the oracle is, and how expensive being wrong is — and the question of how much autonomy to grant answers itself.
| Oracle strength | Cost of being wrong | Appropriate autonomy | What to build first |
|---|---|---|---|
| Strong (total or partial) | Low | Generate freely; the machinery is the review | Nothing — this is where the 26% lives |
| Strong | High | Generate, but gate on the oracle and stage the rollout | Shadow execution and a rollback path |
| Weak | Low | Generate, ship behind a flag, let production be the oracle | Metrics that would actually show the failure |
| Weak | High | Do not delegate. This is where the 19% lives | An oracle — before writing any of the feature |
The bottom row is the one teams get wrong, and they get it wrong in a specific way: they attempt the work anyway, generate a large change quickly, and then discover that review of a 900-line diff in a weak-oracle domain is not a smaller task than writing it. It is a larger one.
The non-obvious move in that row is to treat oracle construction as the first deliverable of the feature, not as test coverage to be added afterwards. Spend the first day building the metamorphic relations, the shadow harness or the invariant monitor. Then the rest of the work moves up a row and you get to use the generator properly.
Where this model breaks
Four honest limits, because a framework that claims no failure modes is marketing.
Oracles cost real money. A shadow pipeline is infrastructure with its own maintenance and its own bugs. For a feature with a two-week lifespan, building one is a straightforward waste. The model justifies investment proportional to how long the code will live and how often it will be changed, not to how important it feels.
A wrong oracle is worse than none. An incorrect metamorphic relation or an over-fitted golden file produces confident, automated, repeated wrongness — and teams trust automated checks more than they trust their own reading. Oracles need review more carefully than the code they check.
Property tests can create false comfort. Passing a thousand generated cases tells you the invariants you wrote are satisfied. It says nothing about the invariant you failed to think of, and the feeling of thoroughness is disproportionate to the actual coverage.
Some work genuinely has no oracle. Whether an API will feel right to use in two years, whether an abstraction is carrying its weight, whether a product decision is correct — no machinery decides these, and pretending otherwise is how teams end up optimising something nobody wanted. This is the residue that stays human, and it is a larger fraction of senior engineering work than the current discourse admits.
What this implies for the next few years
If the constraint is verification rather than generation, several things that currently look like separate trends turn out to be the same trend.
Code review stops scaling and has to be replaced, not expanded. A team that went from fifteen pull requests a release to a hundred and fifty cannot fix that by adding reviewers, because review is a human-oracle activity and human oracles do not scale with headcount — they scale with context, which is the scarcest thing on the team. The change that survives is moving checks leftward into machinery.
Architecture becomes a verification decision. Choosing pure functions over stateful services, narrow interfaces over broad ones, and explicit data flow over implicit coupling used to be defended on maintainability grounds. The stronger argument now is economic: these choices make correctness machine-decidable, and machine-decidable correctness is the thing that converts a generator into throughput.
The junior pipeline is a real problem and it is not solved by banning the tools. If delegating generation costs 17 points of comprehension while asking conceptual questions costs nothing, then the distinction is not AI versus no AI. It is which interaction mode the organisation makes normal. That is a training and review-culture decision, and most companies have not made it deliberately.
The valuable AI products shift. The interesting work is not another model that writes code faster. It is machinery that establishes correctness: generating metamorphic relations from specifications, inferring invariants from production traffic, synthesising differential harnesses for migrations. Verification is the expensive half, and it is where the engineering is still mostly unbuilt.
Frequently asked
Does AI actually make developers faster?
It depends on whether the work has a cheap, automatic correctness check. With a strong oracle — clear specs, real coverage, machine-decidable output — the measured gains are large; Cui and colleagues found a 26% increase in tasks completed across 4,867 developers. Without one, on mature systems where correctness lives in a maintainer's head, METR measured experienced developers running 19% slower while believing they were faster. The average of those two numbers describes nobody.
What is a test oracle?
The mechanism that decides whether an observed behaviour is correct. A compiler, a type checker, an assertion, a reference implementation and a human reviewer are all oracles of different strength. The oracle problem — automating that judgment — is the long-standing bottleneck in test automation, and it is the bottleneck AI assistance has now pushed to the front of ordinary engineering work.
Can I just use AI to review AI-generated code?
As an additional filter, yes; it catches real defects cheaply. As your verification strategy, no. An oracle is valuable in proportion to how independent its errors are from the generator's, and the self-preference research finds model judgment is least reliable precisely on the outputs where the model was wrong. Pair it with checks derived from the specification rather than from the same reasoning.
What is metamorphic testing, and when should I use it?
Testing relations between executions rather than absolute outputs — adding an irrelevant term must not increase results, reordering independent filters must not change the result set. It is the standard answer to the oracle problem for systems where you cannot state the correct output: ranking, pricing, routing, recommendation, simulation. If you have ever said "we can't really test this," this is usually the technique you were missing.
Our AI-generated code passes all the tests and still breaks in production. Why?
Almost always because the tests are example-based and encode the same assumptions the generator made. Add checks whose origin is different from the implementation's: properties rather than examples, invariants that run against live traffic, and differential comparison against the previous behaviour.
Where should a team start?
Take the last three incidents caused by a change that passed review. For each, name the check that would have caught it automatically and ask why it does not exist. That exercise identifies your weakest oracles faster than any audit, and it is grounded in failures your organisation has already paid for.
The short version
AI made producing code close to free. It did not make knowing the code is right any cheaper, and writing by hand had been quietly subsidising that judgment all along. The teams getting large gains are the ones whose work already had machinery to decide correctness. The teams getting slower are generating faster into domains where the only oracle is a human being who now has more to read and less context to read it with.
The useful response is not to review harder or to point a second model at the first. It is to build the machinery — types that make errors unrepresentable, properties instead of examples, metamorphic relations where the right answer is unknowable, differential runs against what you trust, invariants that survive into production, and reversibility where none of that is affordable.
Code got cheap. Correctness did not. Engineering is now mostly the work of closing that distance.
If you are weighing where AI fits in a system you already depend on, or you want a second opinion on where your verification is thinnest, tell us what you are building. Our work on systems integration and the economics of running agents in production comes at the same problem from the other side.
Frequently asked questions
Does AI actually make developers faster?
It depends on whether the work has a cheap, automatic correctness check. With a strong oracle — clear specs, real coverage, machine-decidable output — the measured gains are large; Cui and colleagues found a 26% increase in tasks completed across 4,867 developers. Without one, on mature systems where correctness lives in a maintainer's head, METR measured experienced developers running 19% slower while believing they were faster. The average of those two numbers describes nobody.
What is a test oracle?
The mechanism that decides whether an observed behaviour is correct. A compiler, a type checker, an assertion, a reference implementation and a human reviewer are all oracles of different strength. The oracle problem — automating that judgment — is the long-standing bottleneck in test automation, and it is the bottleneck AI assistance has now pushed to the front of ordinary engineering work.
Can I just use AI to review AI-generated code?
As an additional filter, yes; it catches real defects cheaply. As your verification strategy, no. An oracle is valuable in proportion to how independent its errors are from the generator's, and the self-preference research finds model judgment is least reliable precisely on the outputs where the model was wrong. Pair it with checks derived from the specification rather than from the same reasoning.
What is metamorphic testing, and when should I use it?
Testing relations between executions rather than absolute outputs — adding an irrelevant term must not increase results, reordering independent filters must not change the result set. It is the standard answer to the oracle problem for systems where you cannot state the correct output: ranking, pricing, routing, recommendation, simulation. If you have ever said "we can't really test this," this is usually the technique you were missing.
Our AI-generated code passes all the tests and still breaks in production. Why?
Almost always because the tests are example-based and encode the same assumptions the generator made. Add checks whose origin is different from the implementation's: properties rather than examples, invariants that run against live traffic, and differential comparison against the previous behaviour.
Where should a team start?
Take the last three incidents caused by a change that passed review. For each, name the check that would have caught it automatically and ask why it does not exist. That exercise identifies your weakest oracles faster than any audit, and it is grounded in failures your organisation has already paid for.
Related Posts
Your Model Is Not a Function
A thousand identical requests at temperature zero produced eighty different answers. The usual explanation about floating point is only half of it — the other half is that your output depends on how many strangers were using the server at the same moment.
Your Context Window Is Not Memory
Every frontier model now advertises a million tokens. Published results show the usable fraction is far smaller, and it moves with the shape of your question. Why bigger windows and higher-recall retrieval both aim the wrong way, and how to measure your own ceiling.
Your Agent Has a Permissions Problem, Not a Prompt Problem
Prompt injection is a confused deputy attack, described in 1988 and unfixable at the model layer for the same reason it was unfixable then. The containment architecture already exists, most of it is standardised, and almost nobody ships it.
