Leet Force
Back to Blog
DevelopmentUpdated

Your Model Is Not a Function

A thousand identical requests at temperature zero produced eighty different answers. The usual explanation about floating point is only half of it — the other half is that your output depends on how many strangers were using the server at the same moment.

Abdullah Khan
By Abdullah KhanCEO, Leet Force
Your Model Is Not a Function

Send the same prompt to the same model a thousand times. Temperature zero. No sampling. Identical parameters, identical weights, identical everything.

In September 2025, Horace He and Thinking Machines Lab did exactly that with Qwen3-235B and got 80 different answers. The thousand completions agreed perfectly for 102 tokens. At token 103 they began to split.

Nothing was wrong. No bug, no misconfiguration, no sampling left on by accident. This is what the system does.

Most engineers hold a quiet belief that temperature zero means determinism — that greedy decoding is a function, and a function given the same input returns the same output. That belief is wrong, and the reason it is wrong is stranger and more consequential than the usual explanation suggests. It is not really about floating point. It is about the fact that your result depends on how many other people were talking to the server at the same moment you were.

Which means every number you have about your AI system — every eval score, every benchmark comparison, every "I tested it and it worked" — is a single draw from a distribution whose width almost nobody has measured.

The explanation everyone gives, and why it is incomplete

Ask why LLM outputs vary and you will get a confident answer: GPUs run thousands of threads concurrently, floating point addition is not associative, so (a+b)+c and a+(b+c) give slightly different results, and the order in which concurrent threads accumulate is nondeterministic. Tiny numerical differences propagate into different token choices.

Every clause of that is true. As a complete explanation of what you observe through an API, it is still wrong.

He's analysis makes the correction precisely: atomic adds, the usual villain, are not necessary for the vast majority of kernels in an LLM forward pass. Run the same kernel twice on the same GPU with the same inputs and you generally get bit-identical results. The individual operations are reproducible.

What is not reproducible is the batch.

Inference servers do not process your request alone. They gather concurrent requests into a batch and run them together, because that is the only way GPU economics work. The kernels that implement RMSNorm, matrix multiplication and attention choose their reduction strategy — how to split the work across the hardware — based on how much work there is. A batch of four and a batch of forty get different split strategies, and different split strategies sum the same numbers in a different order.

The result is a property He names batch invariance, and almost no production kernel has it. Your tokens are computed slightly differently depending on batch size. Batch size depends on concurrent load. Concurrent load depends on strangers.

Two requests, byte-identical, sent a second apart, can produce different text — not because the model is creative, but because the second one arrived while more other people were using the service.

A thousand identical requests to Qwen3-235B at temperature zero produce one shared prefix for 102 tokens, then split into 80 distinct completions. Batch-invariant kernels collapse this to a single completion. 1,000 IDENTICAL REQUESTS · TEMPERATURE 0 · QWEN3-235B tokens 1–102: every run agrees token 103 first divergence 80 unique completions THE SAME TEST WITH BATCH-INVARIANT KERNELS 1 unique completion Bit-identical across all 1,000 runs — at roughly 1.6 to 2.1 times the inference cost.
The prefix is stable because the divergence needs somewhere to compound. Hover the lines.

Two papers that look like they disagree

A month after He's analysis, Yuan and colleagues published Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference and identified floating-point non-associativity as the root cause — the explanation He had just argued was incomplete.

The disagreement is apparent rather than real, and sorting it out is worth doing because the internet has mostly collapsed the two into one muddled story.

Floating-point non-associativity is the enabling condition. It is why summation order matters at all. If arithmetic were exact, reordering a reduction would change nothing and none of this would exist.

Batch-invariance failure is the trigger. It is what causes the summation order to differ between your two otherwise identical requests. Non-associativity explains why a different order produces a different number; batch variance explains why you got a different order.

You need both to explain what you see. And the practical consequence of separating them is that they have different fixes at different prices — which is the first useful thing to come out of this literature.

The cost of being sure

Yuan and colleagues quantified the enabling condition by varying numerical precision while holding everything else fixed. The numbers are larger than most people expect.

Standard deviation of benchmark accuracy across repeated runs under greedy decoding, by numerical precision, showing BF16 at 9.15 percent against FP32 at approximately zero. RUN-TO-RUN SWING IN BENCHMARK ACCURACY Standard deviation across repeated runs. Same model, same prompts, greedy decoding. DeepSeek-R1-Distill-Qwen-7B on AIME'24. BF16 BF16: 9.15 percentage points of standard deviation. Two models whose published scores differ by five points cannot be distinguished at this noise level from single runs. 9.15% FP16 FP16: 5.74 percentage points. Better, and still large enough to swamp most reported model-to-model differences. 5.74% FP32 FP32: approximately zero. Deterministic, at roughly double the memory and inference time of BF16. ~0% On the same benchmark, output length varied by a standard deviation of 9,189 tokens under BF16 and essentially zero under FP32. BF16 is the default almost everywhere.
Hover each bar. The precision that nearly every production deployment uses is the one with nine points of run-to-run swing.

Sit with the top bar for a moment. On a reasoning benchmark, the same model evaluated repeatedly under greedy decoding produced accuracy with a standard deviation of 9.15 percentage points in BF16. Output length swung with a standard deviation of over nine thousand tokens.

Almost every model comparison you have ever read reports a single run. If two models are five points apart on a benchmark with this much run-to-run variance, the comparison carries no information. It is a coin landing on its edge and being written down as a law of physics.

Both fixes are purchasable, and both have sticker prices:

FixWhat it removesMeasured costAvailable to you if…
Batch-invariant kernelsThe trigger — output no longer depends on concurrent load~2.1× slower, reduced to ~1.6× with optimised attentionYou run your own inference
FP32 inferenceThe enabling condition — numerics stop being order-sensitiveRoughly double memory and inference time against BF16You run your own inference
LayerCast (BF16 storage, FP32 compute)Same, more cheaplyFP32-level reproducibility at 34% less memory than full FP32You run your own inference
A seed parameterOnly the sampling layer, which was never the problemFreeAnyone — but see below

Three of those four require self-hosting. If you call a hosted API, you do not have access to the batch, the kernels or the precision. You have a seed parameter, and it is worth being clear about what a seed actually buys.

The seed does not do what you think

Providers offer a seed, and it is easy to read that as a determinism switch. OpenAI's own documentation is careful not to say so: with a seed set, the system makes a best effort to sample deterministically, and determinism is explicitly not guaranteed. The docs go further and note that even with the same seed and the same system_fingerprint, it is not uncommon to observe variability.

The mechanism explains why. A seed fixes the random number generator used to sample a token from the probability distribution. But the nondeterminism described above happens below that layer, while the distribution itself is being computed. Seeding the sampler does not help when the thing being sampled from has already shifted.

Two practical notes. The system_fingerprint identifies the backend configuration, and when it changes, a seed that produced one output may produce another — so it belongs in your logs next to the model name. And longer outputs are less reproducible than short ones even with a seed set, which follows directly from the divergence picture: a difference needs tokens to compound in.

The evidence from outside the lab

Everything so far is a mechanism, demonstrated under controlled conditions. The obvious objection is that it might not be visible in production, where effects this small could vanish into ordinary noise.

In April 2026, Paul Tschisgale and Peter Wulff published a study that is unglamorous in design and quietly remarkable in result. They took one multiple-choice physics problem and sent it to GPT-4o — a fixed snapshot, gpt-4o-2024-08-06 — every three hours for roughly three months. Same prompt, same model, 6,930 queries, from August to October 2025.

Mean accuracy was 0.632, with a standard deviation of 0.260. Then they ran a Fourier analysis on the time series, and the variation turned out not to be random.

Roughly 20% of total variance was periodic, with statistically significant peaks at weekly frequencies (7.3 and 5.5 days) and around the daily cycle (21.0 and 30.9 hours). Peak-to-peak variation was 0.139 on a 0–1 scale — about 14% of the full range. The model's measured ability to answer one fixed physics question rose and fell with the day of the week and the hour of the day.

The authors attribute this to server infrastructure dynamics: load follows human rhythms, and load-shedding strategies such as prompt pruning or quantization could translate into periodic performance changes.

Here is where intellectual honesty matters more than a tidy narrative. This is not proof that batch invariance caused it. The authors propose a different mechanism — load-shedding — and both are consistent with the data. It is a single model and a single question. What the two results share is the upstream driver: server load, which follows the clock.

Theory predicts output should depend on concurrent traffic. Three months of measurement finds performance moving on exactly the cycle concurrent traffic follows. That is corroboration, not demonstration, and it is strong enough to act on.

It also produces the most quotable practical instruction in this whole area, from the authors themselves: if you are measuring a model, your data collection should span at least one full week. An A/B test between two models run on a Tuesday and a Saturday is partly measuring the calendar.

The model is not a function

The useful shift is small to state and large in consequence.

Engineers reason about software as functions: input, output, repeat. That intuition is load-bearing for almost everything we do — unit tests, bug reports, A/B tests, SLAs, caching, incident response, audit trails. All of it assumes that the same input produces the same output, or at minimum that a difference means something changed.

An LLM endpoint is not a function. It is a sampler: it draws from a distribution, and the distribution itself shifts with conditions you do not control and cannot observe. Once you hold that model, several things that looked like puzzles resolve into the same answer.

What you observedFunction thinkingSampler thinking
The eval dropped 4 points after a prompt changeThe change made it worseYou have one sample from each of two distributions. You cannot yet tell them apart.
A user reports a bad answer; you cannot reproduce itBad report, or an environment differenceYou drew from the tail. It is real and it will recur at some rate.
The agent passed the test suite yesterday, fails todaySomething regressedPossibly nothing changed. Flakiness is the default, not the exception.
Model B beat Model A by 3 pointsB is betterBelow the noise floor. Run both many times or say nothing.
It works in staging, not in productionConfig driftLoad differs. That is itself an input.

Nothing on the right requires new tooling. It requires asking a different question — not what did it output but what is the shape of what it outputs — and that question has a measurable answer.

Measuring your own variance

This is an afternoon of work and it changes most arguments you are currently having by intuition.

  1. Pick 20 to 30 real inputs with known-good answers. Real ones, from production traffic, not curated examples.
  2. Run each one 20 times at your production settings. Keep every output.
  3. Record three numbers per input: the proportion of runs that were correct, the number of distinct outputs, and the spread in output length.
  4. Spread the runs across the week. Given the periodicity result, a batch fired at 3am on a Sunday is not representative of Tuesday mid-morning.
  5. Log the system_fingerprint alongside each result, where the provider exposes one. When it changes, your baseline is void.

What you now have is a noise floor. Any change smaller than it is not a change. That single number retires a surprising amount of unproductive debate — about whether a prompt tweak helped, whether a model upgrade is worth it, whether last week's regression was real.

The second thing it gives you is a per-input variance profile, and that is where the engineering decisions live. Inputs with high variance are not evenly distributed. They cluster around genuine ambiguity in the task, and they tell you exactly where to constrain the output, add a check, or route to a human.

The determinism ladder

Determinism is not a switch. It is a ladder with a price on every rung, and most teams should not climb to the top.

Five approaches to controlling output variance, ordered from cheapest to most expensive, with what each one costs and what it does not solve. CHEAPEST FIRST. MOST TEAMS SHOULD STOP AT RUNG 2. Measure it. Costs nothing but time, and it is the only rung that tells you whether you need the others. 1 · Measure the variance Repeated runs, spread across a week. You cannot manage a number you have never seen. free Constrain the output space. An enum of four values cannot drift into a fifth. This removes the variance that matters without touching the model. 2 · Shrink the output space Structured output, enums, schemas. Variance it cannot express is variance you never see. free Cache on a semantic key so repeat questions return the stored answer. Converts variance into staleness, which is a problem you already know how to manage. 3 · Cache the answer, not the call Trades variance for staleness — a failure mode your team already has tools for. storage Sample k times and take the majority. Reduces variance at k times the cost, and only works where answers can be compared for agreement. 4 · Vote across k samples Works only where two answers can be compared. Linear cost for sub-linear benefit. k× tokens Batch-invariant kernels, FP32 compute or LayerCast. True bit-level reproducibility, available only if you control the inference stack. 5 · Fix the numerics Batch-invariant kernels, FP32 or LayerCast. Self-hosted only. 1.6–2.1× compute Rung 2 is the one teams skip, and it is the one with the best ratio of variance removed to effort spent.
Hover each rung. Climbing to five is an infrastructure decision; most problems are solved at two.

The second rung deserves more attention than it gets. Most variance that reaches production is variance you invited by asking for free text where a constrained value would do. A classifier that returns one of four enum values cannot drift into a fifth. A field typed as an integer cannot come back as "approximately 40". Every degree of freedom you remove from the output is a degree of freedom the nondeterminism cannot express itself through — and unlike the rungs above it, this one is free and available on hosted APIs.

Where this actually bites

Not everywhere. That matters, because the correct response to this article is not to rebuild everything.

Variance is nearly harmless when the output is consumed by a human who will judge it anyway, when many phrasings are equally acceptable, and when being wrong is cheap and visible. Drafting, summarising, brainstorming, first-pass code a developer will read — you can ignore all of this.

It bites hard in five places:

  • Evaluation. Single-run numbers are samples. If your eval harness reports one figure with no interval, it is reporting a draw and calling it a measurement — and you have been making decisions on it.
  • Agent loops. Variance compounds across steps. A 2% per-step divergence is a different trajectory entirely by step thirty, which is why the same agent on the same task takes different paths on different days.
  • Debugging. "Cannot reproduce" stops being a verdict. The bug was real; you drew a different sample. The useful question is its rate, not its existence.
  • Caching and idempotency. A retried request is not guaranteed to return what the first one did. Anything that assumes retry-safety needs an explicit idempotency key holding the stored answer, not a re-request.
  • Anything you have to defend. Contracts, audits, regulated decisions. "The system produced this output" is a claim you cannot support by re-running it. You support it by having stored it.

That last point connects to a quiet irony. The software industry spent a decade on reproducible builds — bit-identical artifacts from identical sources — because supply-chain security demanded it. Then it adopted, at the centre of its most consequential new systems, a component that is not reproducible by default and whose output depends on other customers' traffic. The logging discipline that follows is not optional: if you cannot re-derive an output, the stored output is the only record that exists.

The counterargument

Three objections worth taking seriously.

"This is tiny in practice." Sometimes. The variance is concentrated in inputs where competing tokens have close probabilities — genuine ambiguity — and most production traffic is not ambiguous. But that is an argument for measuring rather than assuming, because it predicts your variance is unevenly distributed, and the clustered high-variance inputs are exactly the ones you most need to get right. A 9.15-point standard deviation on a reasoning benchmark is not tiny by any definition.

"It is getting fixed." It is. Batch-invariant kernels went from a research artifact to a working vLLM configuration in months, and the overhead has already fallen from 2.1× to 1.6×. Work on cheaper determinism continues. But every available fix today requires controlling the inference stack, and most teams call an API. For them the picture is unchanged.

"The periodicity study proves nothing about batch invariance." Correct, and the article has said so. It is one model, one question, and the authors propose a different mechanism. Its value is not as proof of a mechanism but as evidence that output quality moves on a schedule under real production conditions — which is actionable regardless of which mechanism produced it.

Frequently asked

Does temperature 0 make an LLM deterministic?

No. It removes sampling randomness, which is one source of variation among several. The dominant remaining source is that inference servers batch concurrent requests, and the kernels computing attention and matrix multiplication produce slightly different numerics at different batch sizes. Since batch size tracks server load, identical requests sent moments apart can diverge. One published test produced 80 distinct completions from 1,000 identical temperature-zero requests.

Why do I get different answers to the same prompt?

Three layers, in rough order of impact. Sampling, which temperature zero and a seed address. Batch-dependent numerics from concurrent load, which only the operator of the inference stack can address. And provider-side changes to model weights or serving configuration, which the system_fingerprint field exposes where it is available.

Does the seed parameter guarantee reproducible output?

No. Providers describe it as best effort and explicitly decline to guarantee determinism; variability is documented even with an identical seed and system fingerprint. A seed controls the sampler, while the main source of divergence acts on the distribution the sampler draws from. Longer outputs are less reproducible than shorter ones, because divergence needs tokens to compound in.

How many times should I run an eval?

Enough to see the spread, which usually means at least 20 runs per input, reported as a range rather than a single figure. Spread the runs across a full week: a three-month study of one fixed question found about 20% of performance variance followed daily and weekly cycles, so a benchmark run entirely on one afternoon measures that afternoon.

Can I make a hosted API deterministic?

Not at the numerical level — you do not control batching, kernels or precision. You can make the system behave deterministically by constraining the output space so variance has nowhere to express itself, caching answers on a semantic key, and voting across samples where correctness can be compared. Bit-level reproducibility requires running your own inference.

What does this cost to fix properly?

Batch-invariant kernels were measured at roughly 2.1× slower, improved to about 1.6× with optimised attention. FP32 inference roughly doubles memory and time against BF16; LayerCast reaches FP32-level reproducibility at about 34% less memory than full FP32. All three assume you run the model yourself.

The short version

Temperature zero is not determinism. The usual explanation — concurrent threads and floating-point non-associativity — names the enabling condition but not the trigger. The trigger is that inference servers batch requests, kernels are not batch-invariant, and batch composition depends on other people's traffic. Your output is coupled to strangers.

The consequence is not philosophical. On a reasoning benchmark, the same model under greedy decoding showed over nine points of run-to-run standard deviation at the precision almost everyone deploys. Three months of hourly sampling against a fixed question found a fifth of the performance variance moving on daily and weekly cycles. Most model comparisons, most eval deltas, and most "we improved the prompt" claims are reported from single runs and sit comfortably inside that noise.

Treat the model as a sampler rather than a function and the fixes are ordinary: measure the spread before trusting a number, constrain the output so variance cannot express itself, cache what must stay stable, and store what you may later have to defend. Full determinism is purchasable at roughly 1.6 to 2.1 times the compute, and only if you run the inference yourself. Most teams do not need it. Every team needs to know their noise floor, and almost none of them do.

If you are putting a model behind something where the same question has to produce the same answer — a pricing decision, an eligibility check, anything you might have to justify later — the variance profile is worth measuring before the architecture is settled. Tell us what you are building and we will work through where it needs to be pinned down. Our writing on knowing when AI output is correct and how much context a model can actually use covers the two questions this one usually arrives with.

Frequently asked questions

Does temperature 0 make an LLM deterministic?

No. It removes sampling randomness, which is one source of variation among several. The dominant remaining source is that inference servers batch concurrent requests, and the kernels computing attention and matrix multiplication produce slightly different numerics at different batch sizes. Since batch size tracks server load, identical requests sent moments apart can diverge. One published test produced 80 distinct completions from 1,000 identical temperature-zero requests.

Why do I get different answers to the same prompt?

Three layers, in rough order of impact. Sampling, which temperature zero and a seed address. Batch-dependent numerics from concurrent load, which only the operator of the inference stack can address. And provider-side changes to model weights or serving configuration, which the system_fingerprint field exposes where it is available.

Does the seed parameter guarantee reproducible output?

No. Providers describe it as best effort and explicitly decline to guarantee determinism; variability is documented even with an identical seed and system fingerprint. A seed controls the sampler, while the main source of divergence acts on the distribution the sampler draws from. Longer outputs are less reproducible than shorter ones, because divergence needs tokens to compound in.

How many times should I run an eval?

Enough to see the spread, which usually means at least 20 runs per input, reported as a range rather than a single figure. Spread the runs across a full week: a three-month study of one fixed question found about 20% of performance variance followed daily and weekly cycles, so a benchmark run entirely on one afternoon measures that afternoon.

Can I make a hosted API deterministic?

Not at the numerical level — you do not control batching, kernels or precision. You can make the system behave deterministically by constraining the output space so variance has nowhere to express itself, caching answers on a semantic key, and voting across samples where correctness can be compared. Bit-level reproducibility requires running your own inference.

What does this cost to fix properly?

Batch-invariant kernels were measured at roughly 2.1× slower, improved to about 1.6× with optimised attention. FP32 inference roughly doubles memory and time against BF16; LayerCast reaches FP32-level reproducibility at about 34% less memory than full FP32. All three assume you run the model yourself.

#llm#reliability#evaluation#reproducibility#architecture