Your Model Is Not a Function
A thousand identical requests at temperature zero produced eighty different answers. The usual explanation about floating point is only half of it — the other half is that your output depends on how many strangers were using the server at the same moment.

Send the same prompt to the same model a thousand times. Temperature zero. No sampling. Identical parameters, identical weights, identical everything.
In September 2025, Horace He and Thinking Machines Lab did exactly that with Qwen3-235B and got 80 different answers. The thousand completions agreed perfectly for 102 tokens. At token 103 they began to split.
Nothing was wrong. No bug, no misconfiguration, no sampling left on by accident. This is what the system does.
Most engineers hold a quiet belief that temperature zero means determinism — that greedy decoding is a function, and a function given the same input returns the same output. That belief is wrong, and the reason it is wrong is stranger and more consequential than the usual explanation suggests. It is not really about floating point. It is about the fact that your result depends on how many other people were talking to the server at the same moment you were.
Which means every number you have about your AI system — every eval score, every benchmark comparison, every "I tested it and it worked" — is a single draw from a distribution whose width almost nobody has measured.
The explanation everyone gives, and why it is incomplete
Ask why LLM outputs vary and you will get a confident answer: GPUs run thousands of threads concurrently, floating point addition is not associative, so (a+b)+c and a+(b+c) give slightly different results, and the order in which concurrent threads accumulate is nondeterministic. Tiny numerical differences propagate into different token choices.
Every clause of that is true. As a complete explanation of what you observe through an API, it is still wrong.
He's analysis makes the correction precisely: atomic adds, the usual villain, are not necessary for the vast majority of kernels in an LLM forward pass. Run the same kernel twice on the same GPU with the same inputs and you generally get bit-identical results. The individual operations are reproducible.
What is not reproducible is the batch.
Inference servers do not process your request alone. They gather concurrent requests into a batch and run them together, because that is the only way GPU economics work. The kernels that implement RMSNorm, matrix multiplication and attention choose their reduction strategy — how to split the work across the hardware — based on how much work there is. A batch of four and a batch of forty get different split strategies, and different split strategies sum the same numbers in a different order.
The result is a property He names batch invariance, and almost no production kernel has it. Your tokens are computed slightly differently depending on batch size. Batch size depends on concurrent load. Concurrent load depends on strangers.
Two requests, byte-identical, sent a second apart, can produce different text — not because the model is creative, but because the second one arrived while more other people were using the service.
Two papers that look like they disagree
A month after He's analysis, Yuan and colleagues published Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference and identified floating-point non-associativity as the root cause — the explanation He had just argued was incomplete.
The disagreement is apparent rather than real, and sorting it out is worth doing because the internet has mostly collapsed the two into one muddled story.
Floating-point non-associativity is the enabling condition. It is why summation order matters at all. If arithmetic were exact, reordering a reduction would change nothing and none of this would exist.
Batch-invariance failure is the trigger. It is what causes the summation order to differ between your two otherwise identical requests. Non-associativity explains why a different order produces a different number; batch variance explains why you got a different order.
You need both to explain what you see. And the practical consequence of separating them is that they have different fixes at different prices — which is the first useful thing to come out of this literature.
The cost of being sure
Yuan and colleagues quantified the enabling condition by varying numerical precision while holding everything else fixed. The numbers are larger than most people expect.
Sit with the top bar for a moment. On a reasoning benchmark, the same model evaluated repeatedly under greedy decoding produced accuracy with a standard deviation of 9.15 percentage points in BF16. Output length swung with a standard deviation of over nine thousand tokens.
Almost every model comparison you have ever read reports a single run. If two models are five points apart on a benchmark with this much run-to-run variance, the comparison carries no information. It is a coin landing on its edge and being written down as a law of physics.
Both fixes are purchasable, and both have sticker prices:
| Fix | What it removes | Measured cost | Available to you if… |
|---|---|---|---|
| Batch-invariant kernels | The trigger — output no longer depends on concurrent load | ~2.1× slower, reduced to ~1.6× with optimised attention | You run your own inference |
| FP32 inference | The enabling condition — numerics stop being order-sensitive | Roughly double memory and inference time against BF16 | You run your own inference |
| LayerCast (BF16 storage, FP32 compute) | Same, more cheaply | FP32-level reproducibility at 34% less memory than full FP32 | You run your own inference |
A seed parameter | Only the sampling layer, which was never the problem | Free | Anyone — but see below |
Three of those four require self-hosting. If you call a hosted API, you do not have access to the batch, the kernels or the precision. You have a seed parameter, and it is worth being clear about what a seed actually buys.
The seed does not do what you think
Providers offer a seed, and it is easy to read that as a determinism switch. OpenAI's own documentation is careful not to say so: with a seed set, the system makes a best effort to sample deterministically, and determinism is explicitly not guaranteed. The docs go further and note that even with the same seed and the same system_fingerprint, it is not uncommon to observe variability.
The mechanism explains why. A seed fixes the random number generator used to sample a token from the probability distribution. But the nondeterminism described above happens below that layer, while the distribution itself is being computed. Seeding the sampler does not help when the thing being sampled from has already shifted.
Two practical notes. The system_fingerprint identifies the backend configuration, and when it changes, a seed that produced one output may produce another — so it belongs in your logs next to the model name. And longer outputs are less reproducible than short ones even with a seed set, which follows directly from the divergence picture: a difference needs tokens to compound in.
The evidence from outside the lab
Everything so far is a mechanism, demonstrated under controlled conditions. The obvious objection is that it might not be visible in production, where effects this small could vanish into ordinary noise.
In April 2026, Paul Tschisgale and Peter Wulff published a study that is unglamorous in design and quietly remarkable in result. They took one multiple-choice physics problem and sent it to GPT-4o — a fixed snapshot, gpt-4o-2024-08-06 — every three hours for roughly three months. Same prompt, same model, 6,930 queries, from August to October 2025.
Mean accuracy was 0.632, with a standard deviation of 0.260. Then they ran a Fourier analysis on the time series, and the variation turned out not to be random.
Roughly 20% of total variance was periodic, with statistically significant peaks at weekly frequencies (7.3 and 5.5 days) and around the daily cycle (21.0 and 30.9 hours). Peak-to-peak variation was 0.139 on a 0–1 scale — about 14% of the full range. The model's measured ability to answer one fixed physics question rose and fell with the day of the week and the hour of the day.
The authors attribute this to server infrastructure dynamics: load follows human rhythms, and load-shedding strategies such as prompt pruning or quantization could translate into periodic performance changes.
Here is where intellectual honesty matters more than a tidy narrative. This is not proof that batch invariance caused it. The authors propose a different mechanism — load-shedding — and both are consistent with the data. It is a single model and a single question. What the two results share is the upstream driver: server load, which follows the clock.
Theory predicts output should depend on concurrent traffic. Three months of measurement finds performance moving on exactly the cycle concurrent traffic follows. That is corroboration, not demonstration, and it is strong enough to act on.
It also produces the most quotable practical instruction in this whole area, from the authors themselves: if you are measuring a model, your data collection should span at least one full week. An A/B test between two models run on a Tuesday and a Saturday is partly measuring the calendar.
The model is not a function
The useful shift is small to state and large in consequence.
Engineers reason about software as functions: input, output, repeat. That intuition is load-bearing for almost everything we do — unit tests, bug reports, A/B tests, SLAs, caching, incident response, audit trails. All of it assumes that the same input produces the same output, or at minimum that a difference means something changed.
An LLM endpoint is not a function. It is a sampler: it draws from a distribution, and the distribution itself shifts with conditions you do not control and cannot observe. Once you hold that model, several things that looked like puzzles resolve into the same answer.
| What you observed | Function thinking | Sampler thinking |
|---|---|---|
| The eval dropped 4 points after a prompt change | The change made it worse | You have one sample from each of two distributions. You cannot yet tell them apart. |
| A user reports a bad answer; you cannot reproduce it | Bad report, or an environment difference | You drew from the tail. It is real and it will recur at some rate. |
| The agent passed the test suite yesterday, fails today | Something regressed | Possibly nothing changed. Flakiness is the default, not the exception. |
| Model B beat Model A by 3 points | B is better | Below the noise floor. Run both many times or say nothing. |
| It works in staging, not in production | Config drift | Load differs. That is itself an input. |
Nothing on the right requires new tooling. It requires asking a different question — not what did it output but what is the shape of what it outputs — and that question has a measurable answer.
Measuring your own variance
This is an afternoon of work and it changes most arguments you are currently having by intuition.
- Pick 20 to 30 real inputs with known-good answers. Real ones, from production traffic, not curated examples.
- Run each one 20 times at your production settings. Keep every output.
- Record three numbers per input: the proportion of runs that were correct, the number of distinct outputs, and the spread in output length.
- Spread the runs across the week. Given the periodicity result, a batch fired at 3am on a Sunday is not representative of Tuesday mid-morning.
- Log the
system_fingerprintalongside each result, where the provider exposes one. When it changes, your baseline is void.
What you now have is a noise floor. Any change smaller than it is not a change. That single number retires a surprising amount of unproductive debate — about whether a prompt tweak helped, whether a model upgrade is worth it, whether last week's regression was real.
The second thing it gives you is a per-input variance profile, and that is where the engineering decisions live. Inputs with high variance are not evenly distributed. They cluster around genuine ambiguity in the task, and they tell you exactly where to constrain the output, add a check, or route to a human.
The determinism ladder
Determinism is not a switch. It is a ladder with a price on every rung, and most teams should not climb to the top.
The second rung deserves more attention than it gets. Most variance that reaches production is variance you invited by asking for free text where a constrained value would do. A classifier that returns one of four enum values cannot drift into a fifth. A field typed as an integer cannot come back as "approximately 40". Every degree of freedom you remove from the output is a degree of freedom the nondeterminism cannot express itself through — and unlike the rungs above it, this one is free and available on hosted APIs.
Where this actually bites
Not everywhere. That matters, because the correct response to this article is not to rebuild everything.
Variance is nearly harmless when the output is consumed by a human who will judge it anyway, when many phrasings are equally acceptable, and when being wrong is cheap and visible. Drafting, summarising, brainstorming, first-pass code a developer will read — you can ignore all of this.
It bites hard in five places:
- Evaluation. Single-run numbers are samples. If your eval harness reports one figure with no interval, it is reporting a draw and calling it a measurement — and you have been making decisions on it.
- Agent loops. Variance compounds across steps. A 2% per-step divergence is a different trajectory entirely by step thirty, which is why the same agent on the same task takes different paths on different days.
- Debugging. "Cannot reproduce" stops being a verdict. The bug was real; you drew a different sample. The useful question is its rate, not its existence.
- Caching and idempotency. A retried request is not guaranteed to return what the first one did. Anything that assumes retry-safety needs an explicit idempotency key holding the stored answer, not a re-request.
- Anything you have to defend. Contracts, audits, regulated decisions. "The system produced this output" is a claim you cannot support by re-running it. You support it by having stored it.
That last point connects to a quiet irony. The software industry spent a decade on reproducible builds — bit-identical artifacts from identical sources — because supply-chain security demanded it. Then it adopted, at the centre of its most consequential new systems, a component that is not reproducible by default and whose output depends on other customers' traffic. The logging discipline that follows is not optional: if you cannot re-derive an output, the stored output is the only record that exists.
The counterargument
Three objections worth taking seriously.
"This is tiny in practice." Sometimes. The variance is concentrated in inputs where competing tokens have close probabilities — genuine ambiguity — and most production traffic is not ambiguous. But that is an argument for measuring rather than assuming, because it predicts your variance is unevenly distributed, and the clustered high-variance inputs are exactly the ones you most need to get right. A 9.15-point standard deviation on a reasoning benchmark is not tiny by any definition.
"It is getting fixed." It is. Batch-invariant kernels went from a research artifact to a working vLLM configuration in months, and the overhead has already fallen from 2.1× to 1.6×. Work on cheaper determinism continues. But every available fix today requires controlling the inference stack, and most teams call an API. For them the picture is unchanged.
"The periodicity study proves nothing about batch invariance." Correct, and the article has said so. It is one model, one question, and the authors propose a different mechanism. Its value is not as proof of a mechanism but as evidence that output quality moves on a schedule under real production conditions — which is actionable regardless of which mechanism produced it.
Frequently asked
Does temperature 0 make an LLM deterministic?
No. It removes sampling randomness, which is one source of variation among several. The dominant remaining source is that inference servers batch concurrent requests, and the kernels computing attention and matrix multiplication produce slightly different numerics at different batch sizes. Since batch size tracks server load, identical requests sent moments apart can diverge. One published test produced 80 distinct completions from 1,000 identical temperature-zero requests.
Why do I get different answers to the same prompt?
Three layers, in rough order of impact. Sampling, which temperature zero and a seed address. Batch-dependent numerics from concurrent load, which only the operator of the inference stack can address. And provider-side changes to model weights or serving configuration, which the system_fingerprint field exposes where it is available.
Does the seed parameter guarantee reproducible output?
No. Providers describe it as best effort and explicitly decline to guarantee determinism; variability is documented even with an identical seed and system fingerprint. A seed controls the sampler, while the main source of divergence acts on the distribution the sampler draws from. Longer outputs are less reproducible than shorter ones, because divergence needs tokens to compound in.
How many times should I run an eval?
Enough to see the spread, which usually means at least 20 runs per input, reported as a range rather than a single figure. Spread the runs across a full week: a three-month study of one fixed question found about 20% of performance variance followed daily and weekly cycles, so a benchmark run entirely on one afternoon measures that afternoon.
Can I make a hosted API deterministic?
Not at the numerical level — you do not control batching, kernels or precision. You can make the system behave deterministically by constraining the output space so variance has nowhere to express itself, caching answers on a semantic key, and voting across samples where correctness can be compared. Bit-level reproducibility requires running your own inference.
What does this cost to fix properly?
Batch-invariant kernels were measured at roughly 2.1× slower, improved to about 1.6× with optimised attention. FP32 inference roughly doubles memory and time against BF16; LayerCast reaches FP32-level reproducibility at about 34% less memory than full FP32. All three assume you run the model yourself.
The short version
Temperature zero is not determinism. The usual explanation — concurrent threads and floating-point non-associativity — names the enabling condition but not the trigger. The trigger is that inference servers batch requests, kernels are not batch-invariant, and batch composition depends on other people's traffic. Your output is coupled to strangers.
The consequence is not philosophical. On a reasoning benchmark, the same model under greedy decoding showed over nine points of run-to-run standard deviation at the precision almost everyone deploys. Three months of hourly sampling against a fixed question found a fifth of the performance variance moving on daily and weekly cycles. Most model comparisons, most eval deltas, and most "we improved the prompt" claims are reported from single runs and sit comfortably inside that noise.
Treat the model as a sampler rather than a function and the fixes are ordinary: measure the spread before trusting a number, constrain the output so variance cannot express itself, cache what must stay stable, and store what you may later have to defend. Full determinism is purchasable at roughly 1.6 to 2.1 times the compute, and only if you run the inference yourself. Most teams do not need it. Every team needs to know their noise floor, and almost none of them do.
If you are putting a model behind something where the same question has to produce the same answer — a pricing decision, an eligibility check, anything you might have to justify later — the variance profile is worth measuring before the architecture is settled. Tell us what you are building and we will work through where it needs to be pinned down. Our writing on knowing when AI output is correct and how much context a model can actually use covers the two questions this one usually arrives with.
Frequently asked questions
Does temperature 0 make an LLM deterministic?
No. It removes sampling randomness, which is one source of variation among several. The dominant remaining source is that inference servers batch concurrent requests, and the kernels computing attention and matrix multiplication produce slightly different numerics at different batch sizes. Since batch size tracks server load, identical requests sent moments apart can diverge. One published test produced 80 distinct completions from 1,000 identical temperature-zero requests.
Why do I get different answers to the same prompt?
Three layers, in rough order of impact. Sampling, which temperature zero and a seed address. Batch-dependent numerics from concurrent load, which only the operator of the inference stack can address. And provider-side changes to model weights or serving configuration, which the system_fingerprint field exposes where it is available.
Does the seed parameter guarantee reproducible output?
No. Providers describe it as best effort and explicitly decline to guarantee determinism; variability is documented even with an identical seed and system fingerprint. A seed controls the sampler, while the main source of divergence acts on the distribution the sampler draws from. Longer outputs are less reproducible than shorter ones, because divergence needs tokens to compound in.
How many times should I run an eval?
Enough to see the spread, which usually means at least 20 runs per input, reported as a range rather than a single figure. Spread the runs across a full week: a three-month study of one fixed question found about 20% of performance variance followed daily and weekly cycles, so a benchmark run entirely on one afternoon measures that afternoon.
Can I make a hosted API deterministic?
Not at the numerical level — you do not control batching, kernels or precision. You can make the system behave deterministically by constraining the output space so variance has nowhere to express itself, caching answers on a semantic key, and voting across samples where correctness can be compared. Bit-level reproducibility requires running your own inference.
What does this cost to fix properly?
Batch-invariant kernels were measured at roughly 2.1× slower, improved to about 1.6× with optimised attention. FP32 inference roughly doubles memory and time against BF16; LayerCast reaches FP32-level reproducibility at about 34% less memory than full FP32. All three assume you run the model yourself.
Related Posts
Your Context Window Is Not Memory
Every frontier model now advertises a million tokens. Published results show the usable fraction is far smaller, and it moves with the shape of your question. Why bigger windows and higher-recall retrieval both aim the wrong way, and how to measure your own ceiling.
AI Agent Cost Per Task: Why Your Bill Is a Reliability Problem
Token cost grows with the square of the steps, the chance of needing another attempt grows exponentially with them, and the human who cleans up afterwards costs more than both. Here is the model we use to decide whether an agent is viable, with the arithmetic shown.
Your Agent Has a Permissions Problem, Not a Prompt Problem
Prompt injection is a confused deputy attack, described in 1988 and unfixable at the model layer for the same reason it was unfixable then. The containment architecture already exists, most of it is standardised, and almost nobody ships it.
