Does the same AI reviewer give the same release advice twice?
150 executions of one frozen configuration over 30 identical change packets, comparing what each run raised, ranked and proposed to test.
Question
Give an AI reviewer the same release, twice, with nothing changed between the two — same model, same prompt, same bytes. Does it give you the same advice?
This matters commercially rather than academically. If a review board is going to act on a machine's opinion about a release, somebody has to know how much of that opinion is a property of the release and how much is a property of the run.
Short answer
We ran one frozen configuration five times over each of 30 identical change packets — 150 executions in total. The five runs named the same single most important risk surface 60% of the time on average, and somewhere between 20.6% and 68.8% of the mechanisms they raised appeared in only one run out of the five. What the model surfaces moves a great deal; how it ranks and scores what it does surface barely moves at all.
What we measured
The design. 30 change packets, drawn under a published salt from 475 eligible transitions across 21 distinct servers, with a per-server cap. Each packet was submitted five times to one fixed configuration: same model, same prompt bytes, same evidence. 150 of 150 executions succeeded, and all five outputs of every packet were byte-distinct — so nothing below is a caching artefact.
What was compared, and at what grain. Runs were compared at three levels: the coarse surface a risk points at, the specific identifiers it names, and the assertion tokens of the acceptance tests it proposes. Comparing only at the coarsest level would have flattered the result; comparing only at the finest would have counted two descriptions of one thing as a disagreement. Both are reported.
Why one figure is a range and not a point. Deciding whether two differently-worded risks are the same mechanism is itself a semantic judgment. The matcher that would have made that judgment was never run — the budget stop landed first — so the result is bounded from both sides instead: one anchor over-merges and one under-merges, and the truth is between them. Publishing a midpoint would be inventing a number that no instrument produced.
Result
60%mean over 30 packets
How often two runs of the same packet name the same top risk surface. Per-packet spread is the story: the 10th percentile is 19% and the 90th is 100%. There is no single "how stable is it" number — there are stable packets and unstable ones.
21%–69%of mechanisms appear in 1 of 5 runs
The bracket runs from 20.6% under the anchor that over-merges to 68.8% under the anchor that under-merges. The range is the answer, not an imprecise estimate of a point inside it: both ends are measured, and the judgment that would have narrowed them was never made.
0.83rank correlation, recurring risks
Among risks that do recur across runs, ordering barely moves, and confidence moves less: a standard deviation of 0.046 within a recurring risk against 0.106 across all risks. The instability is in which mechanisms surface, not in how they are ranked once they do.
The finer the grain, the less it agrees
| Compared at | Mean agreement across run pairs |
|---|---|
| the risk surface, as a set | 0.62 |
| the single top-ranked surface | 0.60 |
| the specific identifiers named | 0.33 |
| the assertion tokens of the proposed tests | 0.26 |
The proposed acceptance tests move more than the risks do.
One number that is perfectly stable, and it is the wrong one
The count of proposed tests has a variance of 0: every one of the 150 executions produced exactly five, against a prompt that said in terms that fewer was acceptable and would not be penalised. The count is stable because it is pinned at the cap, not because anything was triaged. A layer that decides which of the five deserves a test is doing work the model demonstrably will not do for itself.
The two ends of the same instrument
A second, smaller experiment scored 24 model-generated risk analyses twice with one frozen scoring instrument. On the concrete question — would this specific proposed test have failed against the new release? — the two runs agreed on every unit: Cohen's kappa 1.00. On the semantic question — is this the same mechanism that actually broke? — kappa was 0.80.
The more a judgment is a semantic match between free text and a terse record, the more the instrument moves. And at the other extreme: the deterministic retrieval step underneath all of this was run 1,500 times across fifteen batteries, including five that shuffled the corpus and rebuilt the index, and disagreed with itself 0 times.
What changed in PRIVI because of it
- The ship/no-ship verdict is produced by named rules over supplied evidence, never by a model's opinion. PROMOTE, REVIEW and BLOCK come from a fixed set of rules, each one printed in the artifact beside the exact conditions that would clear it, with no score anywhere in the path. Re-running it over the same inputs produces the same bytes. That design is a direct consequence of this result.
- Change classification uses a closed vocabulary rather than free-text judgment. What changed is drawn from a fixed list of change categories, each carrying the checks it touches and the nearby safe change it must not fire on. A closed vocabulary can be argued with; a paraphrase cannot.
- We do not ship a confidence signal built from repetition. "This risk was raised in five of five independent reviews" is easy to build and was, for a while, the plan. It was dropped: in the follow-up experiment, risks that turned out to name the actual mechanism recurred slightly less often than risks that did not. Repeating an opinion does not make it true, and we are not going to sell it as though it did.
- Where semantic judgment is unavoidable, it is labelled as such on the artifact rather than presented beside deterministic findings as though the two were the same kind of thing.
What this does not prove
- It does not show the model is wrong. Nothing in this experiment was scored against a real outcome. It measures agreement between runs, which is a different question from accuracy, and a run that is unstable can still be useful.
- It does not show that PRIVI's approach is more accurate. No comparison arm was run. We have not measured any increment from anything we build, and we do not claim one.
- The 21%–69% bracket is not a precision failure to be tidied up later. Both ends are measured, and they bound the answer from either side because the judgment that would narrow them was never made. It should be read as a range, and it is reported as one everywhere.
- One model, one prompt, one family of change packets, on one day. A different prompt, a lower temperature, or a narrower question could all move these numbers, and none of that was tested.
- The scoring figures rest on a small sample. 24 double-scored units across twelve cases; the interval on the 0.80 kappa runs from 0.47 to 1.00. That sample can rule out a broken instrument. It cannot certify a good one.
- The two scoring runs are one instrument executed twice, not two independent judges. It measures repeatability, which is an upper bound on what a genuinely different second scorer would show.
Method
Every execution used one model at one effort setting, verified from the transcript of each run rather than asserted. The sample was drawn under a salt recorded before the draw, from a population that excluded every transition and every server reserved for a future benchmark; no reserved case was touched and no outcome label was read while the analyses were being generated.
The comparison statistics are deterministic: no model is involved in any figure in the first three tables, so they survive independently of any judgment about the matcher. Percentile figures are computed across packets, and the per-packet spread is published rather than summarised away.
What was not done, stated rather than discovered later. The semantic matcher was never executed. The scoring experiment reached 24 of a designed 36 units before its budget stop, and is recorded as incomplete rather than as complete. No adversarial audit of these findings has been run.
Artifacts
noise-floor-v1/prediction-noise/analysis/DETERMINISTIC.json— every agreement statistic above, with its distribution across packets.noise-floor-v1/prediction-noise/sample/SAMPLE.json— the sampling frame, the salt, the strata and the per-server cap.noise-floor-v1/retrieval-noise/DETERMINISM.json— the fifteen retrieval batteries and their comparison.noise-floor-v1/scoring-noise/analysis/— the scoring agreement figures and the two disagreements, described rather than adjudicated.
These artifacts live in the private research repository rather than beside this page:
they carry per-server change classifications this project has reported as unreliable, and naming third
parties beside an unreliable classification is not something a published figure is worth. The figures
themselves are re-derived from them at every build, against
RESEARCH-CLAIMS-INVENTORY.json.