Does the same AI reviewer give the same release advice twice?

150 executions of one frozen configuration over 30 identical change packets, comparing what each run raised, ranked and proposed to test.

150 executions 30 packets one fixed configuration

Question

Give an AI reviewer the same release, twice, with nothing changed between the two — same model, same prompt, same bytes. Does it give you the same advice?

This matters commercially rather than academically. If a review board is going to act on a machine's opinion about a release, somebody has to know how much of that opinion is a property of the release and how much is a property of the run.

Short answer

We ran one frozen configuration five times over each of 30 identical change packets — 150 executions in total. The five runs named the same single most important risk surface 60% of the time on average, and somewhere between 20.6% and 68.8% of the mechanisms they raised appeared in only one run out of the five. What the model surfaces moves a great deal; how it ranks and scores what it does surface barely moves at all.

What we measured

The design. 30 change packets, drawn under a published salt from 475 eligible transitions across 21 distinct servers, with a per-server cap. Each packet was submitted five times to one fixed configuration: same model, same prompt bytes, same evidence. 150 of 150 executions succeeded, and all five outputs of every packet were byte-distinct — so nothing below is a caching artefact.

What was compared, and at what grain. Runs were compared at three levels: the coarse surface a risk points at, the specific identifiers it names, and the assertion tokens of the acceptance tests it proposes. Comparing only at the coarsest level would have flattered the result; comparing only at the finest would have counted two descriptions of one thing as a disagreement. Both are reported.

Why one figure is a range and not a point. Deciding whether two differently-worded risks are the same mechanism is itself a semantic judgment. The matcher that would have made that judgment was never run — the budget stop landed first — so the result is bounded from both sides instead: one anchor over-merges and one under-merges, and the truth is between them. Publishing a midpoint would be inventing a number that no instrument produced.

Result

60%mean over 30 packets

How often two runs of the same packet name the same top risk surface. Per-packet spread is the story: the 10th percentile is 19% and the 90th is 100%. There is no single "how stable is it" number — there are stable packets and unstable ones.

21%–69%of mechanisms appear in 1 of 5 runs

The bracket runs from 20.6% under the anchor that over-merges to 68.8% under the anchor that under-merges. The range is the answer, not an imprecise estimate of a point inside it: both ends are measured, and the judgment that would have narrowed them was never made.

0.83rank correlation, recurring risks

Among risks that do recur across runs, ordering barely moves, and confidence moves less: a standard deviation of 0.046 within a recurring risk against 0.106 across all risks. The instability is in which mechanisms surface, not in how they are ranked once they do.

The finer the grain, the less it agrees

Compared atMean agreement across run pairs
the risk surface, as a set0.62
the single top-ranked surface0.60
the specific identifiers named0.33
the assertion tokens of the proposed tests0.26

The proposed acceptance tests move more than the risks do.

One number that is perfectly stable, and it is the wrong one

The count of proposed tests has a variance of 0: every one of the 150 executions produced exactly five, against a prompt that said in terms that fewer was acceptable and would not be penalised. The count is stable because it is pinned at the cap, not because anything was triaged. A layer that decides which of the five deserves a test is doing work the model demonstrably will not do for itself.

The two ends of the same instrument

A second, smaller experiment scored 24 model-generated risk analyses twice with one frozen scoring instrument. On the concrete question — would this specific proposed test have failed against the new release? — the two runs agreed on every unit: Cohen's kappa 1.00. On the semantic question — is this the same mechanism that actually broke? — kappa was 0.80.

The more a judgment is a semantic match between free text and a terse record, the more the instrument moves. And at the other extreme: the deterministic retrieval step underneath all of this was run 1,500 times across fifteen batteries, including five that shuffled the corpus and rebuilt the index, and disagreed with itself 0 times.

What changed in PRIVI because of it

What this does not prove

Method

Every execution used one model at one effort setting, verified from the transcript of each run rather than asserted. The sample was drawn under a salt recorded before the draw, from a population that excluded every transition and every server reserved for a future benchmark; no reserved case was touched and no outcome label was read while the analyses were being generated.

The comparison statistics are deterministic: no model is involved in any figure in the first three tables, so they survive independently of any judgment about the matcher. Percentile figures are computed across packets, and the per-packet spread is published rather than summarised away.

What was not done, stated rather than discovered later. The semantic matcher was never executed. The scoring experiment reached 24 of a designed 36 units before its budget stop, and is recorded as incomplete rather than as complete. No adversarial audit of these findings has been run.

Artifacts

These artifacts live in the private research repository rather than beside this page: they carry per-server change classifications this project has reported as unreliable, and naming third parties beside an unreliable classification is not something a published figure is worth. The figures themselves are re-derived from them at every build, against RESEARCH-CLAIMS-INVENTORY.json.