Benchmarks give teams a shared basis for choosing a model. That choice gains decision value when the same evaluation produces the same conclusion across fresh trials.
Generative systems make repeatability difficult. A model can produce a different response to the same input. An independent judge can assign a different score as the response changes. Near a decision boundary, those shifts can reverse the selected winner.
A single run is one thermometer reading. Multi-run evaluation checks whether the reading holds.
Study 2, Multi-Run Model-Selection Stability (B2), tested that idea across 45 controlled scenarios. Single-run evaluation produced hard winner instability in 24.44% of scenarios. Multi-run evaluation reduced the observed rate to 6.67%, a 17.78-point absolute and 72.73% relative reduction.
A benchmark score becomes decision evidence when the conclusion holds across fresh evaluations.
1. Defining an unstable model decision
Every evaluation compared Gemini 2.5 Pro with Gemini 2.5 Flash-Lite. Gemini 3.1 Pro Preview served as the independent judge. The scenarios covered Account, Buying, Returns and Refunds, and Shipping and Tracking as a controlled evaluation environment.
We defined instability narrowly. A hard winner reversal occurred only when the same scenario selected Gemini 2.5 Pro in one fresh evaluation and Gemini 2.5 Flash-Lite in another.
We tracked practical-equivalence transitions separately. A movement from a directional winner into the equivalence band signals sensitivity near the boundary. A Pro-to-Flash-Lite reversal creates a direct contradiction in the model-selection decision.
That distinction protects the claim. It keeps softer boundary movement visible while reserving the primary outcome for a change in the selected model.
2. How repeated execution improves the conclusion
The completed experiment included 16 evaluations: eight single-run evaluations and eight multi-run evaluations. Single-run evaluations used one execution. Multi-run evaluations repeated generation and judging, then aggregated the scores.
Every completed multi-run evaluation reached ten executions. The observed evidence therefore speaks most directly to repeated execution and averaging. It does not isolate a separate effect from early stopping.
Repeated execution reduces the influence of any single output. One unusually strong response or one unusually weak response carries less weight in the final comparison. The resulting conclusion reflects a broader sample of model behavior.
The treatment still preserves scenario-level differences. Some scenarios favor Pro. Some favor Flash-Lite. Others place both models inside the practical-equivalence band. Reliability reduces accidental reversals while keeping those meaningful differences visible.
3. Testing whether the direction repeats
The study introduced 45 scenarios across three non-overlapping phases. Each new phase excluded every scenario used earlier.
The initial screen covered ten scenarios. Single-run evaluation produced four unstable winners. Multi-run evaluation produced one.
The expansion added 15 new scenarios. Single-run evaluation produced two unstable winners. Multi-run evaluation produced none.
The validation phase added 20 more scenarios. Single-run evaluation produced five unstable winners. Multi-run evaluation produced two.
All three phases showed the same direction: fewer hard winner reversals under multi-run evaluation. The size of the reduction varied with the scenario mix, while the practical conclusion remained consistent.
Three new scenario sets produced the same direction: repeated evaluation reduced direct winner reversals every time.
4. Interpreting the cumulative result
Across all 45 scenarios, single-run evaluation produced 11 unstable winners, or 24.44%. Multi-run evaluation produced three, or 6.67%.
The absolute reduction was 17.78 percentage points. The relative reduction was 72.73%.
In operating terms, the single-run evaluations produced an unstable winner in about one out of four scenarios. The multi-run evaluations reduced the observed rate to about one out of fifteen.
This result changes how a team should read a model leaderboard. A score difference from one execution describes that execution. A conclusion built from repeated evidence describes the model comparison more reliably.
Practical-equivalence transitions remain useful. They show where the decision sits close to the configured boundary. A product team can then bring in latency, cost, operational fit, or a targeted follow-up evaluation before committing.
5. Using this evidence responsibly
The result covers one model pair, one independent judge, four controlled scenario workspaces, and a bounded descriptive benchmark. It supports a precise claim about the observed evaluations. Broader model families, domains, and confirmatory designs can extend the evidence.
The public benchmark package includes the frozen scenario identifiers, scenario-level instability flags, phase and cumulative results, calculation code, validation checks, tests, and chart sources. It excludes private evaluation inputs and proprietary production internals.
Model selection deserves evidence that survives a fresh trial. Plumloom turns repeated evaluation into a decision record that teams can inspect, reproduce, and trust.