ResearchBlogHelp & SupportLearn EvalsComing Soon
    Pricing
    Back to blog
    AI Evaluation··3 min read

    Why One Evaluation Run Is Often the Wrong Basis for Model Selection

    Across 45 controlled scenarios, multi-run evaluation reduced hard winner instability by 72.73%. A single run is one thermometer reading — repeated evaluation checks whether the reading holds.

    Benchmarks give teams a shared basis for choosing a model. That choice gains decision value when the same evaluation produces the same conclusion across fresh trials.

    Generative systems make repeatability difficult. A model can produce a different response to the same input. An independent judge can assign a different score as the response changes. Near a decision boundary, those shifts can reverse the selected winner.

    A single run is one thermometer reading. Multi-run evaluation checks whether the reading holds.

    Study 2, Multi-Run Model-Selection Stability (B2), tested that idea across 45 controlled scenarios. Single-run evaluation produced hard winner instability in 24.44% of scenarios. Multi-run evaluation reduced the observed rate to 6.67%, a 17.78-point absolute and 72.73% relative reduction.

    Model-Winner Instability · 45 Scenarios
    72.73%
    relative reduction in hard winner reversals under multi-run evaluation
    0%8%16%30%Single-runMulti-run24.44%6.67%
    Single-run unstable
    11 / 45
    24.44%
    Multi-run unstable
    3 / 45
    6.67%
    Source: model-selection-stability-summary.json · Absolute reduction 17.78 pp.

    A benchmark score becomes decision evidence when the conclusion holds across fresh evaluations.

    1. Defining an unstable model decision

    Every evaluation compared Gemini 2.5 Pro with Gemini 2.5 Flash-Lite. Gemini 3.1 Pro Preview served as the independent judge. The scenarios covered Account, Buying, Returns and Refunds, and Shipping and Tracking as a controlled evaluation environment.

    We defined instability narrowly. A hard winner reversal occurred only when the same scenario selected Gemini 2.5 Pro in one fresh evaluation and Gemini 2.5 Flash-Lite in another.

    We tracked practical-equivalence transitions separately. A movement from a directional winner into the equivalence band signals sensitivity near the boundary. A Pro-to-Flash-Lite reversal creates a direct contradiction in the model-selection decision.

    That distinction protects the claim. It keeps softer boundary movement visible while reserving the primary outcome for a change in the selected model.

    2. How repeated execution improves the conclusion

    The completed experiment included 16 evaluations: eight single-run evaluations and eight multi-run evaluations. Single-run evaluations used one execution. Multi-run evaluations repeated generation and judging, then aggregated the scores.

    Every completed multi-run evaluation reached ten executions. The observed evidence therefore speaks most directly to repeated execution and averaging. It does not isolate a separate effect from early stopping.

    Repeated execution reduces the influence of any single output. One unusually strong response or one unusually weak response carries less weight in the final comparison. The resulting conclusion reflects a broader sample of model behavior.

    The treatment still preserves scenario-level differences. Some scenarios favor Pro. Some favor Flash-Lite. Others place both models inside the practical-equivalence band. Reliability reduces accidental reversals while keeping those meaningful differences visible.

    3. Testing whether the direction repeats

    The study introduced 45 scenarios across three non-overlapping phases. Each new phase excluded every scenario used earlier.

    The initial screen covered ten scenarios. Single-run evaluation produced four unstable winners. Multi-run evaluation produced one.

    The expansion added 15 new scenarios. Single-run evaluation produced two unstable winners. Multi-run evaluation produced none.

    The validation phase added 20 more scenarios. Single-run evaluation produced five unstable winners. Multi-run evaluation produced two.

    Instability Rate By Phase
    Three non-overlapping scenario sets, same direction every time
    Single-runMulti-run
    Initial screenExpansionValidation0%15%30%50%40%13.33%25%10%0%10%
    Initial screen
    10 scenarios
    4 single1 multi
    Expansion
    15 scenarios
    2 single0 multi
    Validation
    20 scenarios
    5 single2 multi
    Source: phase-results.csv · Unstable winners under single vs. multi-run evaluation.

    All three phases showed the same direction: fewer hard winner reversals under multi-run evaluation. The size of the reduction varied with the scenario mix, while the practical conclusion remained consistent.

    Three new scenario sets produced the same direction: repeated evaluation reduced direct winner reversals every time.

    4. Interpreting the cumulative result

    Across all 45 scenarios, single-run evaluation produced 11 unstable winners, or 24.44%. Multi-run evaluation produced three, or 6.67%.

    The absolute reduction was 17.78 percentage points. The relative reduction was 72.73%.

    In operating terms, the single-run evaluations produced an unstable winner in about one out of four scenarios. The multi-run evaluations reduced the observed rate to about one out of fifteen.

    This result changes how a team should read a model leaderboard. A score difference from one execution describes that execution. A conclusion built from repeated evidence describes the model comparison more reliably.

    Practical-equivalence transitions remain useful. They show where the decision sits close to the configured boundary. A product team can then bring in latency, cost, operational fit, or a targeted follow-up evaluation before committing.

    5. Using this evidence responsibly

    The result covers one model pair, one independent judge, four controlled scenario workspaces, and a bounded descriptive benchmark. It supports a precise claim about the observed evaluations. Broader model families, domains, and confirmatory designs can extend the evidence.

    The public benchmark package includes the frozen scenario identifiers, scenario-level instability flags, phase and cumulative results, calculation code, validation checks, tests, and chart sources. It excludes private evaluation inputs and proprietary production internals.

    Model selection deserves evidence that survives a fresh trial. Plumloom turns repeated evaluation into a decision record that teams can inspect, reproduce, and trust.

    Measure Your AI Before You Launch It

    See where your model holds up and where it breaks, in minutes.

    No credit card required