Plumloom Research
Evidence for evaluations
you can trust
Studies, benchmarks, and open datasets from the Plumloom team. Every report ships with its data, code, and methodology so you can reproduce it end-to-end.
02
Studies published
7+
Open datasets
100%
Reproducible
Featured Study
Latest from the lab
Aug 18, 2026
Headline result
16/16
Cross-judge paired comparisons separated honest answers from fabricated twins.
RAG EvaluationCorpus BoundaryJudge Replication
Plumloom RAG Evaluation Benchmark
Can a document agent prove it answered from the document? A 20-case corpus-boundary benchmark, matched honest/fabricated controls, second-judge replication, and multi-run Scenario evidence for RAG release decisions.
19 of 20
Expected outcome class observed
5 of 5
Expected refusals observed
16 of 16
Cross-judge paired comparisons
0.003 / 0.068
Scenario replica agreement
Data Code Methodology
Read studyJul 15, 2026
Headline result
72.73%
Relative reduction in model-winner instability with repeated evaluation.
Evaluation ReliabilityModel SelectionRepeated Evaluation
Evaluation Reliability Benchmark
Testing whether an evaluation system responds correctly to score variability, and whether repeated evaluation produces more stable model choices.
38 of 38
Correct reliability decisions
24.44%
Single-evaluation instability
6.67%
Repeated-evaluation instability
72.73%
Relative reduction
Data Code Methodology
Read studyRun reliable evals on your own model
Every result on this page came out of Plumloom. Try the same convergence-checked pipeline on your prompts.
No credit card required