research / arXiv:2607.21735 · 30 July 2026
What AI Red-Team Evaluations Can and Cannot Prove
Abstract
Introduces the evidential ceiling, the largest factor by which one result can move belief under a fixed testing budget, and derives it in closed form. For frequent harms, modest benchmarks can certify safety at stated standards, and negative results carry more information than reproduced failures. For rare catastrophic harms, feasible benchmarks fall several orders of magnitude short. The analysis extends to adaptive and automated red-teaming, and an audit of eight evaluation suites finds them calibrated for frequent harm categories but not for low-probability, high-impact ones.
discrimination between hypotheses, not attack success, sets evidential worth
Key findings
- The evidential ceiling has a closed form: the most one red-team result can move belief under a fixed testing budget.
- For frequent harms, modest benchmarks can certify safety at stated standards.
- Above the crossing, a clean sheet carries more information than one reproduced failure.
- For rare, catastrophic harms, feasible passive benchmarks fall several orders of magnitude short.
- The result extends to adaptive and automated red-teaming.
- An audit of eight evaluation suites finds them calibrated for frequent harm categories, not for low-probability, high-impact ones.
Read and cite
arXiv·PDF·doi:10.48550/arXiv.2607.21735
@misc{kaur2026evidential,
title = {What {AI} Red-Team Evaluations Can and Cannot Prove},
author = {Kaur, Bandana},
year = {2026},
eprint = {2607.21735},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2607.21735},
url = {https://arxiv.org/abs/2607.21735}
}