According to arXiv paper 2411.12990, most evaluated AI benchmarks do not allow their results to be easily replicated.
Notes on verification
Directly confirmed by the paper's abstract and full text, with corroboration from NeurIPS proceedings, OpenReview, and NeurIPS poster page. Specific data (17/24 benchmarks lack easy replication scripts) supports the claim. [tier=unverified indep_score=0.3 clusters=2 claim_tier=notable] [rescored 2026-09-15: curated origin-host map (PR #65); unclassified hosts no longer scored as aggregators]
Sources
- BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices (seed:technology_and_ai)
- https://neurips.cc/virtual/2024/poster/97566 (corroboration)
- https://openreview.net/forum?id=hcOq2buakM (corroboration)
- https://proceedings.neurips.cc/paper_files/paper/2024/hash/26889e8359e7ef8a7f5d77457364ca55-Abstract-Datasets_and_Benchmarks_Track.html (corroboration)