Science2024globallow confidence

According to arXiv paper 2411.12990, most evaluated AI benchmarks do not allow their results to be easily replicated.

Notes on verification

Directly confirmed by the paper's abstract and full text, with corroboration from NeurIPS proceedings, OpenReview, and NeurIPS poster page. Specific data (17/24 benchmarks lack easy replication scripts) supports the claim. [tier=unverified indep_score=0.3 clusters=2 claim_tier=notable] [rescored 2026-09-15: curated origin-host map (PR #65); unclassified hosts no longer scored as aggregators]

Sources

According to arXiv paper 2411.12990, most evaluated AI be… · DeepInquiry