A study (arXiv 2411.12990) found that commonly used AI benchmarks suffer from significant issues and exhibit large differences in quality.
Notes on verification
Directly confirmed by the arXiv abstract and corroborated by NeurIPS proceedings, poster page, and project website; peer-reviewed as NeurIPS 2024 Spotlight. [tier=silver indep_score=0.875 clusters=2 claim_tier=notable]
Sources
- BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices (seed:technology_and_ai)
- https://proceedings.neurips.cc//paper_files/paper/2024/hash/26889e8359e7ef8a7f5d77457364ca55-Abstract-Datasets_and_Benchmarks_Track.html (corroboration)
- https://betterbench.stanford.edu/ (corroboration)
- https://neurips.cc/virtual/2024/poster/97566 (corroboration)