FrontierMath is a math benchmark whose difficult questions remain unpublished so that AI companies cannot train their models against them.
Notes on verification
Confirmed by original arXiv paper (2411.04872), Epoch AI's own benchmark descriptions, and secondary academic sources discussing contamination-prevention protocols. [tier=silver indep_score=0.925 clusters=2 claim_tier=notable]
Sources
- New secret math benchmark stumps AI models and PhDs alike (seed:technology_and_ai)
- https://arxiv.org/abs/2411.04872 (corroboration)
- https://epoch.ai/frontiermath (corroboration)
- https://epoch.ai/benchmarks/frontiermath-tiers-1-3-v2 (corroboration)