AI models evaluated against FrontierMath include OpenAI's GPT-4o and o1-preview.
Notes on verification
Confirmed by Epoch AI's official FrontierMath benchmark documentation and the arXiv paper describing the benchmark methodology and evaluated models. [tier=silver indep_score=0.925 clusters=2 claim_tier=everyday]
Sources
- New secret math benchmark stumps AI models and PhDs alike (seed:technology_and_ai)
- https://epoch.ai/frontiermath/tiers-1-4/the-benchmark (corroboration)
- https://arxiv.org/pdf/2411.04872 (corroboration)