AI models evaluated against FrontierMath include Anthropic's Claude 3.5.
Notes on verification
Confirmed by Epoch AI's official FrontierMath documentation and the original arXiv paper, both independent primary sources listing Claude 3.5 Sonnet among evaluated models. [tier=silver indep_score=0.925 clusters=2 claim_tier=everyday]
Sources
- New secret math benchmark stumps AI models and PhDs alike (seed:technology_and_ai)
- https://epoch.ai/frontiermath/tiers-1-4/the-benchmark (corroboration)
- https://arxiv.org/pdf/2411.04872 (corroboration)