When GPT-4 was tested in 2023 on easy Codeforces problems, Narayanan and Kapoor (2023b) found that it could regularly solve problems added before 5 September 2021 but got no question right among problems added afterward.
Notes on verification
Confirmed by the original Narayanan/Kapoor blog post and corroborated by a later Princeton CRCL paper and a third independent benchmark paper. Minor discrepancy in exact date wording (Sept 5 vs Sept 12 cutoff) does not undermine the core finding of a sharp performance drop-off tied to training data cutoff. [tier=silver indep_score=0.833 clusters=3 claim_tier=notable]
Sources
- Can We Trust AI Benchmarks? An Interdisciplinary Review ... (seed:technology_and_ai)
- https://www.normaltech.ai/p/gpt-4-and-professional-benchmarks (corroboration)
- https://shana.codes/posts/week-notes-4-03-23.html (corroboration)
- https://www.cs.princeton.edu/~sayashk/papers/crcl-kapoor-henderson-narayanan.pdf (corroboration)