OpenAI researchers Yuri Burda and Harri Edwards accidentally discovered the 'grokking' phenomenon after arithmetic-training experiments ran for days instead of hours, finding that models which initially memorized example sums eventually learned to add previously unseen numbers.
Notes on verification
Confirmed by original arXiv paper (Power, Burda, Edwards et al.), Wikipedia, and Quanta Magazine reporting; all key details—accidental long training run, arithmetic task, memorization-to-generalization pattern—corroborate consistently. [cascade flags: wikipedia_dropped_other_sources_exist | tier=silver indep_score=0.925 clusters=2 claim_tier=notable]
Sources
- Large language models can do jaw-dropping things. But nobody ... (seed:technology_and_ai)
- https://www.quantamagazine.org/how-do-machines-grok-data-20240412/ (corroboration)
- https://arxiv.org/abs/2201.02177 (corroboration)