Google's Gemma-2 9b and OpenAI's GPT-4o performed poorly on a Stanford team's descriptive and normative bias benchmarks despite scoring near-perfectly on Anthropic's DiscrimEval.
Notes on verification
Directly confirmed by the original Stanford arXiv paper and Stanford HAI news summary, which explicitly state that Gemma-2 9b and GPT-4o scored near-perfectly on existing fairness benchmarks like DiscrimEval but poorly on the new DiffAware/CtxtAware benchmarks. [tier=silver indep_score=1.0 clusters=2 claim_tier=notable]
Sources
- These new AI benchmarks could help make models less biased (seed:technology_and_ai)
- https://arxiv.org/abs/2502.01926 (corroboration)
- https://arxiv.org/html/2502.01926v2 (corroboration)
- https://hai.stanford.edu/news/ais-fairness-problem-when-treating-everyone-the-same-is-the-wrong-approach (corroboration)
- https://law.stanford.edu/stanford-lawyer/articles/ai-fairness-through-difference-awareness/ (corroboration)