Technology2020sglobalhigh confidence

Google's Gemma-2 9b and OpenAI's GPT-4o performed poorly on a Stanford team's descriptive and normative bias benchmarks despite scoring near-perfectly on Anthropic's DiscrimEval.

Notes on verification

Directly confirmed by the original Stanford arXiv paper and Stanford HAI news summary, which explicitly state that Gemma-2 9b and GPT-4o scored near-perfectly on existing fairness benchmarks like DiscrimEval but poorly on the new DiffAware/CtxtAware benchmarks. [tier=silver indep_score=1.0 clusters=2 claim_tier=notable]

Sources