JPEG compression levels used in API calls can change the ranking lineup of models on visually prompted benchmarks.
Notes on verification
Confirmed by the originating arXiv paper and its associated GitHub repo and project page, which is strong primary-source support. However, this is a single research group's recent (2025) finding not yet independently replicated or peer-reviewed, so corroboration outside the authors' own materials is limited. [tier=unverified indep_score=0.3 clusters=3 claim_tier=notable] [rescored 2026-09-15: curated origin-host map (PR #65); unclassified hosts no longer scored as aggregators]
Sources
- Visually Prompted Benchmarks Are Surprisingly Fragile (seed:technology_and_ai)
- https://www.researchgate.net/publication/398937366_Visually_Prompted_Benchmarks_Are_Surprisingly_Fragile (corroboration)
- https://lisadunlap.github.io/vpbench/ (corroboration)
- https://github.com/TonyLianLong/VPBench (corroboration)