Researchers created a benchmark called VPBench, a visually prompted benchmark containing 16 visual marker variants.
Notes on verification
Confirmed directly by arXiv abstract, official GitHub repo, HuggingFace dataset page, and project website, all consistently describing VPBench's 16 visual marker variants. [tier=unverified indep_score=0.3 clusters=4 claim_tier=notable] [rescored 2026-09-15: curated origin-host map (PR #65); unclassified hosts no longer scored as aggregators]
Sources
- Visually Prompted Benchmarks Are Surprisingly Fragile (seed:technology_and_ai)
- https://www.researchgate.net/publication/398937366_Visually_Prompted_Benchmarks_Are_Surprisingly_Fragile (corroboration)
- https://lisadunlap.github.io/vpbench/ (corroboration)
- https://github.com/TonyLianLong/VPBench (corroboration)
- https://huggingface.co/datasets/longlian/VPBench (corroboration)