CRYSTAL Benchmark Reveals VLM Reasoning Gaps
CRYSTAL is a new benchmark with 6,372 visual questions featuring verified step-by-step reasoning to evaluate multimodal models beyond final answers. Tests on 20 models show high accuracy but poor reasoning recovery, like GPT-5 at 58% accuracy vs 48% steps. CPR Curriculum training boosts reasoning by up to 93% on select VLMs.




