
StateSight Exposes VLM Spatial Reasoning Gaps
StateSight is a procedurally generated benchmark for testing latent spatial-state reconstruction in vision-language models through cube reasoning, occluded counting, and connected-component tasks. GPT-5.5 and Claude Sonnet 5 substantially lagged behind human participants, despite producing perfectly formatted answers.
ArXiv AI · 26d ago














