
DeepSeek's visual primitives fix MLLM spatial gaps
DeepSeek releases GitHub technical report on multimodal model with 'Thinking with Visual Primitives' framework, turning points/bboxes into core reasoning units to solve spatial referencing issues. Overcomes language ambiguity in complex layouts, enabling precise spatial inference. Compact model matches GPT-4o, Claude-Sonnet, Gemini on tough benchmarks.

