GPT-6 Wins a Multimodal Drawing Test by a Nose

💡A 1,350-image test reveals where multimodal models still fail: unusual relations, slang, and even recognizing their own
⚡ 30-Second TL;DR
What Changed
The benchmark tested eight multimodal models on 150 terms spanning common nouns, actions, memes, AI jargon, and counterintuitive combinations.
Why It Matters
The results suggest that multimodal capability is not captured by standard image-generation or VQA scores alone: models must encode unusual relations clearly and resist default semantic assumptions. The benchmark also highlights a practical risk for agentic systems that depend on visual communication, especially when instructions involve slang, role reversals, or abstract concepts.
What To Do Next
Add counterintuitive, slang, and role-reversal prompts to your multimodal model evaluation set, then score both image generation and independent visual recognition.
Key Points
- •The benchmark tested eight multimodal models on 150 terms spanning common nouns, actions, memes, AI jargon, and counterintuitive combinations.
- •GPT-6 Astra scored 75.5, followed by Gemini 3.8 Flash at 74.6 and Claude Fable 5.1 at 74.3; their confidence intervals overlapped.
- •Models were evaluated equally on drawing clarity and guessing accuracy, with answers automatically matched against predefined terms and synonyms.
- •Counterintuitive prompts produced severe failures: none of the counted guesses identified the concept “cat shoveling for a human.”
- •Some models failed to recognize their own drawings; MiniMax and DeepSeek had self-recognition rates below the rates achieved by other models.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅 ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.



