🐯Recentcollected in 17m

GPT-6 Wins a Multimodal Drawing Test by a Nose

GPT-6 Wins a Multimodal Drawing Test by a Nose
PostLinkedIn
🐯Read original on 虎嗅
#svg-generation#visual-reasoning#benchmarkingdraw-and-guess-benchmarkgpt-6 astragemini 3.8 flashclaude fable 5.1minimax-m3deepseek-v4-flash-vision-exp

💡A 1,350-image test reveals where multimodal models still fail: unusual relations, slang, and even recognizing their own

⚡ 30-Second TL;DR

What Changed

The benchmark tested eight multimodal models on 150 terms spanning common nouns, actions, memes, AI jargon, and counterintuitive combinations.

Why It Matters

The results suggest that multimodal capability is not captured by standard image-generation or VQA scores alone: models must encode unusual relations clearly and resist default semantic assumptions. The benchmark also highlights a practical risk for agentic systems that depend on visual communication, especially when instructions involve slang, role reversals, or abstract concepts.

What To Do Next

Add counterintuitive, slang, and role-reversal prompts to your multimodal model evaluation set, then score both image generation and independent visual recognition.

Who should care:Researchers & Academics

Key Points

  • The benchmark tested eight multimodal models on 150 terms spanning common nouns, actions, memes, AI jargon, and counterintuitive combinations.
  • GPT-6 Astra scored 75.5, followed by Gemini 3.8 Flash at 74.6 and Claude Fable 5.1 at 74.3; their confidence intervals overlapped.
  • Models were evaluated equally on drawing clarity and guessing accuracy, with answers automatically matched against predefined terms and synonyms.
  • Counterintuitive prompts produced severe failures: none of the counted guesses identified the concept “cat shoveling for a human.”
  • Some models failed to recognize their own drawings; MiniMax and DeepSeek had self-recognition rates below the rates achieved by other models.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 虎嗅

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.