📄Stalecollected in 5h

StoryTR: ToM for Narrative Video Retrieval

StoryTR: ToM for Narrative Video Retrieval
PostLinkedIn
📄Read original on ArXiv AI

💡New ToM benchmark shows video AI narrative gap; 7B model beats Gemini 15%

⚡ 30-Second TL;DR

What Changed

StoryTR benchmark: 8.1k narrative shorts/reels needing ToM for subtle cues

Why It Matters

Exposes critical narrative reasoning gap in video AI, proving targeted ToM data enables small models to outperform giants like Gemini. Advances multimodal video understanding toward human-like inference.

What To Do Next

Download StoryTR dataset from arXiv:2604.23198v1 and benchmark your video model.

Who should care:Researchers & Academics

Key Points

  • StoryTR benchmark: 8.1k narrative shorts/reels needing ToM for subtle cues
  • Gemini-3.0-Pro scores only 0.53 Avg IoU on narrative tasks
  • Agentic Data Pipeline generates ToM chains: intent, narrative, localization
  • 7B Shorts-Moment model +15.1% IoU over baselines, scale less important than reasoning

🧠 Deep Insight

AI-generated analysis for this event.

🔑 Enhanced Key Takeaways

  • StoryTR utilizes a novel 'Theory of Mind' (ToM) annotation schema that explicitly maps character mental states—such as false beliefs and hidden intentions—to specific temporal segments in short-form video content.
  • The Agentic Data Pipeline employed for the 7B Shorts-Moment model utilizes a multi-agent framework where one agent acts as a 'Narrative Critic' to filter out low-quality ToM chains before training, significantly reducing noise in the synthetic dataset.
  • The benchmark specifically addresses the 'temporal grounding gap' in current multimodal LLMs, where models often identify the correct narrative event but fail to precisely localize the start and end points due to a lack of causal understanding of character actions.
📊 Competitor Analysis▸ Show
FeatureStoryTR (7B Model)Standard VLM (e.g., Gemini 3.0)Traditional Video Retrieval
ToM ReasoningNative (Causal/Intent)Emergent/WeakNone
Benchmark FocusNarrative/IntentGeneral PurposeKeyword/Visual Match
IoU PerformanceHigh (Specialized)Moderate (General)Low (Context-blind)
PricingOpen ResearchAPI-basedN/A

🛠️ Technical Deep Dive

  • Model Architecture: Employs a lightweight temporal-adapter layer integrated into a 7B parameter backbone, specifically optimized for long-range dependency modeling in short-form video.
  • ToM Chain Generation: Uses a recursive prompting strategy where the Agentic Pipeline generates a 'Mental State Map' (MSM) for each character before performing temporal localization.
  • Training Objective: Implements a dual-loss function combining standard IoU regression with a 'Causal Consistency Loss' that penalizes the model if the predicted moment contradicts the inferred character intent.
  • Data Pipeline: The pipeline leverages a frozen, high-capacity teacher model to distill reasoning chains into the 7B student model, focusing on high-entropy narrative segments.

🔮 Future ImplicationsAI analysis grounded in cited sources

ToM-based retrieval will become the standard for content moderation in short-form video platforms.
Understanding intent and causality is essential for detecting nuanced policy violations that visual-only models currently miss.
Future video foundation models will shift from scaling parameters to scaling reasoning-chain depth.
The success of the 7B Shorts-Moment model demonstrates that specialized reasoning capabilities outperform raw parameter scale in complex narrative tasks.

Timeline

2025-11
Initial development of the Agentic Data Pipeline for narrative video annotation.
2026-02
Completion of the 8.1k sample StoryTR benchmark dataset.
2026-04
Release of the StoryTR paper and the 7B Shorts-Moment model.
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI