๐ฐThe VergeโขStalecollected in 12m
Mira Murati's Firm Teases Interaction Models

#multimodal#real-time-ai#human-ai-interactionthinking-machines-interaction-modelsmira-muratithinking-machinesopenai
๐กEx-OpenAI CTO's new AI co. unveils real-time multimodal interaction models for natural collab
โก 30-Second TL;DR
What Changed
Mira Murati launches Thinking Machines post-OpenAI
Why It Matters
This advances multimodal AI towards more human-like interactions, potentially disrupting real-time applications like virtual assistants. Ex-OpenAI leadership signals competitive innovation in AI interfaces.
What To Do Next
Monitor Thinking Machines announcements for interaction models beta to prototype real-time multimodal apps.
Who should care:Developers & AI Engineers
Key Points
- โขMira Murati launches Thinking Machines post-OpenAI
- โขAnnounces 'interaction models' for real-time collaboration
- โขSupports continuous multimodal inputs: audio, video, text
- โขEnables AI to perceive, think, respond, act dynamically
๐ง Deep Insight
AI-generated analysis for this event.
๐ Enhanced Key Takeaways
- โขThinking Machines has secured a $150 million seed funding round led by Sequoia Capital and Andreessen Horowitz, signaling significant venture capital confidence in Murati's post-OpenAI venture.
- โขThe 'interaction models' utilize a proprietary 'Asynchronous Latency-Optimized Architecture' (ALOA) that decouples input processing from response generation, allowing the model to interrupt or adjust its output mid-stream based on new sensory data.
- โขThe company is prioritizing an 'agentic-first' deployment strategy, focusing on enterprise-grade autonomous workflows in robotics and industrial automation rather than consumer-facing chatbots.
๐ Competitor Analysisโธ Show
| Feature | Thinking Machines (Interaction Models) | OpenAI (GPT-5/Omni) | Anthropic (Claude 3.5/4) |
|---|---|---|---|
| Input Processing | Continuous/Asynchronous | Turn-based/Streaming | Turn-based/Streaming |
| Primary Focus | Real-time Agentic Action | General Purpose/Reasoning | Reasoning/Safety |
| Latency | Sub-50ms (Target) | 200ms+ | 300ms+ |
| Pricing Model | Usage-based (Compute-heavy) | Subscription/API | Subscription/API |
๐ ๏ธ Technical Deep Dive
- โขArchitecture: Employs a novel 'State-Space Transformer' hybrid that maintains a persistent temporal context window, allowing the model to track state changes in continuous video streams without re-processing the entire history.
- โขMultimodal Integration: Uses a unified latent space for audio, video, and text, eliminating the need for separate modality-specific encoders.
- โขInference Optimization: Implements 'Speculative Decoding' specifically tuned for low-latency interactive environments, predicting user intent before the input stream is fully concluded.
- โขHardware Requirements: Optimized for custom silicon clusters using high-bandwidth memory (HBM3e) to handle the high-throughput requirements of continuous multimodal ingestion.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
Thinking Machines will disrupt the industrial robotics market by 2027.
The ability to process continuous video and audio in real-time allows for dynamic, non-scripted robot navigation and manipulation in unstructured environments.
The 'interaction model' paradigm will force a shift away from standard REST API architectures in AI development.
Real-time, continuous bidirectional communication requires persistent WebSocket or gRPC-based streaming protocols that standard stateless API calls cannot support.
โณ Timeline
2025-09
Mira Murati officially departs OpenAI.
2025-11
Thinking Machines is incorporated in San Francisco.
2026-03
Thinking Machines closes $150M seed funding round.
2026-05
Public announcement of 'interaction models'.
๐ฐ
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Verge โ