RTCA Workshop Opens NeurIPS 2026 Submissions
๐กFind benchmarks and research directions for making voice and embodied agents feel genuinely real-time.
โก 30-Second TL;DR
What Changed
Submission tracks include full papers up to 8 pages, short papers up to 4 pages, and demo papers.
Why It Matters
The workshop could help establish shared benchmarks and terminology for evaluating conversational agents in live interaction rather than only on offline metrics. It is particularly relevant to teams building voice agents, embodied avatars, or streaming multimodal systems.
What To Do Next
Review the RTCA CFP and submit a full, short, or demo paper through OpenReview before August 29, 2026 AoE.
Key Points
- โขSubmission tracks include full papers up to 8 pages, short papers up to 4 pages, and demo papers.
- โขTopics cover streaming speech and language models, full-duplex agents, avatars, turn-taking, backchannels, and multimodal alignment.
- โขThe workshop welcomes position papers, evaluation critiques, reproducibility studies, and live-system benchmarks.
- โขSubmissions are double-blind, non-archival, single-round review, with no rebuttal; the workshop takes place in Sydney on December 11 or 12, 2026.
๐ง Deep Insight
AI-generated analysis for this event.
๐ Enhanced Key Takeaways
- โขThe RTCA workshop series has evolved from previous NeurIPS workshops focusing on 'Conversational AI' and 'Spoken Dialogue Systems,' specifically shifting focus toward the sub-100ms latency requirements of modern full-duplex LLM agents.
- โขThe workshop organizers have emphasized a 'live-system' requirement, encouraging submissions that include latency-throughput trade-off analysis on edge hardware rather than just cloud-based API performance.
- โขNeurIPS 2026 is being hosted in Sydney, marking a significant return to the Asia-Pacific region for the conference, which influences the workshop's focus on global accessibility and multilingual conversational benchmarks.
- โขThe workshop is explicitly soliciting 'failure analysis' papers, a departure from traditional academic venues that often prioritize state-of-the-art performance metrics over negative results in conversational agent stability.
- โขRTCA 2026 is collaborating with the 'Open-Source Conversational AI Initiative' to provide a standardized evaluation framework for participants to test their models against a common set of interactional datasets.
๐ ๏ธ Technical Deep Dive
- Focus on minimizing Time-To-First-Token (TTFT) in streaming architectures through speculative decoding and KV-cache quantization.
- Emphasis on multimodal alignment techniques, specifically integrating audio-visual features directly into the latent space of the LLM to handle non-verbal cues like backchanneling.
- Exploration of 'interruptibility' mechanisms, where the model must maintain a state-aware buffer to handle user barge-in events without losing context.
- Implementation of adaptive turn-taking algorithms that utilize prosodic features (pitch, energy, and pause duration) rather than simple silence-threshold detection.
๐ฎ Future ImplicationsAI analysis grounded in cited sources
โณ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events โ
๐Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning โ