Benchmark Exposes Voice Agents’ Hidden Instruction Gaps

💡See why voice agents follow personas in content but fail to adapt when managing the conversational floor.
⚡ 30-Second TL;DR
What Changed
The benchmark covers five conditioning protocols: default, explicit rules, persona-only, combined persona-and-rules, and instruction conflict.
Why It Matters
The findings show that persona prompting, conversational timing, and safety-oriented instruction resolution are separate engineering problems rather than one general instruction-following capability. Teams building real-time voice agents should test inferred behavior explicitly instead of relying on role descriptions alone.
What To Do Next
Run your full-duplex voice agent through DuplexSpeechBench-IFEval-style tests for persona-only, explicit-rule, and safety-conflict conditions before deployment.
Key Points
- •The benchmark covers five conditioning protocols: default, explicit rules, persona-only, combined persona-and-rules, and instruction conflict.
- •It uses Instruction Adherence Score for deterministic floor management and Persona Adherence Score for LLM-judged content consistency.
- •F-Actor and PersonaPlex show 9.7% and 4.5% adherence drops, respectively, under persona-only conditioning.
- •GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat maintain persona-consistent content but show limited adaptation in floor behavior and proactive actions.
- •Systems can follow conflicting directives toward a prescribed persona yet still fail to override them during safety conflicts.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: ArXiv AI ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.