๐Ÿ“‹Stalecollected in 2m

Sesame debuts iOS app with 4 personal voice agents

Sesame debuts iOS app with 4 personal voice agents
PostLinkedIn
๐Ÿ“‹Read original on TestingCatalog

๐Ÿ’กExplore how Sesame implements persona-driven voice agents to improve user engagement in mobile AI applications.

โšก 30-Second TL;DR

What Changed

Launched iOS preview version across 39 countries.

Why It Matters

This release highlights the growing trend of specialized, persona-driven AI agents in the consumer space. It challenges existing voice assistants by offering more lifelike, multi-modal interaction capabilities.

What To Do Next

Download the Sesame iOS app to analyze their prompt engineering and conversational flow for building persona-based AI agents.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขLaunched iOS preview version across 39 countries.
  • โ€ขFeatures four distinct AI agents: Maya, Miles, Simone, and Charlie.
  • โ€ขIncludes integrated search cards, note-taking, and privacy-focused incognito mode.

๐Ÿง  Deep Insight

Web-grounded analysis with 17 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขSesame's agents are powered by a Conversational Speech Model (CSM) that processes both text and audio simultaneously, enabling natural conversational flow with ultra-low latency responses, typically within 200-300 milliseconds.
  • โ€ขThe company's core focus is on developing emotionally intelligent voice companions capable of detecting and responding to user emotions, managing natural dialogue flow, and maintaining consistent personalities, aiming for more human-like interactions than traditional AI assistants.
  • โ€ขThe iOS preview, currently available for free in 39 countries with a potential waitlist, is a foundational step towards Sesame's broader roadmap, which includes future Android support and the integration of these AI agents into intelligent eyewear by 2027.
  • โ€ขEach of the four distinct agents โ€“ Maya, Miles, Simone, and Charlie โ€“ is designed with a unique personality, point of view, and individual memory, allowing for personalized and evolving conversational experiences that adapt over time.
๐Ÿ“Š Competitor Analysisโ–ธ Show
Feature / ProductSesame AI (iOS App)ChatGPT (Voice Mode)Google Gemini (Voice)Apple Siri (Apple Intelligence)Microsoft Copilot (Voice)
Core FocusEmotionally intelligent, natural voice companions for daily conversation and thought partnership.General-purpose drafting, brainstorming, Q&A, research.General assistance, Workspace integration, multi-modal.On-device tasks, system integration, privacy-sensitive.Productivity assistant, Microsoft 365 integration.
Voice InteractionUltra-low latency, emotionally intelligent, natural conversational dynamics, interruptible.Natural, interruptible, low-latency.Two-way conversation.Voice recognition, basic commands.Voice-enabled productivity.
Memory/ContextComprehensive, individualized agent memory; context-aware.Custom GPTs, growing connector ecosystem.Remembers user preferences, adapts to context.Adapts to user language, searches, preferences over time.Remembers user preferences, adapts to user context.
Key FeaturesReal-time search cards, note-taking, incognito mode, text mode.Web browsing, image input, custom GPTs.Web search, scheduling, drafting, smart home control.Setting reminders, sending messages, photo cleanup, writing tools, smart replies.Draft emails, summarize meetings, generate reports, enterprise data integration.
PricingFree during preview phase.Free tier, ChatGPT Plus ($20/month).Free tier, Google AI Pro ($19.99/month), Google AI Ultra ($249.99/month).Free with compatible Apple devices.Free with eligible Microsoft 365 subscription; Microsoft 365 Business Standard ($33.50/month).
AvailabilityiOS (39 countries), Android preview coming.iOS, Android, web, desktop.iOS, Android, web, Google ecosystem.iOS, macOS, Apple Watch.Web, Microsoft 365 apps, Windows.

๐Ÿ› ๏ธ Technical Deep Dive

  • Conversational Speech Model (CSM): Sesame's core technology is a multimodal, end-to-end learning model that simultaneously processes both text and audio inputs to generate speech.
  • Architecture Foundation: The CSM builds upon a Llama-based architecture, which serves as the foundation for its language processing capabilities.
  • Audio Processing: It utilizes the Mimi speech encoder, a split-Residual Vector Quantizer (RVQ) tokenizer, to convert continuous audio waveforms into discrete "latent" tokens.
  • Multimodal Input: Text and audio tokens are interleaved and fed sequentially into a multimodal backbone transformer, which predicts the zeroth level of the codebook.
  • Speech Generation: A smaller audio decoder, featuring a distinct linear head for each codebook, then models the remaining N-1 codebooks to reconstruct speech from the backbone's representations, facilitating low-latency generation.
  • Low Latency & Expressivity: The single-stage model design enhances efficiency and expressivity, enabling response times of 200-300 milliseconds.
  • Contextual Awareness: The model leverages the history of the conversation to produce more natural and coherent speech, adapting its tone and style to match the situation.
  • Real-time Search: Sesame agents can execute multiple parallel searches while speaking, seamlessly integrating relevant results into their responses and even pivoting mid-sentence if necessary.
  • Open-sourced Component: A 1B variant of the CSM model was open-sourced in March 2025, designed for efficient operation on consumer-grade hardware with a CUDA-compatible GPU.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Sesame AI will accelerate the adoption of intelligent eyewear as a primary interface for AI interaction.
The company explicitly states its roadmap includes intelligent eyewear coming in 2027, leveraging its voice-first, low-latency agents for hands-free interaction.
The emphasis on emotionally intelligent and personalized AI agents will set a new standard for user expectations in conversational AI.
Sesame's focus on detecting and responding to emotions, maintaining consistent personalities, and offering individualized agent experiences aims to create a more human-like and engaging interaction, pushing beyond transactional assistants.
Sesame's approach to balancing quick responses with thoughtful, contextually rich information will influence future AI assistant design.
By optimizing its stack for ultra-low latency and training agents to understand conversational flow while simultaneously performing parallel searches, Sesame addresses a core challenge in making AI conversations feel natural and informative.

โณ Timeline

2022
Sesame AI founded by Brendan Iribe, Ankit Kumar, and Ryan Brown.
2023-10
Secured Seed Round funding of $10.1 million.
2025-02
Released a research demo of its voice technology (Maya and Miles) and announced its vision for an AI voice companion paired with smart glasses.
2025-03
Open-sourced its Conversational Speech Model (CSM).
2026-05
Launched iOS preview version of its app with four personal voice agents (Maya, Miles, Simone, Charlie) in 39 countries.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: TestingCatalog โ†—