來源較早收集於 2m

Sesame 推出 iOS 應用程式,內建 4 款個人語音助理

Sesame 推出 iOS 應用程式,內建 4 款個人語音助理
PostLinkedIn
📋閱讀原文: TestingCatalog
#voice-ai#ios-app#personal-assistantsesamesesame

💡探索 Sesame 如何透過人格化語音助理來提升行動 AI 應用程式的使用者參與度。

⚡ 30 秒速覽

有什麼變化

於 39 個國家推出 iOS 預覽版。

為什麼重要

此發布凸顯了消費市場中,專用且具備人格特質的 AI 助理趨勢。它透過提供更擬真、多模態的互動能力,對現有的語音助理市場構成挑戰。

下一步行動

下載 Sesame iOS 應用程式,分析其提示工程 (prompt engineering) 與對話流程,以利開發具備人格特質的 AI 助理。

誰應關注:Developers & AI Engineers

關鍵要點

  • 於 39 個國家推出 iOS 預覽版。
  • 包含四款不同的 AI 助理:Maya、Miles、Simone 與 Charlie。
  • 整合搜尋卡片、筆記功能及隱私導向的隱身模式。

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 17 個來源。

🔑 增強重點摘要

  • Sesame's agents are powered by a Conversational Speech Model (CSM) that processes both text and audio simultaneously, enabling natural conversational flow with ultra-low latency responses, typically within 200-300 milliseconds.
  • The company's core focus is on developing emotionally intelligent voice companions capable of detecting and responding to user emotions, managing natural dialogue flow, and maintaining consistent personalities, aiming for more human-like interactions than traditional AI assistants.
  • The iOS preview, currently available for free in 39 countries with a potential waitlist, is a foundational step towards Sesame's broader roadmap, which includes future Android support and the integration of these AI agents into intelligent eyewear by 2027.
  • Each of the four distinct agents – Maya, Miles, Simone, and Charlie – is designed with a unique personality, point of view, and individual memory, allowing for personalized and evolving conversational experiences that adapt over time.
📊 競品分析▸ Show
Feature / ProductSesame AI (iOS App)ChatGPT (Voice Mode)Google Gemini (Voice)Apple Siri (Apple Intelligence)Microsoft Copilot (Voice)
Core FocusEmotionally intelligent, natural voice companions for daily conversation and thought partnership.General-purpose drafting, brainstorming, Q&A, research.General assistance, Workspace integration, multi-modal.On-device tasks, system integration, privacy-sensitive.Productivity assistant, Microsoft 365 integration.
Voice InteractionUltra-low latency, emotionally intelligent, natural conversational dynamics, interruptible.Natural, interruptible, low-latency.Two-way conversation.Voice recognition, basic commands.Voice-enabled productivity.
Memory/ContextComprehensive, individualized agent memory; context-aware.Custom GPTs, growing connector ecosystem.Remembers user preferences, adapts to context.Adapts to user language, searches, preferences over time.Remembers user preferences, adapts to user context.
Key FeaturesReal-time search cards, note-taking, incognito mode, text mode.Web browsing, image input, custom GPTs.Web search, scheduling, drafting, smart home control.Setting reminders, sending messages, photo cleanup, writing tools, smart replies.Draft emails, summarize meetings, generate reports, enterprise data integration.
PricingFree during preview phase.Free tier, ChatGPT Plus ($20/month).Free tier, Google AI Pro ($19.99/month), Google AI Ultra ($249.99/month).Free with compatible Apple devices.Free with eligible Microsoft 365 subscription; Microsoft 365 Business Standard ($33.50/month).
AvailabilityiOS (39 countries), Android preview coming.iOS, Android, web, desktop.iOS, Android, web, Google ecosystem.iOS, macOS, Apple Watch.Web, Microsoft 365 apps, Windows.

🛠️ 技術深入

  • Conversational Speech Model (CSM): Sesame's core technology is a multimodal, end-to-end learning model that simultaneously processes both text and audio inputs to generate speech.
  • Architecture Foundation: The CSM builds upon a Llama-based architecture, which serves as the foundation for its language processing capabilities.
  • Audio Processing: It utilizes the Mimi speech encoder, a split-Residual Vector Quantizer (RVQ) tokenizer, to convert continuous audio waveforms into discrete "latent" tokens.
  • Multimodal Input: Text and audio tokens are interleaved and fed sequentially into a multimodal backbone transformer, which predicts the zeroth level of the codebook.
  • Speech Generation: A smaller audio decoder, featuring a distinct linear head for each codebook, then models the remaining N-1 codebooks to reconstruct speech from the backbone's representations, facilitating low-latency generation.
  • Low Latency & Expressivity: The single-stage model design enhances efficiency and expressivity, enabling response times of 200-300 milliseconds.
  • Contextual Awareness: The model leverages the history of the conversation to produce more natural and coherent speech, adapting its tone and style to match the situation.
  • Real-time Search: Sesame agents can execute multiple parallel searches while speaking, seamlessly integrating relevant results into their responses and even pivoting mid-sentence if necessary.
  • Open-sourced Component: A 1B variant of the CSM model was open-sourced in March 2025, designed for efficient operation on consumer-grade hardware with a CUDA-compatible GPU.

🔮 前景展望基於引用來源的 AI 分析

Sesame AI will accelerate the adoption of intelligent eyewear as a primary interface for AI interaction.
The company explicitly states its roadmap includes intelligent eyewear coming in 2027, leveraging its voice-first, low-latency agents for hands-free interaction.
The emphasis on emotionally intelligent and personalized AI agents will set a new standard for user expectations in conversational AI.
Sesame's focus on detecting and responding to emotions, maintaining consistent personalities, and offering individualized agent experiences aims to create a more human-like and engaging interaction, pushing beyond transactional assistants.
Sesame's approach to balancing quick responses with thoughtful, contextually rich information will influence future AI assistant design.
By optimizing its stack for ultra-low latency and training agents to understand conversational flow while simultaneously performing parallel searches, Sesame addresses a core challenge in making AI conversations feel natural and informative.

時間線

2022
Sesame AI founded by Brendan Iribe, Ankit Kumar, and Ryan Brown.
2023-10
Secured Seed Round funding of $10.1 million.
2025-02
Released a research demo of its voice technology (Maya and Miles) and announced its vision for an AI voice companion paired with smart glasses.
2025-03
Open-sourced its Conversational Speech Model (CSM).
2026-05
Launched iOS preview version of its app with four personal voice agents (Maya, Miles, Simone, Charlie) in 39 countries.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: TestingCatalog

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週電子報

每週一封,可隨時退訂。