sanoTTS Brings Tiny TTS to Microcontrollers

π‘A 337 KB TTS model could enable offline voice features on $3 microcontrollers.
β‘ 30-Second TL;DR
What Changed
The smallest model has 294K parameters and occupies 337 KB after INT8 quantization.
Why It Matters
The project could make offline voice interfaces practical for low-cost embedded devices, reducing reliance on cloud TTS APIs. Its small footprint is especially relevant for toys, sensors, wearables, and privacy-sensitive products.
What To Do Next
Install the sanotts-web package and benchmark the 294K INT8 model on your target ESP32 or 512 KB-SRAM device using your own language, latency, and audio-quality tests.
Key Points
- β’The smallest model has 294K parameters and occupies 337 KB after INT8 quantization.
- β’The model family supports 11 voices and 6 languages, with a recipe for extending languages and voices.
- β’The 1.51M-parameter sanoTTS-Amy reports SCOREQ 4.13 and UTMOS 4.10, while ESP32 inference reaches an RTF of 0.225.
- β’Web deployment is available through the sanotts-web npm package and WebAssembly.
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.
