PulseAugur
EN
LIVE 19:49:14

Apple unveils memory-efficient on-device audio synthesis for Siri

Apple has developed a new memory-efficient architecture for on-device audio synthesis, detailed in a research paper. This system, powering Siri Expressive Voices, uses a Diffusion Transformer (DiT)-style decoder to convert semantic audio tokens into high-fidelity speech with minimal computational resources. The architecture achieves real-time synthesis speeds, requires only about 21MB of runtime memory, and significantly improves audio quality metrics compared to previous on-device systems. AI

IMPACT Enables more sophisticated and efficient on-device voice capabilities, potentially improving user experience across Apple products.

RANK_REASON Research paper detailing a novel architecture for audio synthesis.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Apple unveils memory-efficient on-device audio synthesis for Siri

COVERAGE [2]

  1. Apple Machine Learning Research TIER_1 English(EN) ·

    Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

    Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple’s most powerful on-device foundation model. This work presents the memory-efficient audio synthesis architecture behind that capability: a detoke…

  2. arXiv cs.CL TIER_1 English(EN) · Dongseong Hwang, Prasanth Yadla, Kaan Elgin, Shifas Padinjaru Veettil, Sivanand Achanta, Dipjyoti Paul, Ramya Rasipuram, Tyler Johnson, Emad Soroush, Chung-Cheng Chiu, Zhifeng Chen ·

    Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

    arXiv:2607.23811v1 Announce Type: cross Abstract: Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple's most powerful on-device foundation model. This work presents the memory-efficient audio synth…