Gemma 3n
PulseAugur coverage of Gemma 3n — every cluster mentioning Gemma 3n across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
PhaseCoder enables LLMs to understand spatial audio, regardless of microphone setup
Researchers have developed PhaseCoder, a novel transformer-based encoder designed to process raw multichannel audio and microphone coordinates, enabling multimodal LLMs to understand spatial audio information. Unlike pr…
-
Google's Eloquent dictation app fails benchmarks due to frequent transcription errors
A user attempting to benchmark Google's new on-device dictation app, Eloquent, found it to be largely unusable due to frequent failures to transcribe audio. While the app's accuracy was competitive when it did function,…
-
Gemma 3n fully available in the open-source ecosystem!
Google DeepMind has fully released Gemma 3n, a mobile-first multimodal model designed for on-device applications. This new architecture supports image, audio, video, and text inputs, with text outputs, and is optimized …