Sber has updated its GigaChat Audio model, enhancing its ability to process audio files directly without full transcription and to detect emotional tones in speech. The updated model can now identify speakers, summarize content, and provide timestamps for specific moments within recordings up to three hours long. Additionally, Sber has released an open-source version, GigaChat3.1-Audio-10B, and the GigaAM Multilingual speech recognition family, supporting multiple languages. AI
IMPACT This release enhances audio processing capabilities for developers, potentially improving applications like meeting summarization and call analysis.
RANK_REASON The cluster describes a new model release and open-source availability from a major AI lab (Sber). [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Arena Hard Audio
- arXiv
- CNews
- GigaAM Multilingual
- GigaChat3.1-Audio-10B
- GigaChat Audio
- OpenRouter
- provod.ai
- Sber
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →