ai-sage has released GigaChat3.1-Audio-10B-A1.8B, an audio-native large language model built upon the GigaChat 3.1 Lightning text model. This new model integrates a Conformer speech encoder and a modality adapter, allowing it to process audio embeddings directly into its Mixture-of-Experts decoder. This architecture enables the model to maintain the text quality of its base while gaining speech understanding capabilities, including audio question answering, classification, temporal grounding, and tool-use. AI
IMPACT Introduces an audio-native LLM capable of complex temporal grounding and tool use, potentially advancing multimodal AI applications.
RANK_REASON Release of a new open-source model with a paper and dataset. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →