PulseAugur
EN
LIVE 11:49:23

ai-sage releases GigaChat3.1-Audio-10B-A1.8B, an audio-native LLM

ai-sage has released GigaChat3.1-Audio-10B-A1.8B, an audio-native large language model built upon the GigaChat 3.1 Lightning text model. This new model integrates a Conformer speech encoder and a modality adapter, allowing it to process audio embeddings directly into its Mixture-of-Experts decoder. This architecture enables the model to maintain the text quality of its base while gaining speech understanding capabilities, including audio question answering, classification, temporal grounding, and tool-use. AI

IMPACT Introduces an audio-native LLM capable of complex temporal grounding and tool use, potentially advancing multimodal AI applications.

RANK_REASON Release of a new open-source model with a paper and dataset. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

ai-sage releases GigaChat3.1-Audio-10B-A1.8B, an audio-native LLM

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 (TL) · /u/pmttyji ·

    ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v6zksb/aisagegigachat31audio10ba18b_hugging_face/"> <img alt="ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face" src="https://external-preview.redd.it/szDqPyquF34NnfOa0Nf53wiECe5gbQ7yFvH7Ihy93PQ.png?width=64…