PulseAugur
EN
LIVE 14:15:30

GigaChat Audio 10B model offers time-aware understanding for 2-hour audio

Researchers have developed GigaChat Audio, a novel time-aware large audio language model capable of understanding and answering questions about audio recordings up to two hours long. This model addresses the challenge of temporal grounding in long audio by interleaving periodic time markers with audio tokens. It achieves significant accuracy in temporal localization, outperforming other open-source models, and supports various audio tasks including ASR, translation, and QA. The model weights and datasets are being released to foster further research in time-aware audio understanding. AI

IMPACT This model advances temporal grounding in audio LLMs, potentially improving applications requiring analysis of long audio recordings.

RANK_REASON The item describes a new research paper and model release for audio understanding. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

GigaChat Audio 10B model offers time-aware understanding for 2-hour audio

COVERAGE [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    GigaChat Audio: Time-aware Large Audio Language Model

    Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens usi…