A new 1-billion parameter audio-language model called Sori-1B has been developed by a single researcher from SNU. Unlike typical models, Sori-1B's decoder is trained entirely from scratch on audio-paired text, aiming to ground its responses in audio rather than relying on text-only priors. The model reuses NVIDIA's frozen Audio Flamingo Next encoder and was trained on approximately 7.4k hours of data using three RTX 4090 GPUs. Sori-1B supports various modes including multiple-choice questions, open-ended questions, captioning, and automatic speech recognition, though its weights are restricted to non-commercial and academic use. AI
IMPACT Introduces a novel training methodology for audio-grounded language models, potentially improving their ability to connect language with auditory information.
RANK_REASON Release of a new, specialized language model with a novel training approach. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →