A new multimodal large language model, Agnes-3.0-Flash, has been released with a 33 billion parameter count and a substantial 262,144 token context window. This model features a hybrid attention architecture, incorporating both delta-rule recurrent layers and standard global attention layers to manage its extensive context. The model also supports adjustable reasoning effort, tool calling, and understanding of text, image, and video inputs. AI
IMPACT Introduces a novel hybrid attention architecture that could influence future LLM designs for handling long contexts.
RANK_REASON Release of a new open-source model with novel architectural details. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →