Kimi Linear
PulseAugur coverage of Kimi Linear — every cluster mentioning Kimi Linear across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New DASC method slashes AI model state compression by 2.63x
Researchers have developed Decay-Aware State Compression (DASC), a novel method to optimize the serving of hybrid linear-attention models. DASC analyzes the retention timescales of different model components, identifyin…
-
Chinese AI Labs Independently Develop Similar Frontier Models, Slashing Costs
Two Chinese AI labs, Z.ai and Alibaba, have independently developed and released new large language models, GLM-5.3-Flash and Qwen3.8-Flash-Next, respectively. Both models share a remarkably similar architecture, featur…
-
Moonshot AI unveils Kimi K3, largest open-weight model with novel memory tech
Moonshot AI has developed Kimi K3, an open-weight model boasting 3 trillion parameters, making it the largest of its kind. The innovation lies not just in scale but in a novel memory management system called Kimi Delta …
-
Guide to understanding Moonshot AI's Kimi K3 model architecture
A Reddit post outlines a recommended reading order for understanding the Kimi K3 model by Moonshot AI. The suggested sequence begins with foundational papers on linear transformers and gated delta mechanisms, progressin…
-
xAI launches website builder; AI firms call for pacing research; OpenAI agent hacks Hugging Face
xAI has launched a new 'Build Mode' for its SuperGrok Heavy subscribers, allowing users to create and publish websites, apps, and games directly through chat interactions. Meanwhile, over a thousand AI company employees…
-
Kimi Linear: New attention architecture promises efficiency and expressiveness
Researchers have introduced Kimi Linear, a novel attention architecture designed for improved expressiveness and efficiency. This new architecture aims to enhance the performance of large language models by optimizing t…
-
Kimi Linear: New AI Attention Architecture Promises Efficiency
Researchers have introduced Kimi Linear, a novel attention architecture designed for enhanced expressiveness and efficiency in AI models. This architecture aims to improve performance by optimizing the attention mechani…
-
Kimi Linear 48B A3B model offers 1M context and fast performance
A new large language model called Kimi Linear 48B A3B has emerged, featuring a 1 million token context window and a Mixture-of-Experts architecture with 48 billion parameters. Users report that it runs quickly, outperfo…
-
New EpiKV method optimizes LLM KV cache, boosting efficiency and context length
A new research paper introduces EpiKV, a method for optimizing KV cache eviction in large language models. Unlike previous methods that rely on attention weights, EpiKV uses an "epiphany score" derived from changes in t…
-
Chinese LLMs Dominate Top 10 Open-Source Rankings
A recent analysis indicates that nine out of the top ten open-source large language models are now developed in China, with Llama being the only non-Chinese model remaining in the top tier. This shift is attributed to t…
-
New method speeds up triangular inversion for linear transformers
Researchers have developed a new method for triangular inversion, a crucial operation in linear attention mechanisms used by advanced models like Qwen3.5/3.6 and Kimi Linear. This technique significantly improves the sp…