Kimi K2.5
PulseAugur coverage of Kimi K2.5 — every cluster mentioning Kimi K2.5 across labs, papers, and developer communities, ranked by signal.
- developed by Composer 2.5 90%
- used by Inkling 90%
- used by Fireworks AI 90%
- uses Composer 2.5 90%
- developed Kimi k3 90%
- used by Composer 2.5 70%
- developed Fireworks AI 70%
- used by jalapeño 70%
- developed by Fireworks AI 70%
- competes with Thinking machines 70%
- competes with Claude Sonnet 4.5 70%
- competes with GLM-5.2 60%
9 day(s) with sentiment data
-
Real-world email test reveals LLM formatting failures in cheaper models
A company that uses AI to generate cold emails found that cheaper models like DeepSeek V4 Flash, Gemini Flash Lite, and GLM failed to maintain proper email formatting, specifically collapsing paragraphs into a single bl…
-
Developer's Kimi model API costs surge due to OpenRouter aggregator
A developer building a coding assistant found that using the OpenRouter aggregator for Moonshot AI's Kimi models led to unexpectedly high API costs. While convenient for accessing multiple models like Kimi K2.5, K2.6, a…
-
New Context Compilation Architecture Boosts LLM In-Context Learning
Researchers have introduced a new Context Compilation Architecture (CCA) designed to improve how large language models handle in-context learning (ICL). The CCA aims to address the brittleness of current models in tasks…
-
New DVBench benchmark evaluates MLLMs on data videos
Researchers have introduced DVBench, a new benchmark designed to evaluate multimodal large language models (MLLMs) on their ability to understand data videos. These videos combine dynamic charts with narrative elements,…
-
MineBench 4.0 released with community gallery, iOS app, and private model testing
MineBench, a benchmark for evaluating AI models' ability to generate 3D structures, has released version 4.0. This update includes a community gallery for users to showcase and upvote custom prompts, and introduces A/B …
-
Kimi K2.5 achieves strong benchmark scores with competitive pricing
Kimi K2.5 has achieved notable performance on several benchmarks, including GPQA, HLE, Long Context, and SciCode. The model offers competitive pricing at 30 integer points per dollar across these evaluations. These resu…
-
New PeakBench benchmark reveals AI agent execution failures due to resource limits
A new benchmark called PeakBench has been introduced to evaluate the execution capabilities of AI agents, moving beyond simple planning accuracy. This benchmark highlights that agents can correctly identify parallelizab…
-
Anthropic targets $30T market for IPO, signaling AI's broad economic reach
Anthropic is reportedly preparing for an IPO and plans to present investors with a total addressable market (TAM) estimate exceeding $30 trillion. This figure, which encompasses all potential work AI models could perfor…
-
OpenAI's custom 'Jalapeno' chip reportedly beats NVIDIA Blackwell in performance
OpenAI has developed a custom AI chip, codenamed "Jalapeno," which reportedly outperforms NVIDIA's latest Blackwell architecture in performance and efficiency. SemiAnalysis, an independent research firm, tested the chip…
-
OpenAI unveils custom Jalapeño ASIC for inference workloads
OpenAI has developed an in-house inference ASIC named Jalapeño, designed in collaboration with Broadcom. This custom chip aims to provide the optimal compute platform for OpenAI's inference workloads, focusing on perfor…
-
OpenAI's Jalapeño ASIC benchmarks show performance gains over Nvidia GPUs
OpenAI has developed its own 700W inference ASIC, codenamed Jalapeño, in collaboration with Broadcom. Benchmarks presented by OpenAI suggest that Jalapeño outperforms Nvidia's GB200 and GB300 GPUs in throughput per kilo…
-
New research optimizes visual token processing for long-video MLLMs
Researchers are exploring methods to optimize how multimodal large language models (MLLMs) process visual information, particularly for long videos. Several papers introduce techniques for selecting, compressing, and pr…
-
Moonshot AI unveils Kimi K3, largest open-weight model with novel memory tech
Moonshot AI has developed Kimi K3, an open-weight model boasting 3 trillion parameters, making it the largest of its kind. The innovation lies not just in scale but in a novel memory management system called Kimi Delta …
-
Speculative Decoding Matures, Accelerating LLM Inference
Speculative decoding, a technique for accelerating LLM inference, has matured significantly, with frameworks adopting it and users reporting impressive performance gains. While the core concept has existed for years, it…
-
AI Labs Launch New Models and Infrastructure Amidst Rapid Development
Several AI labs have released new models and infrastructure updates. Google launched Gemini 3.7 Flash, emphasizing improved coding and agentic capabilities with a significant price cut. Meta released Muse Glimmer, an op…
-
Kimi K2.5 multimodal agent model released with Agent Swarm framework
Researchers have introduced Kimi K2.5, an open-source multimodal agentic model designed to enhance general agentic intelligence through joint optimization of text and vision. The model incorporates techniques like joint…
-
New frameworks and methods tackle bias in LLM judges · 4 sources tracked
Researchers are developing new methods to address scoring bias in Large Language Models (LLMs) when they are used as judges for evaluating text quality. One approach involves instructing LLMs to generate random numbers …
-
AMD MI355X kernel optimizations show 4x performance boost, still trails competitors
A recent kernel hackathon organized by AMD and the GPU_MODE community has led to a significant performance improvement for AMD's MI355X graphics card. The Readonflow Team's optimized kernels reportedly boosted end-to-en…
-
Chinese AI researchers flock to X for global discourse
Chinese AI researchers are increasingly using X (formerly Twitter) to share their work and engage in global discussions about AI development. Platforms like Moonshot AI, Minimax, and DeepSeek have seen their researchers…
-
AI agents tackle deception and reasoning in social deduction games · 2 sources tracked
Researchers have developed new AI agents capable of playing complex social deduction games, which require nuanced skills like deception and reasoning. One agent, CaM-Wolf, integrates multimodal perception, processing vi…