MiniMax M2.5
PulseAugur coverage of MiniMax M2.5 — every cluster mentioning MiniMax M2.5 across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
CI check manages Chinese LLM model names and token budgets
A developer has created a CI check to manage the rapidly changing landscape of Chinese LLM model names and their associated token budgets. This tool helps ensure production stability by treating model catalogs as deploy…
-
Chinese AI models undercut Western pricing by up to 50x, offering competitive performance · 4 sources tracked
A comparison of AI API pricing in 2026 reveals that Chinese providers like Zhipu AI, Baidu, DeepSeek, and Alibaba Group offer significantly lower costs than Western counterparts such as OpenAI, Anthropic, and Google. Mo…
-
AI self-evolution may start with external systems, not model weights
Wonyong Li, former OpenAI safety VP, proposes a new path for AI self-evolution, suggesting it should begin with the external operating system (Harness) rather than directly modifying model weights. This Harness system m…
-
AI system improves H. pylori detection in biopsy reports
Researchers have developed the Nimblemind Multi-Agent System (nMAS) to improve the extraction of evidence for Helicobacter pylori positivity from gastric biopsy reports. In a pilot study using 54 de-identified reports f…
-
MiniMax models now available on Amazon Bedrock for agentic workloads
Amazon Bedrock is now offering three open-weight foundation models from MiniMax, a global AI technology company. These models, part of the MiniMax M2 family, are designed for software engineering and agentic use cases, …
-
LLM Pricing Fluctuates: NVIDIA, Qwen, and Z.ai See Changes; New Models Added · 10 sources tracked
The Token Ledger has released daily updates on LLM pricing changes throughout early August 2026. Several models saw price adjustments, including NVIDIA Nemotron 3 Super and Ultra, Qwen variants, and Z.ai's GLM 5.2, with…
-
New algorithm TASKER improves video understanding and agentic tasks
Researchers have developed TASKER, a novel keyframe extraction algorithm designed to improve performance in both Video Question Answering (VideoQA) and video-guided agentic tasks. This algorithm, detailed in a new paper…
-
Model Gateway enables MiniMax models in Claude Code and opencode
A third-party service called Model Gateway has been developed to allow developers to integrate MiniMax models, such as MiniMax-M3, into existing coding workflows that typically support OpenAI or Anthropic APIs. This gat…
-
New research explores advanced RL for agent survival, navigation, and explainability · 7 sources tracked
Researchers are exploring advanced techniques in reinforcement learning (RL) to enhance agent performance and interpretability. One study introduces programmatic policies (PERL) as an alternative to neural policies (NER…
-
oMLX boosts Apple Silicon LLM performance with KV cache
oMLX, an open-source LLM inference server for Apple Silicon, has demonstrated significant performance improvements, particularly in handling large models and complex workflows. Community benchmarks and local tests highl…
-
Self-Harness enables LLM agents to improve their own operational harnesses
Researchers have developed a novel method called Self-Harness, enabling LLM-based agents to autonomously improve their own operational harnesses. This iterative process involves identifying model-specific failure patter…
-
MiniMax launches M3 with 1M context, beats GPT-5.5 on SWE-Bench
MiniMax, a Chinese AI startup, has released its M3 model, boasting a 1 million token context window and native multimodality, outperforming GPT-5.5 on the SWE-Bench Pro benchmark. The company also offers its M2.5 model …
-
Open-Source LLMs Evolve: Attention, Multimodality, and Efficiency Gains
The open-source LLM landscape has seen significant shifts in recent months, with Sliding Window Attention becoming mainstream, enabling much larger context windows. QK-Norm is also gaining traction as a training stabili…
-
AI Model Pricing Revolution: Chinese Labs Undercut GPT-5, Gemini on Code Benchmarks
The AI model market has seen a significant shift in pricing and performance, particularly in coding benchmarks like SWE-bench. Models from Chinese labs such as DeepSeek, Kimi, and MiniMax are offering comparable or even…
-
MiniMax-M2 Models Achieve Frontier Performance with Efficient Activations
Researchers have introduced the MiniMax-M2 series, a new family of Mixture-of-Experts language models designed for agentic deployment. The flagship M2 model boasts 229.9 billion total parameters but activates only 9.8 b…
-
Fireworks AI: AI agent reliability, not intelligence, is key bottleneck
A new benchmark by Fireworks AI reveals that the reliability of AI model execution, not just intelligence, is a critical bottleneck for agentic AI systems. In 720 browser automation tasks, one model failed to produce va…
-
New agent framework boosts LLM clinical reasoning with active evidence seeking
Researchers have developed ClinSeekAgent, a novel framework designed to enhance clinical reasoning in large language models by enabling them to actively seek and synthesize multimodal evidence. Unlike previous approache…
-
LLM benchmark shows routing strategy outperforms single model selection
A recent benchmark tested 15 LLMs on 38 real-world coding tasks, revealing that a routing strategy combining different models is more effective than selecting a single top-tier model. The study found that cheaper models…
-
LLM benchmarking issues fixed by adjusting 'thinking mode' parameters
A developer encountered issues benchmarking three large language models, Kimi K2.5, MiniMax M2.5, and Gemma 4, initially deeming them broken due to low scores or errors. The root cause was identified as a default "think…
-
Low-cost AI model beats top performers on coding benchmark with new context engine
A new method called Xanther Context Engine (XCE) has enabled the MiniMax M2.5 model to achieve a 78.2% score on the SWE-bench Verified benchmark, outperforming all other models. This achievement is notable because MiniM…