qwen-3.7-max
PulseAugur coverage of qwen-3.7-max — every cluster mentioning qwen-3.7-max across labs, papers, and developer communities, ranked by signal.
- 2026-05-25 research_milestone Alibaba's Qwen 3.7 Max completed a 35-hour autonomous task run with 1,158 tool calls. source
- 2026-05-25 research_milestone Qwen 3.7 Max completed a 35-hour autonomous task run with 1,158 tool calls. source
- 2026-05-25 research_milestone Alibaba's Qwen 3.7 Max achieved a tenfold speedup on kernel execution after a 35-hour autonomous optimization task on unfamiliar hardware. source
- 2026-05-23 product_launch Qwen has launched the Qwen 3.7-Max model, designed for autonomous AI agents. source
- 2026-05-19 product_launch Alibaba released a preview of its Qwen 3.7 Max model. source
3 day(s) with sentiment data
Qwen 3.7-Max to be integrated into Alibaba Cloud's AI agent services within 60 days
Alibaba has previewed and launched Qwen 3.7-Max, emphasizing its capabilities for autonomous agents and long-duration tasks. The concurrent launch with the Zhenwu M890 AI chip, designed for AI agents, suggests a strategic push to integrate these models into Alibaba's cloud offerings for agent-based solutions.
Alibaba to release Qwen 3 Ultra model targeting high-performance benchmarks within 30 days
The preview of both Qwen 3.7 Max and Qwen 3 Ultra suggests a dual-pronged release strategy. Given that 3.7 Max is positioned for agent tasks, the Ultra variant is likely being developed to compete at the top of general LLM benchmarks, potentially challenging existing leaders.
Qwen 3.7-Max demonstrates strong performance in tool use and autonomous task execution
Recent evidence shows Qwen 3.7-Max successfully handling 1,000 tool calls in a single autonomous agent task, reducing p99 latency to below 400ms over nine hours. This highlights the model's robustness and efficiency in complex, long-running agent operations.
-
Motif Technologies unveils 314B parameter Motif 3 LLM
Motif Technologies has released Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. The model features a novel Grouped Differential Latent At…
-
Qwen 3.8 Max improves debate performance but increases cost
Qwen 3.8 Max has shown improvement over its predecessor, Qwen 3.7 Max, on the Debate Benchmark, increasing its score from 1462 to 1588. However, this enhanced performance came at a cost, with the average cost per debate…
-
OpenAI flags Astra model as critical; Meta's Muse Spark shows gains · 4 sources tracked
OpenAI has escalated its Astra model to a "critical" cyber status due to advancements in agentic coding and cybersecurity, prompting stricter internal controls and a pause on non-essential activities. This move, alongsi…
-
Tencent releases open-source Hunyuan Hy3 LLM, claims performance edge
Tencent has released the final version of its Hunyuan Hy3 large language model under an Apache 2.0 license. This model boasts 295 billion parameters with 21 billion active per token and a 256K context window. While Tenc…
-
US policy shift proposed to boost open AI models against China · 2 sources tracked
Ben Thompson proposes that the U.S. should enact legislation to clarify data collection for AI training as fair use and prohibit terms of service that forbid model distillation. This move aims to bolster U.S. open model…
-
Alibaba releases Qwen 3.8, second only to Claude Fable 5, with open weights · 5 sources tracked
Alibaba has launched its latest AI model, Qwen 3.8, boasting 2.4 trillion parameters. The company positions this model as the second most powerful available, trailing only Anthropic's Claude Fable 5. Notably, Alibaba ha…
-
OpenAI's GPT-5.6 lineup: Sol excels, Luna offers value, user reports autonomy concerns
A user shared their experience with OpenAI's new GPT-5.6 series, detailing three models: Terra, Sol, and Luna. Sol is highlighted as the top performer, excelling in benchmarks. Luna is praised for its affordability and …
-
Anthropic files for $965B IPO, valuing agent infrastructure over intelligence
Anthropic has confidentially filed for an IPO at a staggering $965 billion valuation, based on a $47 billion revenue run-rate and projections of $10.9 billion in Q2 2026. This valuation, representing approximately a 20x…
-
Ollama Cloud Models: DeepSeek V4 Flash Offers Major Cost Savings Over V4 Pro
A recent analysis of Ollama Cloud models reveals significant cost discrepancies based on GPU compute usage per task, rather than just token count. The study found that DeepSeek V4 Flash, despite having fewer active para…
-
Google's Android Bench adds new LLMs; Fable 5 leads, Gemini lags
Google has updated its Android Bench benchmark for evaluating large language models (LLMs) in Android development tasks. The updated leaderboard includes eight new models, such as Claude Fable 5, Claude Sonnet 5, and Qw…
-
Tencent launches Hy3, a 295B MoE model focused on agent performance
Tencent has officially launched its Hunyuan Hy3, a 295 billion parameter Mixture-of-Experts (MoE) model with 21 billion active parameters and a 256K context window. The model is licensed under Apache 2.0 and emphasizes …
-
11 LLMs evaluated on code refactoring and proposal evaluation
An experiment evaluated eleven large language models on their ability to refactor a complex "god node" within a LangGraph agent. The models were tasked with proposing solutions to untangle the node's logic and then eval…
-
TotalEnergies sells Malaysian gas field stake for $350M; Alibaba Cloud offers AI model discounts
TotalEnergies has agreed to sell its 85% stake in the Marjoram gas field in Malaysia to INPEX for $350 million. This move allows TotalEnergies to monetize a non-operated minority interest and refocus on its core operati…
-
Alibaba Cloud's Meoo offers Night Plan with deep discounts on Qwen models
Meoo, an AI service from Alibaba Cloud, has launched a "Night Plan" offering discounted rates for specific models during off-peak hours. The Qwen 3.7-Max model will see discounts as low as 20%, while Qwen 3.7-Plus will …
-
Chinese AI models tested on coding tasks: MiniMax and Kimi lead
A comparative analysis of five Chinese AI models—MiniMax M3, Kimi K2.6, DeepSeek V4 Pro, Qwen 3.7 Max, and GLM 5.1—evaluated on real-world engineering tasks revealed significant differences in their coding capabilities.…
-
AI models struggle to manage virtual companies; Claude Fable 5 leads with $47M profit · 1 source tracked
A recent CEO-Bench competition, designed to test AI's ability to run a virtual SaaS startup, revealed mixed results. While many advanced AI models like GLM 5.1 and Gemini 3 Flash went bankrupt, Claude Fable 5 emerged as…
-
New WUBRG-Bench tests LLMs on complex Magic: The Gathering rules
A new benchmark, WUBRG-Bench, has been developed to test the reasoning capabilities of large language models on complex rule-based systems, specifically using questions from the game Magic: The Gathering. The creator fo…
-
UC Berkeley benchmark reveals massive AI model cost and speed disparities
A new benchmark from UC Berkeley, the ALE benchmark, has revealed significant cost and runtime disparities between various AI models across 55 industries. The benchmark highlights that custom harnesses can outperform co…
-
Fireworks AI launches Qwen 3.7 Plus and Max models
Fireworks AI has announced the availability of Qwen 3.7 Plus and Qwen 3.7 Max models on its inference infrastructure. These models are designed for long-horizon agent loops and offer features like preserved reasoning hi…
-
GPT 5.5 and rivals tested in Tamagotchi game creation
A user conducted a comparative test of several large language models, including GPT 5.5, Claude Opus 4.8, Fable/Mythos 5, Gemini 3.5 Flash, Deepseek V4 Pro, and Qwen 3.7 Max. The models were tasked with creating an inte…