Qwen3.5 35B A3B
PulseAugur coverage of Qwen3.5 35B A3B — every cluster mentioning Qwen3.5 35B A3B across labs, papers, and developer communities, ranked by signal.
- 2026-08-20 research_milestone The Qwen3.5 35B A3B model achieved a high score on the GPQA benchmark and demonstrated strong cost-efficiency. source
2 day(s) with sentiment data
-
MetaRSI-v1 advances AI self-improvement capabilities · 1 source tracked
CosmosMind, in collaboration with several universities, has introduced MetaRSI-v1, a novel meta-recursive architecture designed to improve the process of recursive self-improvement (RSI) in AI models. This new framework…
-
Open-weight AI models advance with new architectures and agent capabilities · 1 source tracked
This week's AI developments highlight advancements in open-weight models, with Qwen and Z.ai introducing new sparse/hybrid architectures focused on inference cost and quality. Apodex released a practical 36B/3B-active M…
-
Qwen3.5 35B A3B model achieves high GPQA score and cost-efficiency
The Qwen3.5 35B A3B model has demonstrated impressive performance, achieving an 81.9% score on the GPQA benchmark. It also offers a high intelligence-to-cost ratio, delivering 35.3 intelligence points per dollar. The mo…
-
AI API Digest: Qwen prices surge, Z.ai adds 1M context, AI21 & Mancer models removed
The AI API Digest for August 20, 2026, highlights significant pricing changes and model updates across various providers. Qwen's Qwen3.6 27B model saw a substantial price increase for both prompt and completion tokens, …
-
Researchers develop Self-Harness for LLM agents to autonomously improve their own systems
A new research paper introduces "Self-Harness," a method allowing LLM-based agents to autonomously improve their own operating harnesses. This iterative process involves identifying model-specific failure patterns, gene…
-
Flash-MoE technique allows large AI models to run on 16GB Macs
A new technique called anemll-flash-llama.cpp enables large Mixture-of-Experts (MoE) models to run on Macs with as little as 16GB of RAM. This method stores model experts on an SSD and only loads necessary experts into …
-
DeepGrove unveils Maple-Preview AI for iPhones, 13x faster than Bonsai 27B
AI research firm DeepGrove has announced Maple-Preview, a new AI model designed for efficient operation on mobile devices like the iPhone. This model boasts 13 times the processing speed of Bonsai 27B, another iPhone-co…
-
RAG outperforms GraphRAG for textbook QA, study finds · arXiv research
A new arXiv paper compares Retrieval-Augmented Generation (RAG) and GraphRAG for question answering on a math textbook, using a dataset of 477 question-answer pairs. The study found that embedding-based RAG models, part…
-
SWE-rebench adds multilingual coding tasks, GLM-5.2 leads leaderboard
The SWE-rebench leaderboard has been updated with a new multilingual slice that evaluates software engineering tasks across five programming languages: Go, Java, Python, Rust, and TypeScript. The update includes perform…
-
OpenAI Slashes GPT-5.6 Prices; Qwen Adds New Model
OpenAI has significantly reduced prices for its GPT-5.6 series, with Terra and Terra Pro models seeing prompt and completion costs slashed by approximately 40-50%. The company also removed the GPT-5 Chat and GPT-4o Sear…
-
New research suggests MoE AI routing mimics Huffman coding
A new research paper proposes that Mixture-of-Experts (MoE) architectures in AI models function similarly to Huffman coding, a data compression technique. The study introduces the Frequency-Diversity Law, which suggests…
-
Shanghai AI Lab releases 35B Agents-A1 model for agentic AI tasks
Shanghai Artificial Intelligence Laboratory has released Agents-A1, a 35-billion parameter Mixture-of-Experts model built on Qwen3.5-35B-A3B. The model, available under the Apache 2.0 license, is designed for complex, m…
-
Qwen3.5-122B model fits 64GB RAM, offering better quality at slower speeds
A user on r/LocalLLaMA shared their experience running the Qwen3.5-122B model with UD-Q2_K_XL quantizations on a system with 64GB of RAM. This setup allows the larger model to fit into memory, offering significantly bet…
-
UMoE pipeline enhances domain-specific MoE model training
Researchers have introduced UMoE, a novel pipeline designed to optimize Mixture-of-Experts (MoE) models for domain-specific tasks. This method involves pruning underperforming experts, regrowing the expert pool to its o…
-
VIDRAFT ships dual LLM serving engines for GPU throughput and CPU reach
VIDRAFT has developed two distinct serving engines for large language models, addressing separate optimization targets. VKAE is a kernel-level acceleration engine designed to maximize throughput on GPUs, achieving up to…
-
Reddit user proposes "Local LLM Survival Kit" for offline AI
A user on Reddit's r/LocalLLaMA forum is proposing the concept of a "Local LLM Survival Kit." This kit would be a portable USB drive containing essential components for running large language models offline. The propose…
-
NVIDIA unveils Audex, a unified audio-text LLM that preserves text intelligence
NVIDIA has introduced Audex (Nemotron-Labs-Audex-30B-A3B), a unified audio-text large language model capable of understanding and generating both audio and speech. Unlike many multimodal models that experience a decline…
-
llama.cpp flag boosts Qwen 35B model speed by 2.8x on RTX 4070
A technical guide demonstrates how to achieve a 2.8x speedup when running the Qwen3.5-35B-A3B model on an RTX 4070 GPU with 12GB of VRAM. The key to this performance increase lies in using the `llama.cpp` framework with…
-
New research tackles LLM reasoning, long-context, and tool integration
Multiple research papers explore advancements in large language model (LLM) reasoning capabilities, focusing on improving performance in long-horizon tasks and tool integration. Apple's research introduces LEAD, a metho…
-
New tuning method boosts LLM coding agent performance
Researchers have developed a new method called probe-and-refine tuning to improve the performance of large language model (LLM) coding agents. This technique focuses on enhancing the guidance files that direct agents to…