Qwen3.5 35B A3B
PulseAugur coverage of Qwen3.5 35B A3B — every cluster mentioning Qwen3.5 35B A3B across labs, papers, and developer communities, ranked by signal.
7 day(s) with sentiment data
-
Flash-MoE technique allows large AI models to run on 16GB Macs
A new technique called anemll-flash-llama.cpp enables large Mixture-of-Experts (MoE) models to run on Macs with as little as 16GB of RAM. This method stores model experts on an SSD and only loads necessary experts into …
-
DeepGrove unveils Maple-Preview AI for iPhones, 13x faster than Bonsai 27B
AI research firm DeepGrove has announced Maple-Preview, a new AI model designed for efficient operation on mobile devices like the iPhone. This model boasts 13 times the processing speed of Bonsai 27B, another iPhone-co…
-
RAG outperforms GraphRAG for textbook QA, study finds · arXiv research
A new arXiv paper compares Retrieval-Augmented Generation (RAG) and GraphRAG for question answering on a math textbook, using a dataset of 477 question-answer pairs. The study found that embedding-based RAG models, part…
-
SWE-rebench adds multilingual coding tasks, GLM-5.2 leads leaderboard
The SWE-rebench leaderboard has been updated with a new multilingual slice that evaluates software engineering tasks across five programming languages: Go, Java, Python, Rust, and TypeScript. The update includes perform…
-
OpenAI Slashes GPT-5.6 Prices; Qwen Adds New Model
OpenAI has significantly reduced prices for its GPT-5.6 series, with Terra and Terra Pro models seeing prompt and completion costs slashed by approximately 40-50%. The company also removed the GPT-5 Chat and GPT-4o Sear…
-
New research suggests MoE AI routing mimics Huffman coding
A new research paper proposes that Mixture-of-Experts (MoE) architectures in AI models function similarly to Huffman coding, a data compression technique. The study introduces the Frequency-Diversity Law, which suggests…
-
Shanghai AI Lab releases 35B Agents-A1 model for agentic AI tasks
Shanghai Artificial Intelligence Laboratory has released Agents-A1, a 35-billion parameter Mixture-of-Experts model built on Qwen3.5-35B-A3B. The model, available under the Apache 2.0 license, is designed for complex, m…
-
Qwen3.5-122B model fits 64GB RAM, offering better quality at slower speeds
A user on r/LocalLLaMA shared their experience running the Qwen3.5-122B model with UD-Q2_K_XL quantizations on a system with 64GB of RAM. This setup allows the larger model to fit into memory, offering significantly bet…
-
UMoE pipeline enhances domain-specific MoE model training
Researchers have introduced UMoE, a novel pipeline designed to optimize Mixture-of-Experts (MoE) models for domain-specific tasks. This method involves pruning underperforming experts, regrowing the expert pool to its o…
-
VIDRAFT ships dual LLM serving engines for GPU throughput and CPU reach
VIDRAFT has developed two distinct serving engines for large language models, addressing separate optimization targets. VKAE is a kernel-level acceleration engine designed to maximize throughput on GPUs, achieving up to…
-
Reddit user proposes "Local LLM Survival Kit" for offline AI
A user on Reddit's r/LocalLLaMA forum is proposing the concept of a "Local LLM Survival Kit." This kit would be a portable USB drive containing essential components for running large language models offline. The propose…
-
NVIDIA unveils Audex, a unified audio-text LLM that preserves text intelligence
NVIDIA has introduced Audex (Nemotron-Labs-Audex-30B-A3B), a unified audio-text large language model capable of understanding and generating both audio and speech. Unlike many multimodal models that experience a decline…
-
llama.cpp flag boosts Qwen 35B model speed by 2.8x on RTX 4070
A technical guide demonstrates how to achieve a 2.8x speedup when running the Qwen3.5-35B-A3B model on an RTX 4070 GPU with 12GB of VRAM. The key to this performance increase lies in using the `llama.cpp` framework with…
-
New research tackles LLM reasoning, long-context, and tool integration
Multiple research papers explore advancements in large language model (LLM) reasoning capabilities, focusing on improving performance in long-horizon tasks and tool integration. Apple's research introduces LEAD, a metho…
-
New tuning method boosts LLM coding agent performance
Researchers have developed a new method called probe-and-refine tuning to improve the performance of large language model (LLM) coding agents. This technique focuses on enhancing the guidance files that direct agents to…
-
oMLX significantly outperforms Ollama in Mac LLM inference speed
A performance comparison between oMLX and Ollama for running LLMs locally on Mac devices revealed significant speed differences. oMLX, utilizing Apple Silicon's MLX framework, demonstrated a 35% faster token generation …
-
oMLX boosts Apple Silicon LLM performance with KV cache
oMLX, an open-source LLM inference server for Apple Silicon, has demonstrated significant performance improvements, particularly in handling large models and complex workflows. Community benchmarks and local tests highl…
-
Self-Harness enables LLM agents to improve their own operational harnesses
Researchers have developed a novel method called Self-Harness, enabling LLM-based agents to autonomously improve their own operational harnesses. This iterative process involves identifying model-specific failure patter…
-
Qwen3.5-35B-A3B router shows specific expert for self-reflection
A researcher has documented experiments with the Qwen3.5-35B-A3B model, focusing on how its Mixture-of-Experts (MoE) router behaves when the model generates first-person self-examination text. The findings suggest that …
-
MIRA framework improves LLM mid-training data selection
Researchers have developed MIRA, a novel framework for selecting data during the mid-training phase of large language model development. This method addresses the challenge of heterogeneous data sources by discovering a…