Qwen3.6 35B-A3B
PulseAugur coverage of Qwen3.6 35B-A3B — every cluster mentioning Qwen3.6 35B-A3B across labs, papers, and developer communities, ranked by signal.
8 day(s) with sentiment data
-
Anthropic launches Claude Code Projects; Google updates Gemini agents
Anthropic has launched "Projects" for Claude Code, enabling users to manage multiple parallel cloud sessions within a single conversation, with context passing and continued execution after user departure. Google has up…
-
Qwen3.6 35B A3B vs. Nex-N2.5-mini for coding tasks debated
A user on the r/LocalLLaMA subreddit is seeking advice on choosing between two language models, Qwen3.6 35B A3B and Nex-N2.5-mini, specifically for coding tasks. The user currently employs Qwen3.6 35B A3B with MTP but d…
-
FreeToken engine enables large MoE models on personal PCs
FreeToken is an open-source engine designed to run large Mixture-of-Experts (MoE) models on personal hardware by treating the entire PC as a heterogeneous inference system. It manages MoE models by storing the full expe…
-
ACE framework optimizes MoE LLMs by skipping redundant expert computations
Researchers have developed ACE, a novel framework designed to optimize Mixture-of-Experts (MoE) large language models by adaptively skipping redundant expert computations. This training-free method utilizes a Global Spe…
-
Occamy-1.0: New 35B Co-work Agent Prioritizes Cost-Efficiency
Researchers have introduced Occamy-1.0, a new 35-billion parameter co-work agent model designed for efficiency in complex, multi-step tasks. By further training the Qwen3.6-35B-A3B checkpoint with execution-grounded dat…
-
Perplexity open-sources Lily AI engine for Apple Silicon
Perplexity is open-sourcing its Lily AI engine, designed for local artificial intelligence processing on Apple Silicon. The engine is built using Rust and supports the Qwen3.6-35B-A3B model, aiming to provide faster AI …
-
Perplexity open-sources Lily inference engine for Apple Silicon
Perplexity has open-sourced Lily, a specialized inference engine built with Rust and Metal for running the Qwen3.6-35B-A3B model on Apple Silicon. This engine is designed for narrow hardware optimization, achieving up t…
-
New environment evolution method boosts terminal agent performance · 4 sources tracked
Researchers have developed a new method called "environment evolution" to improve the training of terminal agents. This technique incrementally increases the difficulty of training environments off-policy, providing con…
-
Perplexity open-sources Lily inference engine for Apple Silicon
Perplexity has open-sourced Lily, a local inference engine designed for hybrid compute within its Perplexity Computer product. Lily is specifically optimized for running Qwen3.6-35B-A3B models on Apple silicon, treating…
-
Dual-model literary translation pipeline achieves 2-3 books/day on Tesla P40s
A user has detailed a two-model pipeline for literary book translation, utilizing two Tesla P40 GPUs. The pipeline employs Gemma 4 - 26B-A4B for translation at approximately 40 tokens/second and Qwen3.6 35B-A3B for proo…
-
New LLM benchmarks show Qwen3.8 Flash Next leading on DGX Sparks
A user on r/LocalLLaMA shared performance benchmarks for several new large language models, including DeepSeek V4 Flash, Qwen3.8 Flash Next, Qwen3.8-27B, and Qwen3.6-35B-A3B. The tests were conducted on NVIDIA DGX Spark…
-
Open-source kernel boosts Qwen LLM performance on AMD GPUs
A team has developed and open-sourced an optimized kernel for the Qwen3.6 35B-A3B large language model, specifically targeting AMD MI350X GPUs. Their benchmark results show that 8x MI350X GPUs can achieve over 78,000 ou…
-
Ornith 1.5 and Tiel-Coder lead tool-calling benchmark, outperforming Qwen variants
A benchmark comparing several large language models on tool-calling capabilities reveals that Ornith 1.5 and Tiel-Coder performed best. These models, designed for VRAM-limited hardware, outperformed original Qwen3.6-35B…
-
Raspberry Pi 5 powers local car AI with Qwen model
A developer has created a local AI system for cars using a Raspberry Pi 5 and the Qwen 3.6-35B model. This system operates entirely offline, providing features like departure and arrival notifications, trip summaries, a…
-
AI research explores advanced distillation techniques for model efficiency
Two new research papers explore advanced techniques for knowledge distillation in AI models. The first paper, D$^3$-MOPD, introduces an adaptive scheduling method to dynamically adjust the mixture of domains during mult…
-
New Tiel-Coder 35B model excels at coding and long conversations
A new open-source model, Tiel-Coder-35B-A3B, has been released, optimized for coding tasks and long conversations. It achieves strong performance on the SWE-bench-Live benchmark, fixing 12 out of 25 problems, which is c…
-
TSWAP: AI wellness advisor uses retrieval-augmented Thai medicine knowledge · 2 sources tracked
Researchers have developed TSWAP, an eight-language conversational wellness advisor that uses retrieval-augmented generation to access a verified knowledge base of Thai traditional medicine and wellness providers. The s…
-
llama.cpp PR boosts IQ model prompt processing with AVX2 optimizations · 1 source tracked
A pull request for the llama.cpp project introduces AVX2 optimizations to significantly accelerate prompt processing for IQ models, particularly at large batch sizes. Benchmarks show dramatic speed increases, with some …
-
Ornith-1.5 family of open-source LLMs released, rivals Claude Opus 4.8
AI research organization Ornith has released Ornith-1.5, a family of open-source large language models. The models come in three sizes: Ornith-1.5-397B, Ornith-1.5-35B-A3B, and Ornith-1.5-9B. The largest model, Ornith-1…
-
Depth-Aware Analysis Reveals Sensitivity in Qwen MoE Model Layers
Researchers have developed a depth-aware sensitivity analysis method for Mixture-of-Experts (MoE) models, specifically applied to the Qwen3.6-35B-A3B model. Their findings indicate that early and middle layers are highl…