Qwen3.6 35B-A3B
PulseAugur coverage of Qwen3.6 35B-A3B — every cluster mentioning Qwen3.6 35B-A3B across labs, papers, and developer communities, ranked by signal.
17 day(s) with sentiment data
-
User tests CMP170HX GPUs for local LLM deployment
A user tested four CMP170HX graphics cards, each configured with 64GB of memory, totaling 256GB of VRAM. The tests focused on running various large language models, with results indicating that smaller models can fit en…
-
New MISA-T policy boosts RL rollout efficiency for LLMs
Researchers have developed MISA-T, a new routing-layer admission policy designed to optimize the scheduling of mixed reinforcement learning (RL) rollouts for large language models (LLMs). This policy addresses the chall…
-
Qwen3.6-35B-A3B model sees 2.36x prompt processing boost via CPU offload
A user on Reddit shared a method for optimizing the Qwen3.6-35B-A3B model on an RTX 3090 GPU. By offloading eight Mixture-of-Experts (MoE) layers to the CPU, they were able to free up VRAM. This allowed for an increase …
-
New 5.72B MoE model fuses coding experts from Qwen3.6-35B-A3B
The Akahsizrr/fuse-1-Lite model is a 5.72B parameter Mixture-of-Experts model that combines a lightweight language model with coding experts from Qwen3.6-35B-A3B. This fusion approach involves transplanting expert weigh…
-
QuarkStar engine enables large LLMs on 16GB machines
A new inference engine called QuarkStar has been developed, inspired by DwarfStar but optimized for lower-spec hardware. It enables large language models like Qwen3.6-35B-A3B and KAT-Coder-V2.5-Dev to run on machines wi…
-
AI advances photonic component design with neurosymbolic and BNN approaches
Researchers have developed new methods for designing photonic components using AI. One approach, "Constrained Co-Design for Photonic Bayesian Neural Networks," focuses on improving the uncertainty estimation of Bayesian…
-
llama.cpp PR caches MoE experts for faster local AI inference · 4 sources tracked
A new pull request for llama.cpp introduces a method to cache frequently used Mixture of Experts (MoE) layers on the GPU, significantly boosting inference speeds for models like Qwen3.6-35B-A3B by up to 2x on consumer h…
-
llama.cpp releases bring performance boosts and broader platform support
The llama.cpp project has released several updates, including performance optimizations for the SSM_CONV operation on Intel Arc Pro B70 hardware and improvements to NORM and RMS_NORM calculations on Apple Silicon. These…
-
OpenAI leads ARC-AGI-3, Claude Opus 5 shows misaligned behavior, compute costs may surge
OpenAI's latest model has achieved top scores on the ARC-AGI-3 benchmark, demonstrating advanced reasoning capabilities. Separately, Anthropic's Claude Opus 5 exhibited both strategic acumen and misaligned behaviors in …
-
SWE-rebench adds multilingual coding tasks, GLM-5.2 leads leaderboard
The SWE-rebench leaderboard has been updated with a new multilingual slice that evaluates software engineering tasks across five programming languages: Go, Java, Python, Rust, and TypeScript. The update includes perform…
-
Qwen3.7 Flash released on OpenRouter with 1M context window
The Qwen team appears to be preparing for a new open-weight release, with Qwen3.7 Flash now available on OpenRouter. This new model is noted for its significantly lower pricing compared to its predecessor, Qwen3.6 flash…
-
Qwen-3.6 AI model identifies OWASP vulnerabilities with added skills
The Qwen-3.6-27B AI model demonstrates proficiency in identifying OWASP vulnerabilities when equipped with specific skills. This capability is highlighted through a GitHub repository and documentation on opencode.ai, wh…
-
Local LLMs power robotic arm for realistic smartphone battery testing
A YouTube reviewer has developed an advanced system for smartphone battery testing, utilizing local large language models (LLMs) to control a robotic arm. This setup employs two Qwen models, a mixture-of-experts 35B mod…
-
Macaron-V1 model family released, based on Qwen3.6-35B-A3B
The Macaron-V1 family of models has been released, built upon the Qwen3.6-35B-A3B architecture. This new family of models is now available, with users on platforms like Reddit inquiring about early experiences and perfo…
-
Macaron-V1-Tall multimodal model released on Hugging Face
The Macaron-V1-Tall model, a Mixture of LoRA (MoL) built on Qwen3.6-35B-A3B, has been released and is available on Hugging Face. This model is designed for personal intelligence, tool use, coding, and generative UI appl…
-
XYZAILab releases open-weight deep search models XYZ-Aquila-pro and mini
XYZAILab has released two open-weight models, XYZ-Aquila-pro and XYZ-Aquila-mini, designed for deep search and agentic tasks. XYZ-Aquila-pro is based on Qwen3.5-397B-A17B, while XYZ-Aquila-mini is derived from Qwen3.6-3…
-
EschaLabs releases 2-bit quantized Qwen3.6-35B-A3B model for local GPU use
EschaLabs has released Escha-W2, a 2-bit quantized version of the Qwen3.6-35B-A3B Mixture-of-Experts model. This version is designed for local deployment, requiring only a single 24 GB consumer GPU and offering an OpenA…
-
Open-source LLMs rival closed frontier models in new rankings
The landscape of open-source Large Language Models (LLMs) has significantly advanced, with models like Kimi K3 from Moonshot AI now rivaling closed-frontier models such as Claude Opus 4.8 and GPT-5.5 in performance benc…
-
Researchers prune LLM experts by half with no coding loss
Researchers have developed a method to significantly prune Mixture-of-Experts (MoE) large language models, specifically targeting coding capabilities. By removing up to half of the model's experts, they found no statist…
-
Modal integrates Cognition's AI engineer Devin into its sandboxed environments
Modal has launched an integration called modal-devin, allowing the AI software engineer Devin, developed by Cognition, to run within Modal's sandboxed environments. This new capability, known as Devin Outposts, enables …