Qwen3.6 35B
PulseAugur coverage of Qwen3.6 35B — every cluster mentioning Qwen3.6 35B across labs, papers, and developer communities, ranked by signal.
7 day(s) with sentiment data
-
User fine-tunes Mistral AI model with Qwen for email classification
A user created an email classifier using Mistral AI's models and n8n for automation. The system, fine-tuned by the user due to Mistral's errors, utilizes Qwen3.6-35B for classifying emails into Archive, Trash, or Spam f…
-
Users seek hardware advice for faster Qwen3.6 35B model inference
A user on Reddit is seeking hardware configurations to achieve high inference speeds with the Qwen3.6 35B model. They are currently experiencing around 270-300 tokens/second for prefill and 30 tokens/second for decode o…
-
Qwen3.6 35B KV cache quantization trade-offs debated
A discussion on the r/LocalLLaMA subreddit explores the trade-offs of quantizing the KV cache for the Qwen3.6 35B model. Users are debating whether reducing the quantization level below Q8 is beneficial, considering the…
-
AI Enthusiast Seeks GPU Upgrade Advice for Larger Local Models
A user on the r/LocalLLaMA subreddit is seeking advice on how to expand their local AI model capabilities beyond the limitations of their current NVIDIA GeForce RTX 4080's VRAM. They are considering adding a budget-frie…
-
LLM users discuss optimal models for 20GB VRAM and 64GB RAM setups
A user on the r/LocalLLaMA subreddit is seeking advice on the best large language model (LLM) for their specific hardware configuration, which includes a laptop with 64GB of DDR5 RAM and an external 20GB VRAM GPU. They …
-
New uncensored Qwen3.6-35B model released on Hugging Face
A new, uncensored version of the Qwen3.6-35B model, named LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V5-GGUF, has been released on Hugging Face. This model is designed to be compatible with various inference …
-
User faces LLM loading issues after adding second RTX 3060 GPU
A user on the r/LocalLLaMA subreddit is experiencing issues loading large language models after adding a second RTX 3060 graphics card. Previously, with a single 3060, the user could load models like Qwen3.6 27B, Qwen3.…
-
LLM User Seeks Advice on Upgrading to 40B+ Parameter Models for Speed and Knowledge
A user on the r/LocalLLaMA subreddit is seeking recommendations for large language models (LLMs) with over 40 billion parameters. They are currently using Qwen3.6 35B but find it lacks general knowledge and acts more as…
-
NVIDIA releases new Nemotron and Qwen AI models on Hugging Face
NVIDIA has released several new AI models and checkpoints, including the Nemotron-3 Nano 30B A3B and quantized versions of Qwen models. These releases, primarily announced on Hugging Face, feature Apache 2.0 licensing a…
-
Adaptive MoE Gating Applied Post-Hoc to Qwen3.6-35B Shows Limited Gains
Researchers have developed a post-hoc adaptive Mixture of Experts (MoE) gating method for the Qwen3.6-35B model, aiming to improve efficiency without retraining. Their approach, implemented as an inference-time patch fo…
-
Ornith 35B benchmarked against Gemma4 31B and Qwen3.6 35B
A new language model, Ornith 35B, has been benchmarked against Gemma4 31B and Qwen3.6 35B using the WebBrain's frozen browser-agent planner benchmark. While Ornith 35B shows promise and slightly outperforms Qwen3.6 35B …
-
Multi-tier MoE caching discussed as future of LLM inference
A discussion on Reddit explores the concept of multi-tier Mixture of Experts (MoE) caching as a potential future direction for MoE model inference. The idea involves strategically distributing model experts across CPU a…
-
User criticizes Hermes agent's UI and UX as slow and ugly
A user on Reddit's r/LocalLLaMA subreddit expressed dissatisfaction with the user interface and user experience of the Hermes agent. Despite its promised features and positive reports from others, the user found the web…
-
LLM enthusiasts debate best CPU inference models and software
Users on the r/LocalLLaMA subreddit are discussing the current state of CPU inference for large language models. Participants are seeking advice on optimal models, quantization methods, and specific software versions li…
-
Users seek coding tools for local LLMs like Qwen3.6
A user is seeking recommendations for coding harnesses that work well with local Large Language Models (LLMs), specifically mentioning Qwen3.6 35B. They have found success with GitHub Copilot for general coding assistan…
-
User seeks Qwen3.6 MoE speedup with MTP optimization
A user on the r/LocalLLaMA subreddit is seeking assistance regarding the performance of the Qwen3.6-35B MoE model when using the MTP (Mixture-of-Tensors) optimization. Despite following the unsloth guide and adjusting v…
-
LocalLLaMA user questions mmproj file compatibility for MTP models
A user on the r/LocalLLaMA subreddit is inquiring about the compatibility of mmproj files between MTP (Multi-Turn Prompting) and non-MTP models. They are specifically asking if these files, which appear to be related to…
-
Gemma4-26B beats Qwen3.6-35B in speed despite slower token output
A user compared the performance of Qwen3.6-35B and Gemma4-26B on a Radeon 7900 XTX GPU, finding that Gemma4-26B was approximately 20% faster in end-to-end task completion despite Qwen3.6-35B having a significantly faste…
-
Qwen3.6 model shows markdown best for quality, HTML for token bloat
A user tested the Qwen3.6 35B model to compare output quality and efficiency across different formatting styles: raw text, markdown, unstyled HTML, and styled HTML. The experiment revealed that while markdown produced t…
-
AMD R9700 GPU runs local LLMs like Qwen3.6:35b surprisingly fast
A user shared their experience running local AI models on a new setup featuring an AMD R9700 GPU with 32 GB of VRAM. They successfully operated models such as Qwen3.6:35b using Ollama and Openwebui, noting the surprisin…