Qwen-3.6 27B
PulseAugur coverage of Qwen-3.6 27B — every cluster mentioning Qwen-3.6 27B across labs, papers, and developer communities, ranked by signal.
- used by Multi Token Prediction 90%
- used by PrismML 90%
- instance of Qwen 3.6:35B 90%
- instance of Gemma 4.31B 70%
- competes with GLM-5.2 70%
- used by BeeLlama.cpp 70%
- competes with Qwen 3.6:35B 70%
- competes with Gemma 4: 26b 70%
- developed by PrismML 70%
- competes with Opus 4.8 70%
- competes with Gemma 4 60%
- competes with Gemma 4.31B 50%
- 2026-06-21 product_launch A modified version of the Qwen 3.6 27B model, with reduced safety alignment, has been released. source
19 day(s) with sentiment data
-
LocalLLaMA user shares custom agent upgrades for context and speed
A user on Reddit's r/LocalLLaMA community shared several custom quality-of-life upgrades they implemented for their local AI agents. These enhancements aim to optimize context window usage, improve prompt processing spe…
-
Muse-Glimmer 30B achieves 287 t/s with DFlash speculative decoding
The Muse-Glimmer 30B model, when paired with DFlash speculative decoding, achieved impressive performance metrics in real-world coding tasks. Running on a single RTX 5090 GPU with llama.cpp, the model demonstrated gener…
-
New web design benchmark tests local LLMs: Muse Glimmer, Qwen, DeepSeek
A user on Reddit has developed a new benchmark specifically for evaluating large language models' (LLMs) capabilities in web design tasks. The benchmark was used to compare the performance of three local models: Muse Gl…
-
Meta releases open-weight Muse Glimmer model for agentic tasks
Meta has released Muse Glimmer, a new 30B parameter open-weight model licensed under Apache 2.0. The model is designed for end-to-end agentic task completion, reliable tool use, and multi-step reasoning, showing strong …
-
AI Enthusiasts Seek Uncensored, Smart LLMs for Role-Playing
A user on r/LocalLLaMA is seeking recommendations for large language models that are both intelligent and uncensored, specifically for erotic role-playing (ERP). Their current daily driver, Qwen3-235b-a22b-instruct-2507…
-
Qwen-3.6 27B model configuration shared for llama.cpp
A user on Reddit's r/LocalLLaMA subreddit shared their specific configuration settings for running the Qwen-3.6 27B model using llama.cpp. They detailed parameters such as context size, batch size, and reasoning budget,…
-
AI model release pace questioned, user seeks hardware compatibility info
A user on Mastodon is questioning the rapid release of AI models, asking for a detailed table to understand which models can run on their own hardware. They specifically inquire if anyone has tried the Qwen 3.6 27B model.
-
Dual 3090 GPUs achieve 1600+ tps with Qwen 3.6 27B by switching split modes
A user on Reddit's r/LocalLLaMA community shared their experience optimizing performance for the Qwen 3.6 27B model on a dual 3090 GPU setup. Initially, using `--split-mode tensor` resulted in prompt processing occurrin…
-
Gemma 4's top ranking on SciCode benchmark questioned by users
A user on Reddit's r/LocalLLaMA community is questioning the ranking of Gemma 4 above Qwen-3.6 27B on the SciCode benchmark, as reported by artificialanalysis.ai. The user expresses surprise, stating that this ranking c…
-
Qwen v4 model shows increased hallucination, GLM 5.2 preferred for accuracy
A user on Reddit's r/LocalLLaMA subreddit shared their experience comparing Qwen models, noting that the newer Qwen v4 flash model tends to hallucinate more than its predecessor, Qwen 3.6 27b, despite similar benchmark …
-
User seeks best local LLM for 128GB Mac coding tasks
A user is seeking recommendations for the best local large language model (LLM) to utilize on a company-provided 128GB Mac for daily coding tasks. They are considering Qwen 3.6 27B or potentially a newer 3.8 version, wh…
-
RTX 5090 power limit reduction yields minimal inference performance loss
A user on r/LocalLLaMA shared findings on reducing the power limit of an NVIDIA RTX 5090 graphics card for AI inference. By lowering the power limit to 480W, the card experienced only a negligible performance decrease o…
-
Qwen-3.6 27B model performance optimization discussed on Reddit
A user on Reddit's r/LocalLLaMA subreddit is seeking to optimize the performance of the Qwen-3.6 27B model running on a setup with three NVIDIA 2080 Ti GPUs. They are currently achieving 55 tokens per second using llama…
-
Meituan releases LongCat-Flash-Lite-Sparse with 1M token context
Meituan has released LongCat-Flash-Lite-Sparse, a new language model built upon LongCat-Flash-Lite. This sparse model replaces dense MLA with LongCat Sparse Attention and supports context lengths up to 1 million tokens,…
-
NVIDIA GB10/DGX Spark users debate best AI model performance
A user on Reddit's r/LocalLLaMA community is seeking recommendations for the best performing and most stable AI model that can run on a single NVIDIA GB10/DGX Grace Blackwell Superchip with Apache Spark. The discussion …
-
New MyoCardBench benchmark evaluates LLMs in cardiovascular care
A new benchmark called MyoCardBench has been developed to evaluate large language models (LLMs) in realistic cardiovascular care scenarios. The benchmark, comprising 2,263 items across 13 datasets, was used to test seve…
-
New SINT-Flow framework automates schema integration using LLMs
Researchers have introduced SINT-Flow, a novel framework designed for automated schema integration using large language models. This system employs five LLM-based operators that can be combined into workflows to unify d…
-
Qwen-3.6 AI model identifies OWASP vulnerabilities with added skills
The Qwen-3.6-27B AI model demonstrates proficiency in identifying OWASP vulnerabilities when equipped with specific skills. This capability is highlighted through a GitHub repository and documentation on opencode.ai, wh…
-
Dual GPU bottleneck explained: Layer-by-layer model splitting limits performance
A Reddit user discovered that their dual Nvidia RTX 5060 Ti setup was only reaching about 50% utilization when running the Qwen 3.6 27B model due to a layer-by-layer model splitting method. This method causes the GPUs t…
-
User reports significant improvements in Qwen-3.6 27B fine-tune
A user on Reddit's r/LocalLLaMA subreddit has shared positive experiences with a fine-tuned version of the Qwen-3.6 27B model, developed collaboratively by DavidAU and others. The user noted significant improvements in …