Qwen 3.6:35B
PulseAugur coverage of Qwen 3.6:35B — every cluster mentioning Qwen 3.6:35B across labs, papers, and developer communities, ranked by signal.
- 2026-05-24 product_launch Release of the uncensored Genesis APEX MTP version of the Qwen 3.6-35B model. source
8 day(s) with sentiment data
-
Inkling-Small 276B-A12B model optimized for low-memory consumer hardware
A new conversion of the Inkling-Small 276B-A12B model, which has approximately 12 billion active parameters, has been optimized to run on consumer-grade hardware with less than 10GB of memory. Benchmarks show the model …
-
Mach-1 Additive model rivals Qwen 3.6 35B performance at 10x smaller size
A new model called Mach-1 Additive is reportedly achieving 95% of the performance of Qwen 3.6 35B while being 10 times smaller. This development has sparked discussion within the local LLM community, with users question…
-
KAT Coder 2.5 praised for speed and accuracy over Qwen, Gemma
A developer is recommending the KAT Coder 2.5 model, highlighting its speed and accuracy improvements over other models like Qwen 3.6 35b and Gemma 4. The developer has shared a GitHub repository detailing their testing…
-
TurboFieldfare engine adapted for Qwen 3.6 35B, reducing RAM usage
A user has successfully adapted the TurboFieldfare engine, originally designed for Gemma models, to support Qwen 3.6 35B. This porting effort resulted in the Qwen model requiring less RAM, approximately 1.4 GB compared …
-
New AfriEconQA benchmark challenges AI on economic report reasoning
Researchers have introduced AfriEconQA, a new benchmark designed to test the quantitative and temporal reasoning capabilities of AI models when processing lengthy institutional documents. This benchmark, derived from 22…
-
Nifer inference engine achieves 720t/s on Qwen 3.6 35B model
A new inference engine called Nifer has been released, reportedly achieving speeds of 550-720 tokens per second on a Qwen 3.6 35B model. This performance, described as "insane" by users, is achieved without complex batc…
-
Kimi Linear 48B A3B model offers 1M context and fast performance
A new large language model called Kimi Linear 48B A3B has emerged, featuring a 1 million token context window and a Mixture-of-Experts architecture with 48 billion parameters. Users report that it runs quickly, outperfo…
-
MadMax Ltx2.3 model detailed with prompt examples
A Reddit user shared details about the MadMax Ltx2.3 model, a quantized version of a text-to-video model, noting its use with a Qwen 3.6 35b prompt generator. The post includes detailed examples of prompts designed to g…
-
Hermes LLM runs on Android via Graphene OS with remote gateway
A user has successfully integrated the Hermes large language model with Android, specifically on Graphene OS, utilizing a remote gateway setup. The system leverages Llama.cpp and Qwen 3.6 35b for its backend, with impre…
-
Strix Halo praised for energy efficiency and value in local LLM deployment
A Reddit user shared their experience with the Strix Halo, highlighting its energy efficiency and value for running local large language models. They reported that the device consumes at most $0.48 per day, even under h…
-
Local AI assistant Bella runs on ESP32 board using Qwen 3.6 model
A user has successfully integrated a local AI assistant, named Belochka or Bella, into their daily life, running on a Mac and accessible via an ESP32 w-10 board. This setup utilizes the Qwen 3.6 35B model and allows Bel…
-
LLM community calls for urgent release of 80-160B parameter models
Users on the r/LocalLLaMA subreddit are expressing a strong need for new large language models (LLMs) in the 80-160 billion parameter range. Current models are either too small for users with high-capacity but slower un…
-
AI system ACIE achieves 96.5% accuracy in clinical data extraction
A new agentic retrieval-augmented generation (RAG) system called ACIE has been developed and deployed at University Medicine Essen for clinical information extraction. This system addresses limitations in standard RAG b…
-
Qwen 3.6 35B model runs on consumer hardware with 32k context
A user on Reddit shared their experience running the Qwen 3.6 35B model on a consumer-grade setup, including an RTX 3080 GPU and 32GB of RAM. They achieved a throughput of 26 tokens/second for generation and 1400 tokens…
-
Cohere releases North-Mini-Code-1.0 coding model
Cohere has released North-Mini-Code-1.0, a 30 billion parameter coding model. While its general artificial analysis score is lower than some competitors, it performs competitively in coding benchmarks. The model is avai…
-
Qwen 3.6 35B model excels with KV cache in agentic tasks
A user on r/LocalLLaMA found that the Qwen 3.6 35B model significantly outperforms the 27B version, particularly in agentic tasks, when using KV cache. This user initially favored the 27B model for its perceived intelli…
-
DDR5 Bandwidth Bottlenecks Dual-LLM Inference on AMD APUs
A developer's experiment revealed that the DDR5 bandwidth on AMD APUs significantly limits the performance of running multiple large language models simultaneously. Despite a 35-billion-parameter model like Qwen 3.6:35B…
-
Qwen 3.6-35B model released with uncensored Genesis APEX MTP version
A new, uncensored version of the Qwen 3.6-35B model, named Genesis APEX MTP, has been released. This model boasts impressive performance, handling up to 200k context without glitches and successfully managing complex, i…
-
LM Studio adds MTP Speculative Decoding for faster local LLM inference
LM Studio has updated to version 0.4.14 Build 2 (Beta), integrating MTP Speculative Decoding to accelerate local large language model inference. This feature allows for faster text generation by predicting multiple toke…
-
Qwen 3.6 27B model shows strong local coding ability
The Qwen 3.6 27B model has demonstrated impressive coding capabilities, marking it as the first local model under 100 billion parameters to perform well on Codex tasks with minimal prompting. While the Qwen 3.6 35B vari…