PulseAugur
EN
LIVE 13:17:18
ENTITY Qwen3.5 35B A3B

Qwen3.5 35B A3B

PulseAugur coverage of Qwen3.5 35B A3B — every cluster mentioning Qwen3.5 35B A3B across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
11
25 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
5
12 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

7 day(s) with sentiment data

RECENT · PAGE 1/2 · 25 TOTAL
  1. TOOL · CL_190629 ·

    Flash-MoE technique allows large AI models to run on 16GB Macs

    A new technique called anemll-flash-llama.cpp enables large Mixture-of-Experts (MoE) models to run on Macs with as little as 16GB of RAM. This method stores model experts on an SSD and only loads necessary experts into …

  2. SIGNIFICANT · CL_184463 ·

    DeepGrove unveils Maple-Preview AI for iPhones, 13x faster than Bonsai 27B

    AI research firm DeepGrove has announced Maple-Preview, a new AI model designed for efficient operation on mobile devices like the iPhone. This model boasts 13 times the processing speed of Bonsai 27B, another iPhone-co…

  3. TOOL · CL_169767 ·

    RAG outperforms GraphRAG for textbook QA, study finds · arXiv research

    A new arXiv paper compares Retrieval-Augmented Generation (RAG) and GraphRAG for question answering on a math textbook, using a dataset of 477 question-answer pairs. The study found that embedding-based RAG models, part…

  4. TOOL · CL_169015 ·

    SWE-rebench adds multilingual coding tasks, GLM-5.2 leads leaderboard

    The SWE-rebench leaderboard has been updated with a new multilingual slice that evaluates software engineering tasks across five programming languages: Go, Java, Python, Rust, and TypeScript. The update includes perform…

  5. TOOL · CL_168210 ·

    OpenAI Slashes GPT-5.6 Prices; Qwen Adds New Model

    OpenAI has significantly reduced prices for its GPT-5.6 series, with Terra and Terra Pro models seeing prompt and completion costs slashed by approximately 40-50%. The company also removed the GPT-5 Chat and GPT-4o Sear…

  6. TOOL · CL_160704 ·

    New research suggests MoE AI routing mimics Huffman coding

    A new research paper proposes that Mixture-of-Experts (MoE) architectures in AI models function similarly to Huffman coding, a data compression technique. The study introduces the Frequency-Diversity Law, which suggests…

  7. SIGNIFICANT · CL_149991 ·

    Shanghai AI Lab releases 35B Agents-A1 model for agentic AI tasks

    Shanghai Artificial Intelligence Laboratory has released Agents-A1, a 35-billion parameter Mixture-of-Experts model built on Qwen3.5-35B-A3B. The model, available under the Apache 2.0 license, is designed for complex, m…

  8. COMMENTARY · CL_145230 ·

    Qwen3.5-122B model fits 64GB RAM, offering better quality at slower speeds

    A user on r/LocalLLaMA shared their experience running the Qwen3.5-122B model with UD-Q2_K_XL quantizations on a system with 64GB of RAM. This setup allows the larger model to fit into memory, offering significantly bet…

  9. RESEARCH · CL_141148 ·

    UMoE pipeline enhances domain-specific MoE model training

    Researchers have introduced UMoE, a novel pipeline designed to optimize Mixture-of-Experts (MoE) models for domain-specific tasks. This method involves pruning underperforming experts, regrowing the expert pool to its o…

  10. TOOL · CL_139122 ·

    VIDRAFT ships dual LLM serving engines for GPU throughput and CPU reach

    VIDRAFT has developed two distinct serving engines for large language models, addressing separate optimization targets. VKAE is a kernel-level acceleration engine designed to maximize throughput on GPUs, achieving up to…

  11. MEME · CL_137539 ·

    Reddit user proposes "Local LLM Survival Kit" for offline AI

    A user on Reddit's r/LocalLLaMA forum is proposing the concept of a "Local LLM Survival Kit." This kit would be a portable USB drive containing essential components for running large language models offline. The propose…

  12. SIGNIFICANT · CL_131038 ·

    NVIDIA unveils Audex, a unified audio-text LLM that preserves text intelligence

    NVIDIA has introduced Audex (Nemotron-Labs-Audex-30B-A3B), a unified audio-text large language model capable of understanding and generating both audio and speech. Unlike many multimodal models that experience a decline…

  13. TOOL · CL_120180 ·

    llama.cpp flag boosts Qwen 35B model speed by 2.8x on RTX 4070

    A technical guide demonstrates how to achieve a 2.8x speedup when running the Qwen3.5-35B-A3B model on an RTX 4070 GPU with 12GB of VRAM. The key to this performance increase lies in using the `llama.cpp` framework with…

  14. RESEARCH · CL_128948 ·

    New research tackles LLM reasoning, long-context, and tool integration

    Multiple research papers explore advancements in large language model (LLM) reasoning capabilities, focusing on improving performance in long-horizon tasks and tool integration. Apple's research introduces LEAD, a metho…

  15. RESEARCH · CL_96671 ·

    New tuning method boosts LLM coding agent performance

    Researchers have developed a new method called probe-and-refine tuning to improve the performance of large language model (LLM) coding agents. This technique focuses on enhancing the guidance files that direct agents to…

  16. RESEARCH · CL_88570 ·

    oMLX significantly outperforms Ollama in Mac LLM inference speed

    A performance comparison between oMLX and Ollama for running LLMs locally on Mac devices revealed significant speed differences. oMLX, utilizing Apple Silicon's MLX framework, demonstrated a 35% faster token generation …

  17. RESEARCH · CL_88575 ·

    oMLX boosts Apple Silicon LLM performance with KV cache

    oMLX, an open-source LLM inference server for Apple Silicon, has demonstrated significant performance improvements, particularly in handling large models and complex workflows. Community benchmarks and local tests highl…

  18. TOOL · CL_79558 ·

    Self-Harness enables LLM agents to improve their own operational harnesses

    Researchers have developed a novel method called Self-Harness, enabling LLM-based agents to autonomously improve their own operational harnesses. This iterative process involves identifying model-specific failure patter…

  19. TOOL · CL_75622 ·

    Qwen3.5-35B-A3B router shows specific expert for self-reflection

    A researcher has documented experiments with the Qwen3.5-35B-A3B model, focusing on how its Mixture-of-Experts (MoE) router behaves when the model generates first-person self-examination text. The findings suggest that …

  20. TOOL · CL_68731 ·

    MIRA framework improves LLM mid-training data selection

    Researchers have developed MIRA, a novel framework for selecting data during the mid-training phase of large language model development. This method addresses the challenge of heterogeneous data sources by discovering a…