Qwen3.5 4B
PulseAugur coverage of Qwen3.5 4B — every cluster mentioning Qwen3.5 4B across labs, papers, and developer communities, ranked by signal.
- 2026-07-07 research_milestone Researchers presented a method for efficient inference of the Qwen3.5-4B model, achieving significant speedups. source
13 day(s) with sentiment data
-
New method improves LLM agents by combining harness evolution and weight training
Researchers have developed a new method for improving long-horizon Large Language Model (LLM) agents by strategically combining harness evolution and weight training. This approach first evolves the agent's harness to a…
-
New benchmark RH-Detect improves reward hacking detection in LLMs
Researchers have introduced RH-Detect, a unified benchmark designed to improve the detection of reward hacking in language models. This benchmark consolidates data from eleven public datasets into a common schema, creat…
-
New Spatial Latent Reasoning Framework Enhances Visual Grounding Accuracy
Researchers have developed a new framework called Spatial Latent Reasoning (SLR) to improve the accuracy of pointing-gesture visual grounding. SLR structures supervision around an ordered sequence of geometric and visua…
-
VisionWeave enables MLLMs to adaptively allocate visual representations, saving tokens and boosting performance
Researchers have developed VisionWeave, a new method for multimodal large language models (MLLMs) that allows them to adaptively allocate visual representations based on content. This approach contrasts with current met…
-
EmpirioLabsAI releases Aplomb 1, a 5.3B decision model with 1M context
EmpirioLabsAI has released Aplomb 1, an open-weights decision model with 5.3 billion parameters. This model boasts a 1 million token context window and can process text, images, video, and audio in a single request. It …
-
llama.cpp optimizes hexagon architecture for faster AI model inference
The llama.cpp project has released an update, b11430, focusing on performance improvements for the hexagon architecture. This update introduces head-parallel partitioning for flash-attention, allowing each core to proce…
-
AI agents design their own inference hardware with openTPU project
An open-source project called openTPU demonstrates AI's capability to design its own inference hardware. The project, detailed in a single monorepo, includes hardware design in SystemVerilog, a simulator, and host softw…
-
Local LLM vs. Claude Code: 96% of requests still need frontier model
A developer explored running the Claude Code AI agent locally using a Qwen3.5 4B model on an RTX 4070, aiming to reduce costs and enhance privacy. The experiment revealed that while 97% of individual agent steps could b…
-
Rubric Response Theory enhances RL rewards with item response modeling
Researchers have introduced Rubric Response Theory (RRT), a novel approach for generating rewards in reinforcement learning tasks where human judgment is required. Unlike traditional methods that sum points from rubric …
-
New AI agent research focuses on training, reliability, and multi-agent coordination
Several research efforts are focusing on improving AI agents' capabilities and reliability. Hugging Face has introduced AutoSynthData for generating training data for enterprise agents and Holo4, a series of agentic mod…
-
Microsoft releases FrogNano-4B-2609 for software engineering tasks
Microsoft has introduced FrogNano-4B-2609, a specialized 4-billion parameter language model derived from Qwen/Qwen3.5-4B. This model has undergone additional text-only post-training focused on software engineering tasks…
-
New RL method boosts MLLM visual perception with fewer tokens
Researchers have developed Vision-RL2, a novel reinforcement learning approach to enhance fine-grained visual perception in multimodal large language models (MLLMs). This method optimizes a region proposal network by tr…
-
New VLM framework uses specialized tools for improved remote sensing analysis
Researchers have developed a new framework for Change Visual Question Answering (Change VQA) in remote sensing that improves accuracy by enabling a Vision Language Model (VLM) to selectively use specialized tools. This …
-
User seeks Qwen3.5 4B alternative for fast local AI assistant
A user on Reddit is seeking recommendations for a local AI assistant model that can outperform Qwen3.5 4B in terms of general conversation, multilingual capabilities, and tool-use reliability. The user prioritizes speed…
-
New methods enhance video question answering accuracy and reliability · 3 sources tracked
Researchers are developing advanced methods to improve the reliability and accuracy of video question answering (VideoQA) models. One approach focuses on refining answer-level reliability scores by analyzing response gr…
-
Lightning Weave framework boosts AI reasoning accuracy and efficiency
Researchers have developed Lightning Weave, a post-training framework designed to enhance both the accuracy and efficiency of reasoning models. This method composes distinct capabilities from independently trained model…
-
Microsoft researchers unveil FrogNano, a small AI model trained on synthetic tasks
Researchers from Microsoft Research Montréal, Mila, and UC San Diego have developed FrogNano, a 4-billion parameter model built upon Qwen3.5-4B. This model was refined using reinforcement learning with synthetic softwar…
-
New research probes acoustic information loss in audio-conditioned LLMs
Researchers have investigated why audio-conditioned language models often fail to utilize crucial acoustic cues like prosody and emotion. Their study, detailed on arXiv, tested various audio encoders including Whisper-T…
-
TokenRhythm launches NeoHorse-1, an Agent-Native model trained on agent feedback
TokenRhythm, in collaboration with Wuxinqiong, Tsinghua University, Peking University, and Alibaba Group, has introduced NeoHorse-1, an Agent-Native model. This model, available in 4B and 9B versions, aims to integrate …
-
Wang Yunhe's startup releases NeoHorse, an Agent-Native model trained on execution experience
JiYuan LüDòng, a startup founded by former Huawei Noah's Ark Lab director Wang Yunhe, has released its first Agent-Native model, NeoHorse. This model, developed with support from Wuxin Qiong and research contributions f…