SGLang
PulseAugur coverage of SGLang — every cluster mentioning SGLang across labs, papers, and developer communities, ranked by signal.
- 2026-09-05 product_launch SGLang released version 0.5.19, featuring a large number of contributions. source
- 2026-08-22 product_launch SGLang released version v0.5.18, featuring 710 pull requests from 212 contributors. source
- 2026-08-22 product_launch SGLang released version 0.5.18. source
- 2026-07-25 product_launch SGLang released version 0.5.16, including the new DSpark speculative decoding algorithm. source
- 2026-07-23 controversy A critical vulnerability, CVE-2026-14890, was disclosed in SGLang, enabling unauthenticated remote code execution. source
- 2026-06-20 product_launch SGLang and MUSA merge backends, enabling native GPU support for China's open-source AI ecosystem. source
- 2026-01-09 product_launch SGLang released version 0.3.1 of its model gateway, featuring performance and memory improvements. source
19 day(s) with sentiment data
-
Nvidia releases GLM-5.3 with 1M context and MoE architecture
Nvidia has released GLM-5.3, a new model utilizing a Mixture-of-Experts (MoE) architecture with 753 billion total parameters and 40 billion active parameters. This model features sparse attention mechanisms, enabling a …
-
LLM tuner PolyServe reveals bugs, boosts performance with quantization
An open-source LLM tuner called PolyServe was developed to optimize model serving configurations. Benchmarking revealed several flaws in the tuner's assumptions, including a quality gate that failed to enforce its inten…
-
XingChen-AGI releases Xing4.0-29B-A4B with 256K context length
XingChen-AGI has released Xing4.0-29B-A4B, a new large language model in the Xing series, formerly known as TeleChat. This model boasts 29 billion parameters with only 4 billion activated per token, enabling a native co…
-
AMD MI355X closes performance gap with GB300 via SGLang and UMBP integration · 3 sources tracked
SemiAnalysis reports that AMD's MI355X is rapidly improving in agentic inference performance and total cost of ownership, nearing parity with GB300. This advancement is attributed to AMD's SGLang team and their MoRI lib…
-
LLMeter CLI measures LLM performance on local hardware
LLMeter is a new command-line interface tool designed to measure the performance of large language models (LLMs) on a user's specific hardware and configuration. Unlike traditional leaderboards that test models on optim…
-
AgentKV improves LLM efficiency with phase-aware KV eviction
Researchers have developed AgentKV, a new method for managing KV cache in agentic Large Language Models (LLMs). AgentKV addresses the issue that traditional KV eviction methods, which rely on recent tokens, fail to acco…
-
Ant Group launches SingProbe, an in-process AI safety guardrail
Ant Group has introduced SingProbe, an internally developed safety system designed to integrate risk detection directly into the AI model's generation process. This approach aims to identify unsafe content and factual i…
-
ROCm vs Vulkan: Choosing AMD GPU Accelerators for Local LLM Hosting
This guide compares ROCm and Vulkan for accelerating AMD GPUs in local LLM hosting, highlighting their distinct roles. ROCm serves as AMD's compute platform for frameworks like PyTorch and engines such as vLLM and SGLan…
-
Agnes-AI releases 33B multimodal model with 262K context window
Agnes-AI has released Agnes-3.0-Flash, a 33 billion parameter multimodal model with a 262,144 token context window. The model features a hybrid-attention architecture, combining gated delta-rule recurrent layers with st…
-
RunningHub accelerates MiniMax H3 video model by 12x with open-source optimizations
RunningHub has developed an open-source optimization called H3 Lightning that significantly accelerates the MiniMax H3 AI video generation model. This new method can speed up video generation by up to 12 times, reducing…
-
DeepSeek-V4.1-Flash model released with safety guardrails removed
The dealignai team has released an uncensored version of the DeepSeek-V4.1-Flash model, named DeepSeek-V4.1-Flash-UNCENSORED-FP8. This version features proprietary weight-level abliteration, surgically removing safety g…
-
Ollama alternatives: A practical guide to llama.cpp, vLLM, SGLang, and LocalAI
A guide offers practical advice for selecting an alternative to Ollama, focusing on isolating failing layers and workloads. It details options such as llama.cpp, vLLM, SGLang, and LocalAI, providing a framework for user…
-
KV-Cache Side Channel Reliability Collapses Under LLM Serving Load
Researchers have investigated the reliability of KV-cache timing side channels in multi-tenant LLM serving environments. Their experiments revealed that contention from multiple users significantly degrades the reliabil…
-
LLM inference engines optimize prompt processing with prefix caching
Large Language Models (LLMs) often recompute the same initial prompt tokens repeatedly, leading to inefficiency. This article explains that the KV cache, which stores intermediate states during token generation, is the …
-
DeepSeek releases V4.1 Flash with efficient MoE architecture
DeepSeek has officially released its V4.1 Flash model, a 552 billion parameter Mixture-of-Experts (MoE) model featuring a Causal-Encoder-Decoder (CED) architecture and native multimodal capabilities. This new model is d…
-
New tool diagnoses LLM serving issues on NVIDIA Blackwell GPUs
A new tool called `blackwell-doctor` has been developed to help diagnose and resolve issues when serving large language models on NVIDIA Blackwell GPUs. The tool systematically tests various configurations, including di…
-
Hugging Face hosts new multimodal Qwen models with broad integration support
The ukisai/Swift-Qwen3.8-27B-GGUF and ukisai/Swift-Qwen3.8-27b models are now available on Hugging Face, offering multimodal capabilities. These models can be integrated with various libraries and inference providers, i…
-
SGLang-Omni tackles multi-model orchestration, not just speed
SGLang-Omni has been developed not to optimize individual models, but as a system for orchestrating multiple cooperating models. This new runtime decouples independent stages of a multi-modal AI, allowing each stage to …
-
Nex AGI unveils Nex-N2.5 agentic model family with multimodal and trillion-parameter options
Nex AGI has released its new family of agentic models, Nex-N2.5, designed for long-horizon tasks and real-world environments. The models are available in three sizes: mini, Pro, and Max. The mini and Pro versions build …
-
OUI-1 generative UI model released on Hugging Face
The OUI-1 model, a fine-tuned version of Google's DiffusionGemma 26B-A4B-it, has been released on Hugging Face. This model is designed for generative UI tasks, capable of writing user interface screens in the openui-lan…