AI news — August 4, 2026
The 20 top stories PulseAugur surfaced that day, ranked by signal across labs, papers, and developer communities.
-
LiquidAI releases LFM2.5-2.6B for efficient local agent deployment
LiquidAI has released LFM2.5-2.6B, a new language model designed for efficient local agent deployment. Despite its small size, the model demonstrates competitive performance against significantly larger models on tasks such as tool use and instruction following. LFM2.5-2.6B util…
-
Alibaba releases Qwen3.8 model with enhanced coding and agent skills · 1 source tracked
Alibaba has released its new Qwen3.8 large model, boasting significant improvements in programming and agent capabilities, positioning it among the top global models. The model supports a 1 million token context window and visual understanding, with strong performance in coding …
-
Huawei releases openPangu-2.0-Pro, first 505B model trained on Ascend NPUs
Huawei has released openPangu-2.0-Pro, a large language model with 505 billion parameters and a 512K context window. This model is notable for being the first frontier model fully trained on Ascend NPUs, marking a significant step in reducing reliance on NVIDIA hardware. The rel…
-
DeepSeek V4 Flash processes 8T tokens daily, highlighting cost-efficiency in agent economics
DeepSeek's V4 Flash model demonstrated significant efficiency by processing 8 trillion tokens in a single day on the OpenCode platform. This performance highlights a substantial price-performance gap compared to other models, with an 85x difference noted. The economics of AI age…
-
Kimi K3 launches with 1M context and always-on reasoning
Kimi K3, a new large language model, distinguishes itself by always engaging in reasoning without requiring user prompts or configuration. It boasts a 1 million token context window, significantly larger than competitors like Claude and GPT-4o, and can process entire codebases f…
-
Google DeepMind's DiffusionGemma achieves 1500 tokens/sec via discrete diffusion
Google DeepMind has released DiffusionGemma, an open-weight language model that utilizes discrete diffusion for text generation, offering significantly faster output speeds compared to traditional autoregressive models. While DiffusionGemma achieves approximately 1,500 output to…
-
AssemblyAI launches integrated Voice Agent API for simpler development
AssemblyAI has introduced a new Voice Agent API designed to simplify the development of voice-based AI agents. The API integrates speech-to-text (STT), large language model (LLM), and text-to-speech (TTS) functionalities into a single connection, aiming to reduce the complexity …
-
Alibaba launches Qwen3.8-Max with 2.4T parameters, focuses on agent harness reliability
Alibaba Group has launched Qwen3.8-Max, a 2.4 trillion parameter mixture-of-experts model with approximately 95 billion active parameters. The model supports multimodal input and is available via QwenCloud, with open weights planned for release. While Alibaba reports impressive …
-
AssemblyAI details AI scribe for therapy notes
AssemblyAI has detailed a method for constructing an AI scribe capable of generating progress notes from therapy sessions. The process involves accurate clinical transcription, distinguishing between the therapist and patient's speech, and utilizing a large language model (LLM) …
-
AssemblyAI touts Universal-3.5 Pro over Qwen3-ASR for production speech-to-text
AssemblyAI has compared its Universal-3.5 Pro speech-to-text model against Alibaba's Qwen3-ASR, highlighting the advantages of its proprietary solution for production environments. While Qwen3-ASR is recognized as a capable multilingual open-source model, AssemblyAI argues that …
-
AssemblyAI touts Universal-3.5 Pro over Whisper for production speech-to-text
AssemblyAI has published a comparison highlighting the advantages of its Universal-3.5 Pro model over OpenAI's Whisper Large-v3 for production speech-to-text applications. While Whisper is effective for clean audio and prototyping, AssemblyAI argues that its own models offer sup…
-
AssemblyAI Universal-3.5 Pro outperforms ElevenLabs Scribe v2 in key speech-to-text benchmarks
AssemblyAI has released a comparison of its Universal-3.5 Pro model against ElevenLabs' Scribe v2, highlighting Universal-3.5 Pro's superior performance in key areas for production systems. The comparison, conducted by AssemblyAI's Voice AI lead, indicates that Universal-3.5 Pro…
-
Pixel-Native RAG system indexes visual documents using multimodal embeddings
This tutorial details the creation of a "Pixel-Native RAG" system for visual document indexing. The process involves rendering web pages and PDFs as images, segmenting them into tiles, and generating multimodal embeddings using models like SigLIP or Qwen3-VL. These embeddings ar…
-
DeepSeek leads global AI model calls; OpenAI's Astra shows math prowess; Alibaba launches Qwen3.8
DeepSeek has risen to the top globally in AI model call volume, with Chinese AI models collectively handling a significant portion of global usage. Meanwhile, OpenAI's next-generation model, Astra, has demonstrated impressive capabilities by solving complex mathematical problems…
-
DeepSeek claims top global spot for AI model call volume
DeepSeek has achieved the top position globally in terms of API call volume, surpassing other major AI providers. This surge in usage highlights the growing demand for advanced AI models and DeepSeek's increasing prominence in the field. The news also touches upon a reduction in…
-
MiniMax-H3 omni-modal system ported to Apple Silicon via MLX
A new open-source Python package, MiniMax-H3-MLX, has been released to enable the MiniMax-H3 omni-modal generative system to run on Apple Silicon. This system can process text, images, audio, and video to generate up to 15-second video clips with audio. The package was successfu…
-
TokenMizer adds persistent memory to LLMs via graph database
TokenMizer is a new tool designed to give large language models persistent memory across conversation sessions. Unlike traditional methods that stuff more history into the context window, TokenMizer acts as a proxy that analyzes and stores conversation data in a graph database. …
-
Claude Code Artifacts offer new sharing options, complementing independent preview URLs
Anthropic's Claude Code now supports native Artifacts, offering a new way to share AI-generated content directly from a session. This feature is ideal for single HTML or Markdown pages that remain within the Claude session's scope and constraints. For more complex outputs like f…
-
Kandinsky 5.0 vs. GPT Image 2 & Nano Banana 2: Russian prompt understanding compared
A comparison of text-to-image models reveals that while GPT Image 2 and Nano Banana 2 lead in general image generation based on blind human ratings, they face accessibility issues for users in Russia. Kandinsky 5.0, developed by Sber AI, is highlighted for its native understandi…
-
Tencent Hunyuan releases Hy ASR 3.0 preview with enhanced context understanding
Tencent Hunyuan has released Hy ASR 3.0 preview, a new speech recognition model that integrates advanced language understanding capabilities from its Hy3 large language model. This update significantly improves accuracy, context awareness, robustness across various scenarios, an…