PulseAugur
实时 00:06:58
简报 · 2026-08-04

AI 新闻 —— August 4, 2026

PulseAugur 当天浮现的 20 条头条故事 —— 综合实验室、论文及开发者社区的信号进行排序。

  1. SIGNIFICANT · · 100

    LiquidAI releases LFM2.5-2.6B for efficient local agent deployment

    LiquidAI has released LFM2.5-2.6B, a new language model designed for efficient local agent deployment. Despite its small size, the model demonstrates competitive performance against significantly larger models on tasks such as tool use and instruction following. LFM2.5-2.6B util…

  2. SIGNIFICANT · · 100

    Alibaba releases Qwen3.8 model with enhanced coding and agent skills · 1 source tracked

    Alibaba has released its new Qwen3.8 large model, boasting significant improvements in programming and agent capabilities, positioning it among the top global models. The model supports a 1 million token context window and visual understanding, with strong performance in coding …

  3. SIGNIFICANT · · 100

    Huawei releases openPangu-2.0-Pro, first 505B model trained on Ascend NPUs

    Huawei has released openPangu-2.0-Pro, a large language model with 505 billion parameters and a 512K context window. This model is notable for being the first frontier model fully trained on Ascend NPUs, marking a significant step in reducing reliance on NVIDIA hardware. The rel…

  4. SIGNIFICANT · · 100

    DeepSeek V4 Flash processes 8T tokens daily, highlighting cost-efficiency in agent economics

    DeepSeek's V4 Flash model demonstrated significant efficiency by processing 8 trillion tokens in a single day on the OpenCode platform. This performance highlights a substantial price-performance gap compared to other models, with an 85x difference noted. The economics of AI age…

  5. SIGNIFICANT · · 98

    Kimi K3 launches with 1M context and always-on reasoning

    Kimi K3, a new large language model, distinguishes itself by always engaging in reasoning without requiring user prompts or configuration. It boasts a 1 million token context window, significantly larger than competitors like Claude and GPT-4o, and can process entire codebases f…

  6. SIGNIFICANT · · 90

    Google DeepMind's DiffusionGemma achieves 1500 tokens/sec via discrete diffusion

    Google DeepMind has released DiffusionGemma, an open-weight language model that utilizes discrete diffusion for text generation, offering significantly faster output speeds compared to traditional autoregressive models. While DiffusionGemma achieves approximately 1,500 output to…

  7. TOOL · · 86

    AssemblyAI launches integrated Voice Agent API for simpler development

    AssemblyAI has introduced a new Voice Agent API designed to simplify the development of voice-based AI agents. The API integrates speech-to-text (STT), large language model (LLM), and text-to-speech (TTS) functionalities into a single connection, aiming to reduce the complexity …

  8. SIGNIFICANT · · 86

    Alibaba launches Qwen3.8-Max with 2.4T parameters, focuses on agent harness reliability

    Alibaba Group has launched Qwen3.8-Max, a 2.4 trillion parameter mixture-of-experts model with approximately 95 billion active parameters. The model supports multimodal input and is available via QwenCloud, with open weights planned for release. While Alibaba reports impressive …

  9. TOOL · · 85

    AssemblyAI details AI scribe for therapy notes

    AssemblyAI has detailed a method for constructing an AI scribe capable of generating progress notes from therapy sessions. The process involves accurate clinical transcription, distinguishing between the therapist and patient's speech, and utilizing a large language model (LLM) …

  10. TOOL · · 83

    AssemblyAI touts Universal-3.5 Pro over Qwen3-ASR for production speech-to-text

    AssemblyAI has compared its Universal-3.5 Pro speech-to-text model against Alibaba's Qwen3-ASR, highlighting the advantages of its proprietary solution for production environments. While Qwen3-ASR is recognized as a capable multilingual open-source model, AssemblyAI argues that …

  11. TOOL · · 83

    AssemblyAI touts Universal-3.5 Pro over Whisper for production speech-to-text

    AssemblyAI has published a comparison highlighting the advantages of its Universal-3.5 Pro model over OpenAI's Whisper Large-v3 for production speech-to-text applications. While Whisper is effective for clean audio and prototyping, AssemblyAI argues that its own models offer sup…

  12. TOOL · · 78

    AssemblyAI Universal-3.5 Pro outperforms ElevenLabs Scribe v2 in key speech-to-text benchmarks

    AssemblyAI has released a comparison of its Universal-3.5 Pro model against ElevenLabs' Scribe v2, highlighting Universal-3.5 Pro's superior performance in key areas for production systems. The comparison, conducted by AssemblyAI's Voice AI lead, indicates that Universal-3.5 Pro…

  13. TOOL · · 76

    Pixel-Native RAG system indexes visual documents using multimodal embeddings

    This tutorial details the creation of a "Pixel-Native RAG" system for visual document indexing. The process involves rendering web pages and PDFs as images, segmenting them into tiles, and generating multimodal embeddings using models like SigLIP or Qwen3-VL. These embeddings ar…

  14. RESEARCH · · 76

    DeepSeek leads global AI model calls; OpenAI's Astra shows math prowess; Alibaba launches Qwen3.8

    DeepSeek has risen to the top globally in AI model call volume, with Chinese AI models collectively handling a significant portion of global usage. Meanwhile, OpenAI's next-generation model, Astra, has demonstrated impressive capabilities by solving complex mathematical problems…

  15. RESEARCH · · 75

    DeepSeek claims top global spot for AI model call volume

    DeepSeek has achieved the top position globally in terms of API call volume, surpassing other major AI providers. This surge in usage highlights the growing demand for advanced AI models and DeepSeek's increasing prominence in the field. The news also touches upon a reduction in…

  16. TOOL · · 70

    MiniMax-H3 omni-modal system ported to Apple Silicon via MLX

    A new open-source Python package, MiniMax-H3-MLX, has been released to enable the MiniMax-H3 omni-modal generative system to run on Apple Silicon. This system can process text, images, audio, and video to generate up to 15-second video clips with audio. The package was successfu…

  17. TOOL · · 68

    TokenMizer adds persistent memory to LLMs via graph database

    TokenMizer is a new tool designed to give large language models persistent memory across conversation sessions. Unlike traditional methods that stuff more history into the context window, TokenMizer acts as a proxy that analyzes and stores conversation data in a graph database. …

  18. TOOL · · 68

    Claude Code Artifacts offer new sharing options, complementing independent preview URLs

    Anthropic's Claude Code now supports native Artifacts, offering a new way to share AI-generated content directly from a session. This feature is ideal for single HTML or Markdown pages that remain within the Claude session's scope and constraints. For more complex outputs like f…

  19. TOOL · · 67

    Kandinsky 5.0 vs. GPT Image 2 & Nano Banana 2: Russian prompt understanding compared

    A comparison of text-to-image models reveals that while GPT Image 2 and Nano Banana 2 lead in general image generation based on blind human ratings, they face accessibility issues for users in Russia. Kandinsky 5.0, developed by Sber AI, is highlighted for its native understandi…

  20. FRONTIER RELEASE · · 67

    Tencent Hunyuan releases Hy ASR 3.0 preview with enhanced context understanding

    Tencent Hunyuan has released Hy ASR 3.0 preview, a new speech recognition model that integrates advanced language understanding capabilities from its Hy3 large language model. This update significantly improves accuracy, context awareness, robustness across various scenarios, an…