PulseAugur
实时 22:58:47
English(EN) v0.33.3-rc2: gemma4: image and audio input support

Ollama 为 Gemma4 模型添加了图像和音频支持

Ollama 发布了 v0.33.3-rc2 版本,为 Gemma4 模型引入了图像和音频输入支持。此次更新利用 MLX 引擎处理多模态输入,其中图像由 Transformer 塔和统一嵌入器架构处理,音频则通过一个接受 WAV 字节和音频转录的专用输入通道处理。系统会智能地将模态路由到兼容的检查点,确保没有特定音频或视觉功能的模型仅作为文本模型运行并拒绝不支持的输入类型。 AI

影响 增强了本地 AI 模型部署的多模态能力。

排序理由 这是用于集成 AI 模型的工具的软件发布,而不是前沿模型本身的发布。

在 Ollama — Releases 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Ollama 为 Gemma4 模型添加了图像和音频支持

本文如何被排名

Signal score
33 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是用于集成 AI 模型的工具的软件发布,而不是前沿模型本身的发布。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. Ollama — Releases TIER_1 English(EN) · jessegross ·

    v0.33.3-rc2: gemma4: 支持图像和音频输入

    <p>Safetensors gemma4 imports served by the MLX engine now answer image<br /> and audio chats. Images run through both vision architectures: the<br /> transformer tower (26B, 31B, e-series) and the 12B's encoder-free<br /> unified embedder. Audio arrives through the same intake t…