PulseAugur
EN
LIVE 05:27:57
Deutsch(DE) RT @ollama: Kimi K3 ist jetzt auf Ollamas Cloud verfügbar. Um es mit Claude Code zu nutzen, führen Sie aus: ollama launch claude --model kimi-k3:cloud Derzeit e

Moonshot AI releases Kimi K3 model; AMD integrates GPU optimizations

Moonshot AI has released its Kimi K3 model, a 2.8T MoE model featuring native visual understanding and a 1 million token context window. The model is available on Ollama's cloud, though it currently requires a premium subscription and incurs additional usage credits. Concurrently, AMD has introduced ATOM, a plugin for vLLM that integrates native Instinct GPU optimizations for models like Kimi K3 and Qwen3.5, aiming to enhance performance without altering the vLLM API. AI

IMPACT Kimi K3's large context window and MoE architecture push the boundaries of model capabilities, while AMD's ATOM plugin aims to optimize inference for such models on their hardware.

RANK_REASON Cluster includes announcement of new model Kimi K3 from Moonshot AI, with technical details and availability on Ollama.

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

Moonshot AI releases Kimi K3 model; AMD integrates GPU optimizations

COVERAGE [3]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @ollama: Kimi K3 is now available on Ollama's Cloud. To use it with Claude Code, run: ollama launch claude --model kimi-k3:cloud Currently e

    RT @ollama: Kimi K3 ist jetzt auf Ollamas Cloud verfügbar. Um es mit Claude Code zu nutzen, führen Sie aus: ollama launch claude --model kimi-k3:cloud Derzeit erfordert Kimi K3 ein Pro- oder Max-Abonnement und verbraucht zusätzliche Nutzungsguthaben. Wir arbeiten schnell daran, d…

  2. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @AIatAMD: COMING TOMORROW to our deep dive into ATOM, AMD's high-performance plugin that brings native Instinct GPU optimizations directly into vLLM

    RT @AIatAMD: KOMMT MORGEN zu unserer tiefgehenden Einführung in ATOM, das Hochleistungs-Plugin von AMD, das native Instinct-GPU-Optimierungen direkt in den vLLM-Serving-Stack integriert — ohne Forking, ohne Neuschreibungen. ATOM bietet AMD-optimierte Attention-Backends (über AITE…

  3. Mastodon — mastodon.social TIER_1 Deutsch(DE) · [email protected] ·

    RT @LLMJunky: I will NOT download the Kimi K3 models. Primarily because I don't have 47 TB of VRAM to run them, and in 3 weeks something better

    RT @LLMJunky: Ich werde die Kimi K3-Modelle NICHT herunterladen. Vor allem, weil ich nicht 47 TB VRAM habe, um sie zu betreiben, und in 3 Wochen etwas Besseres verfügbar sein wird, lol. mehr auf Arint.info # AI # DeepLearning # KimiK3 # MachineLearning # VRAM # arint_info https:/…