A new inference runtime called Eider has been developed for NVIDIA DGX Spark and similar hardware, built from scratch in Rust and CUDA. Eider is designed to leverage the NVFP4 capabilities of SM121 GPUs and does not rely on existing libraries like llama.cpp or vLLM. It supports various models including Qwen, Agents-A1, Step-3.7-Flash, Gemma 4, and NVIDIA's Nemotron 3, with a notable feature for paging experts from disk to fit larger models in memory. The runtime includes an OpenAI-compatible server and aims to provide a faster startup for local coding tasks by focusing on SM121 optimizations. AI
IMPACT This new runtime could enable faster local AI model inference on specific NVIDIA hardware, potentially improving developer workflows.
RANK_REASON The item describes a new software tool for running AI models on specific hardware, not a core AI model release or research.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →