PulseAugur
EN
LIVE 19:26:17

Swiss AI's Apertus 1.5 8B runs locally on 8GB GPU

Swiss AI has developed Apertus 1.5 8B, a language model capable of running entirely on a laptop with an 8GB GPU, specifically an RTX PRO 1000. This model achieves this by using W4 quantization for its embedding and output head, allowing for local operation without cloud reliance or CPU offloading. The model demonstrates solid chat capabilities and achieves a speed of 14-22 tokens/s, though its tool-calling functionality requires further improvement. AI

IMPACT Enables running advanced language models on consumer-grade hardware, potentially democratizing AI access and local deployment.

RANK_REASON The cluster describes the release and technical details of a new open-source language model, fitting the research category. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Swiss AI's Apertus 1.5 8B runs locally on 8GB GPU

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Shrinking Apertus 1.5 8B for an 8 GB Laptop GPU. Swiss AI's Apertus 1.5 8B fully local on an RTX PRO 1000, 8 GB VRAM, no cloud, no CPU offloading, 32K context.

    Shrinking Apertus 1.5 8B for an 8 GB Laptop GPU. Swiss AI's Apertus 1.5 8B fully local on an RTX PRO 1000, 8 GB VRAM, no cloud, no CPU offloading, 32K context. W4 quantization of embedding and output head. 14-22 tokens/s, up from 5. Chat: solid. Tool calling: needs improvement. M…