PulseAugur
EN
LIVE 22:30:02

Self-hosted server achieves 63 tokens/sec with Qwen3.6 LLM on budget hardware

A user is reporting positive results running the Qwen3.6-35B-A3B-UD-Q4_K_M large language model on a solar-powered, self-hosted server. The setup, utilizing older DDR3 hardware with three 2080 Ti GPUs, achieved approximately 63 tokens per second with no per-token cost. This demonstrates a budget-friendly and efficient approach to running advanced AI models locally. AI

IMPACT Demonstrates cost-effective local LLM deployment on older hardware.

RANK_REASON User report on running an LLM on self-hosted hardware.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Self-hosted server achieves 63 tokens/sec with Qwen3.6 LLM on budget hardware

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    The solar powered LLM server seems to be happy with Qwen3.6-35B-A3B-UD-Q4_K_M. It gets around 63 tokens per second and costs me $0.00 per token. And the server

    The solar powered LLM server seems to be happy with Qwen3.6-35B-A3B-UD-Q4_K_M. It gets around 63 tokens per second and costs me $0.00 per token. And the server itself is an ancient DDR3 platform with 3 2080 ti GPUs in it! # solarai # decentralized # smarthome # homeassistant # he…