A user experienced a significant performance drop in their local LLM setup after updating LM Studio, with their RTX 3090's processing speed decreasing from 60-70 tokens/sec to 12-18 tokens/sec. The issue was traced to only 54 layers being offloaded to the GPU, despite settings indicating maximum offload. The problem was resolved by switching the runtime from CUDA12 to either CUDA or Vulkan, with CUDA offering slightly better performance. The user noted that CUDA12 appeared to be forcing CPU offload. AI
IMPACT A software update for local LLM deployment tools can unexpectedly degrade performance, requiring users to adjust runtime settings.
RANK_REASON User-reported issue with a specific software update affecting hardware performance.
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →