A user on Reddit's r/LocalLLaMA subreddit is seeking assistance with configuring llama.cpp to effectively utilize multiple GPUs for running large language models. They are experiencing issues with performance when attempting to split a 120GB Deepseek4 flash model across a Blackwell 5000 (48GB) and a 3090 (24GB) GPU, noting that performance is not improving as expected and sometimes degrades. The user has detailed their setup, the model size, and various command-line flag combinations they have tried, including different split modes and manual layer assignments, but has not achieved optimal multi-GPU performance. AI
IMPACT Optimizing multi-GPU setups for local LLM inference can improve performance and accessibility for users running large models on consumer hardware.
RANK_REASON User query about optimizing existing software for hardware configuration.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →