A user on Reddit is seeking advice on the best local large language model for agentic coding tasks, given specific hardware constraints. They have access to a workstation with a substantial amount of RAM (256GB) but limited VRAM across two GPUs (12GB and 20GB). The user is considering using llama.cpp and is debating between a quantized version of Qwen 3.8 27B that fits within the VRAM or a larger model that leverages the high RAM. They are also inquiring about the practical utility of their hardware for generating useful outputs within an 8-hour workday. AI
IMPACT Guidance for users on optimizing local LLM deployment for coding tasks with specific hardware limitations.
RANK_REASON User query seeking advice on LLM deployment given hardware constraints.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →