This article provides practical strategies for users to run both large language models (LLMs) via Ollama and image generation models like Stable Diffusion using ComfyUI on a single RTX GPU without encountering Out-of-Memory (OOM) errors. It details the approximate VRAM consumption for various LLMs (such as Qwen2.5 Coder 14B and Llama3.1 8B) and Stable Diffusion models (SDXL, SD 1.5 with ControlNet), offering guidance on what combinations are feasible on a 16GB RTX 5060 Ti. The author proposes two primary scheduling methods: time-based scheduling, which allocates specific time windows to each application, and priority-based access, where applications request GPU time based on predefined priority levels. AI
IMPACT Enables users to run multiple AI models concurrently on limited hardware, maximizing resource utilization.
RANK_REASON The article provides practical advice and technical strategies for optimizing the use of existing hardware for AI-related tasks, rather than announcing a new product or research breakthrough.
- ComfyUI
- ControlNet
- DeepSeek-R1:14b
- Llama3.1:8b
- Nomic-Embed-Text
- Ollama
- Qwen2.5 Coder 14B
- RTX 5060 Ti
- SD Negeri 1.5 Belimbing
- SDXL
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →