{
  "id": 13798179,
  "title": "Ollama GPU Scheduling: Running Inference and ComfyUI on One RTX Without OOM",
  "url": "https://urgent.news/2026/10/11/ollama-gpu-scheduling-running-inference-and-comfyui-on-one-rtx",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-11T21:00:00.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/ayraix/ollama-gpu-scheduling-running-inference-and-comfyui-on-one-rtx-without-oom-1j75"
  },
  "original_language": "en",
  "account": "A single RTX 5060 Ti GPU with 16GB of VRAM can struggle to run Ollama's 14b model alongside Stable Diffusion's SDXL 1024x1024 model simultaneously, as their combined VRAM requirements exceed the available memory. However, pairing Ollama 14b with SD 1.5 or SD XL 512x512, and adding ControlNet units, leaves enough VRAM for the system to function without running out of memory. Various strategies exist for managing GPU access between Ollama and ComfyUI, such as allocating specific time slots to each application or granting GPU access based on priority levels. The time-based scheduling approach divides the day into time windows for each application, while the priority-based system allows higher priority tasks like interactive chat or generating images to access the GPU more quickly.",
  "summary": "Strategies for sharing a single RTX GPU between Ollama LLM inference and ComfyUI Stable Diffusion on the same homelab machine. The VRAM Reality Check: What Actually Fits on a 16GB RTX 5060 Ti Understanding your hardware limits is the first step to successful GPU sharing: Ollama VRAM Consumption (Approximate) Qwen2.5-Coder:14b (Q4_K_M): ~7.5 GB VRAM DeepSeek-R1:14b (Q4_K_M): ~7.5 GB VRAM…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}