Urgent.News

What's breaking now, across thousands of outlets.

AI

Ollama GPU Scheduling: Running Inference and ComfyUI on One RTX Without OOM

Strategies for sharing a single RTX GPU between Ollama LLM inference and ComfyUI Stable Diffusion on the same homelab machine. The VRAM Reality Check: What Actually Fits on a 16GB RTX 5060 Ti Understanding your hardware limits is the first step to successful GPU sharing: Ollama VRAM Consumption (Approximate) Qwen2.5-Coder:14b (Q4_K_M): ~7.5 GB VRAM DeepSeek-R1:14b (Q4_K_M): ~7.5 GB VRAM…

A single RTX 5060 Ti GPU with 16GB of VRAM can struggle to run Ollama's 14b model alongside Stable Diffusion's SDXL 1024x1024 model simultaneously, as their combined VRAM requirements exceed the available memory. However, pairing Ollama 14b with SD 1.5 or SD XL 512x512, and adding ControlNet units, leaves enough VRAM for the system to function without running out of memory.

Various strategies exist for managing GPU access between Ollama and ComfyUI, such as allocating specific time slots to each application or granting GPU access based on priority levels. The time-based scheduling approach divides the day into time windows for each application, while the priority-based system allows higher priority tasks like interactive chat or generating images to access the GPU more quickly.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Sunday 11 October →