Homelab census: 41 containers, one 6 GB GPU, and where my AI agents run
Today I asked my LLM box what it was doing, and it told me it had a 27-billion-parameter model loaded. Total footprint 18.3 GB. Amount of that on the graphics card: 0.5 GB. So the 27B was "running on the GPU" in the same sense that I am running a marathon when I walk to the shop. The other 17.8 GB sat in system RAM and did its arithmetic on the CPU, one patient token at a time. That seemed like a…
On 24 September, a comprehensive census was taken of a homelab setup consisting of four boxes. Across these machines, 41 Docker containers and 9 LXC containers were found to be running. One 6 GB graphics card played a pivotal role in determining where AI tasks were processed. The Proxmox node 1, equipped with 4 cores and 8 GB RAM, hosted 16 Docker containers.
Meanwhile, Proxmox node 2, with 4 cores and 16 GB RAM, housed 11 containers. The ZimaBlade, a fanless machine with 2 cores and 16 GB RAM, ran 14 Docker containers. The i7-10875H system, boasting 16 threads, 31 GB RAM, and a 6 GB VRAM RTX 2060 graphics card, hosted Ollama with 21 model tags and ComfyUI, but no containers. This configuration allowed the AI agents to operate efficiently, with the 6 GB VRAM card making the majority of routing decisions for the AI tasks.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.