When GPU Resources Run Dry, It Still Looks Like Everything Is Working: A Tale of Two AI Competing for VRAM
📝 Originally published (in Japanese) at forge.workstyle.tech . On the same machine (1 GPU, 12GB VRAM), two Claude Code sessions were running separate tasks concurrently. This one: Generating lip-sync caricature videos using InfiniteTalk (10–12 minutes per video) The other: Estimating hand poses from real-life video footage (WiLoR) One day, the lip-sync task for 6 videos was still pending after…
Two AI systems were running concurrently on a single GPU with 12GB of VRAM. One task involved generating lip-sync caricature videos, while the other estimated hand poses from real-life video footage. After an hour and a minute of waiting, the lip-sync task remained pending. At first glance, everything appeared to be functioning normally as the GPU was fully utilized and memory was near its limit.
However, the lip-sync generation was experiencing significant delays, taking around 1 hour for 300 frames that would typically take 10-12 minutes.
The hand estimation system, on the other hand, was returning a result of "no hands detected" in 431 seconds per frame, which is 80,000 times slower than normal. Both tasks were fully utilizing the GPU, but the lip-sync generation was clearly struggling due to the lack of available VRAM. When VRAM was running low, the hand estimation system would return a seemingly normal result of zero detections, making it difficult to identify the issue simply by looking at the metrics.
To distinguish between normal and abnormal performance, it is recommended to compare processing times per unit rather than relying on display status. By sharing the processing time per unit and the start time with the other party, it becomes possible to determine when a task has become stuck. Establishing clear communication and guidelines for GPU usage can help prevent such issues from arising in the first place.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
