{
  "id": 10566863,
  "title": "When GPU Resources Run Dry, It Still Looks Like Everything Is Working: A Tale of Two AI Competing for VRAM",
  "url": "https://urgent.news/2026/09/29/when-gpu-resources-run-dry-it-still-looks-like-everything-is-working",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-29T00:19:01.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/orca_forge/when-gpu-resources-run-dry-it-still-looks-like-everything-is-working-a-tale-of-two-ai-competing-ejc"
  },
  "original_language": "en",
  "account": "Two AI systems were running concurrently on a single GPU with 12GB of VRAM. One task involved generating lip-sync caricature videos, while the other estimated hand poses from real-life video footage. After an hour and a minute of waiting, the lip-sync task remained pending. At first glance, everything appeared to be functioning normally as the GPU was fully utilized and memory was near its limit. However, the lip-sync generation was experiencing significant delays, taking around 1 hour for 300 frames that would typically take 10-12 minutes.\n\nThe hand estimation system, on the other hand, was returning a result of \"no hands detected\" in 431 seconds per frame, which is 80,000 times slower than normal. Both tasks were fully utilizing the GPU, but the lip-sync generation was clearly struggling due to the lack of available VRAM. When VRAM was running low, the hand estimation system would return a seemingly normal result of zero detections, making it difficult to identify the issue simply by looking at the metrics.\n\nTo distinguish between normal and abnormal performance, it is recommended to compare processing times per unit rather than relying on display status. By sharing the processing time per unit and the start time with the other party, it becomes possible to determine when a task has become stuck. Establishing clear communication and guidelines for GPU usage can help prevent such issues from arising in the first place.",
  "summary": "📝 Originally published (in Japanese) at forge.workstyle.tech . On the same machine (1 GPU, 12GB VRAM), two Claude Code sessions were running separate tasks concurrently. This one: Generating lip-sync caricature videos using InfiniteTalk (10–12 minutes per video) The other: Estimating hand poses from real-life video footage (WiLoR) One day, the lip-sync task for 6 videos was still pending after…",
  "key_points": [
    "Two AI systems running on single GPU with 12GB VRAM",
    "Lip-sync generation delayed by 1 hour for 300 frames",
    "Hand estimation system returned 'no hands detected' 80,000x slower"
  ],
  "editors_take": "Low VRAM allows one AI system to masquerade as functioning normally while severely delaying another, highlighting the need for nuanced performance metrics and communication guidelines to prevent such issues.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}