{
  "id": 12156366,
  "title": "Clockwork.io bags $31M in funding to keep AI inference and training workloads running like … clockwork",
  "url": "https://urgent.news/2026/10/05/clockwork-io-bags-31m-in-funding-to-keep-ai-inference-and-training",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-05T13:00:22.000Z",
  "source": {
    "name": "SiliconANGLE",
    "slug": "siliconangle",
    "url": "https://siliconangle.com/2026/10/05/clockwork-io-bags-31m-in-funding-to-keep-ai-inference-and-training-workloads-running-like-clockwork/"
  },
  "original_language": "en",
  "account": "Clockwork Systems Inc., a data center infrastructure startup, has secured $31 million in new funding to enhance the efficiency of artificial intelligence chip clusters. The round was co-led by Seligman Ventures, Wing Ventures, and Premji Invest, with existing backers New Enterprise Associates and e& Capital also participating. This brings Clockwork's total funding raised to $73 million.\n\nAs AI compute costs rise, organizations are shifting focus from acquiring raw graphics processing unit (GPU) resources to optimizing existing clusters. Large distributed AI workloads in GPU clusters often experience hardware failures, which can take up to 90 minutes to recover from and result in thousands of healthy GPUs sitting idle. To address this, Clockwork developed TorchSnap, a feature that captures multinode snapshots of distributed AI inference workloads across each node within a cluster, enabling quick restarts without requiring developer code modifications.\n\nThe new software layer sits between GPUs and running AI workloads, synchronizing GPU clusters and delivering nanosecond-accurate telemetry to identify failures before they lead to full cluster restarts. Clockwork's Chief Executive Suresh Vasudevan notes that fault tolerance is crucial for maintaining productivity during GPU failures. The company's existing LinkPass network failover tool and TorchPass GPU migration software are complemented by this new feature.\n\nWith the launch of TorchSnap, Clockwork is further enhancing its cluster resilience capabilities. The technology has already seen adoption across public cloud infrastructure providers, neoclouds, and enterprise fleets. Microsoft's LinkedIn, for instance, deployed Clockwork's LinkPass functionality to eliminate thousands of GPU-hours of downtime monthly. Together AI integrates TorchPass as a service on its GPU clusters, while WhiteFiber leverages Clockwork's technology to audit and validate cluster reliability before new AI workloads enter production.",
  "summary": "Clockwork Systems Inc., the data center infrastructure startup that helps to maximize the efficiency of artificial intelligence chip clusters, has raised $31 million in fresh funding and announced the launch of a new feature called TorchSnap that helps to minimize wasted compute. Today’s round was co-led by Seligman Ventures, Wing Ventures and Premji Invest and […] The post Clockwork.io bags $31M…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}