{
  "id": 12156365,
  "title": "Exclusive: Iterate.ai’s Lifeboat runs up to six times more AI agent sessions per GPU",
  "url": "https://urgent.news/2026/10/05/exclusive-iterate-ais-lifeboat-runs-up-to-six-times-more-ai-agent",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-05T13:00:47.000Z",
  "source": {
    "name": "SiliconANGLE",
    "slug": "siliconangle",
    "url": "https://siliconangle.com/2026/10/05/exclusive-iterate-ais-lifeboat-runs-up-to-six-times-more-ai-agent-sessions-per-gpu/"
  },
  "original_language": "en",
  "account": "Iterate.ai has introduced a new inference engine called Lifeboat designed to handle large language models with enhanced capabilities. Lifeboat stands out by fitting significantly more concurrent AI agent sessions on each graphics processing unit (GPU) than standard inference engines, with performance ranging from two to six times higher. This technology addresses the memory challenge that arises as AI agents process increasingly complex tasks, which can exhaust shared hardware resources. Lifeboat achieves this by implementing fair scheduling and admission control, allocating GPU resources to each session proportionally to prevent any single agent from monopolizing the hardware. The company also optimizes the key-value cache within the engine, which effectively doubles its capacity while maintaining full precision of model weights. This optimization is particularly beneficial for mixture-of-experts models, where only the relevant experts are loaded for a given session. The safety and security of data are ensured through compartmentalized execution environments, with each session running in its own security capsule that includes token budgets, filtering, and sandboxed execution. During internal testing, Lifeboat demonstrated its capability on an Nvidia RTX PRO 6000 Blackwell GPU, managing 2,048 concurrent sessions without compromising request completion times. When the optimizations were disabled, the engine's capacity dropped to half, handling only 4,965 tokens per second. Lifeboat's performance is further distinguished by its ability to maintain a 99th-percentile time to the first token at just 1.5 seconds during a memory-pressure test with 128 sessions sending 18,000-token requests, contrasting with 189 seconds on the baseline system. The company emphasizes that enterprises should assess the full potential of their existing GPUs for AI workloads before making additional investments. Lifeboat offers a Confidential Computing edition that includes hardware attestation, ensuring that model weights remain encrypted and secure throughout the execution process. This premium version is available at $499.99 per month, with a free Developer License available for non-commercial use on up to two inference servers. The overall launch has been met with strong interest, drawing thousands of downloads within days of its release.",
  "summary": "Enterprise artificial intelligence software company Iterate Studio Inc. today launched Lifeboat, an inference engine for large language models that has confidential computing built in. Iterate.ai says the software fits two to six times as many concurrent AI agent sessions on each graphics processing unit. Lifeboat is seeking to take on a memory problem that agents […] The post Exclusive:…",
  "key_points": [
    "Lifeboat inference engine handles 2-6x more AI agent sessions per GPU than standard engines.",
    "Implements fair scheduling and admission control to prevent GPU resource monopolization.",
    "Optimizes key-value cache, doubling capacity while preserving model weight precision."
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}