Exclusive: Iterate.ai’s Lifeboat runs up to six times more AI agent sessions per GPU
Enterprise artificial intelligence software company Iterate Studio Inc. today launched Lifeboat, an inference engine for large language models that has confidential computing built in. Iterate.ai says the software fits two to six times as many concurrent AI agent sessions on each graphics processing unit. Lifeboat is seeking to take on a memory problem that agents […] The post Exclusive:…
Iterate.ai has introduced a new inference engine called Lifeboat designed to handle large language models with enhanced capabilities. Lifeboat stands out by fitting significantly more concurrent AI agent sessions on each graphics processing unit (GPU) than standard inference engines, with performance ranging from two to six times higher.
This technology addresses the memory challenge that arises as AI agents process increasingly complex tasks, which can exhaust shared hardware resources. Lifeboat achieves this by implementing fair scheduling and admission control, allocating GPU resources to each session proportionally to prevent any single agent from monopolizing the hardware.
The company also optimizes the key-value cache within the engine, which effectively doubles its capacity while maintaining full precision of model weights. This optimization is particularly beneficial for mixture-of-experts models, where only the relevant experts are loaded for a given session. The safety and security of data are ensured through compartmentalized execution environments, with each session running in its own security capsule that includes token budgets, filtering, and sandboxed execution.
During internal testing, Lifeboat demonstrated its capability on an Nvidia RTX PRO 6000 Blackwell GPU, managing 2,048 concurrent sessions without compromising request completion times. When the optimizations were disabled, the engine's capacity dropped to half, handling only 4,965 tokens per second. Lifeboat's performance is further distinguished by its ability to maintain a 99th-percentile time to the first token at just 1.5 seconds during a memory-pressure test with 128 sessions sending 18,000-token requests, contrasting with 189 seconds on the baseline system.
The company emphasizes that enterprises should assess the full potential of their existing GPUs for AI workloads before making additional investments. Lifeboat offers a Confidential Computing edition that includes hardware attestation, ensuring that model weights remain encrypted and secure throughout the execution process. This premium version is available at $499.99 per month, with a free Developer License available for non-commercial use on up to two inference servers.
The overall launch has been met with strong interest, drawing thousands of downloads within days of its release.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.