Urgent.News

What's breaking now, across thousands of outlets.

AI

10M Batch LLM Inference at $0 Cloud Cost: O(1) Memory Clamped Architecture

High cloud API costs and Out-Of-Memory (OOM) failures in large-scale data pipelines are architectural defects, not hardware constraints. This report documents the performance of a 10,000,000-record batch LLM inference workload executed locally on an HP Z4 G4 workstation (Intel Xeon W-2155, 128GB ECC DDR4 RAM, NVIDIA RTX 5060 Ti 16GB VRAM, NVMe PCIe Gen4 SSD). Verified Metrics Success Rate :…

Large-scale language models (LLMs) often encounter high cloud API costs and out-of-memory (OOM) failures during data processing. This report examines the performance of a 10 million-record batch LLM inference workload executed on a local HP Z4 G4 workstation. The HP workstation features an Intel Xeon W-2155 processor, 128GB of ECC DDR4 RAM, NVIDIA RTX 5060 Ti GPU with 16GB VRAM, and an NVMe PCIe Gen4 SSD.

The results demonstrate a 100% success rate with an average throughput of approximately 254.9 requests per second, completing the task in 10.9 hours. Memory usage was clamped between 6.72GB and 9.4GB, maintaining a constant space complexity (O(1)). The database state was a 3.4GB SQLite WAL file, with integrity verified through PRAGMA checks. No compute costs were incurred as the entire operation was performed locally.

The technical implementation involved three key strategies. Firstly, substituting the glibc malloc allocator with jemalloc via LD_PRELOAD significantly reduced heap fragmentation under allocation and deallocation cycles. By using LD_PRELOAD with active background thread decay (MALLOC_CONF= background_thread:true, dirty_decay_ms:2000, muzzy_decay_ms:2000), the process's resident set size (RSS) remained stable, preventing growth and maintaining a consistent RAM footprint.

Secondly, Polars disk-backed Parquet streaming was employed to avoid the O(N) spatial memory complexity associated with in-memory array construction. By streaming chunked outputs directly to disk using Polars (collect(engine=streaming) and sink_parquet), memory usage remained independent of the row count.

Lastly, atomic 2-phase transaction logging using aiosqlite and WAL (Write-Ahead Logging) ensured state recovery and record tracking. Explicit BEGIN IMMEDIATE transaction blocks were executed every 50,000 items, with periodic WAL truncation to guarantee consistency and integrity in the event of a system crash. Implementing this setup on local, air-gapped infrastructure, with network disabled ( --network none ), further protected against regulatory risks associated with transmitting unmasked enterprise datasets to external APIs, such as those outlined by HIPAA and GDPR. All code and the repository are available at https://github.com/Matsubara-CEO.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

What happens when you reject an AI agent's work in Bees

Most agent tools give you two choices when the output is wrong. Start over, or edit the prompt and hope. We wanted something in between for Bees, the open source desktop app we build that runs a team…

  • Bees uses a reviewer agent to check AI work against goals and evidence.
  • Rejecting AI work requires a reason, scoped by context (single goal vs scheduled job).
  • Reviewer's decision doesn't alter base agent for organization or its functionality.

I asked 63 models the same 76 questions, with and without web search

The 2026 standard deduction for a single filer is $16,100. I asked 63 models what it was. Asked from memory, 8 of 61 got it right. 24 invented a number, three of them landing on $8,300.

  • Only 8 out of 63 models answered 2026 standard deduction correctly from memory
  • 24 models fabricated numbers, including $8,300, when given from memory
  • 14 out of 15 models with web search capability provided correct answers

More from Thursday 8 October →