How Fast Can a 421M-Parameter Decision Model Run? I Benchmarked Laya Across NVIDIA GPUs
One H100 NVL. A 421M-parameter decision model. 15.1 million decisions per day while staying inside a p99 ≤ 130 ms latency budget. That number sounds impressive—but raw throughput is the easy number to publish. The useful question is harder: How many typed decisions can one GPU sustain when tail latency, correctness, and cost all matter? I built an independent, fully reproducible benchmark to…
We haven't written up this one. Dev.to has the full story — the link below goes straight to it.