My agent's p50 was 29s. Its p95 was 182s. That ratio decided the product.
I have 67 timed turns against a live LLM agent, captured in a single batch run and written to disk: seconds p50 29.2 p95 182.0 max 210.3 n 67 completed turns The median says background job, that's fine . The p95 says you may never put a human in front of this . Those are not two performance notes. They are a product spec — and I found that out the expensive way, by designing the product first and…
A live LLM agent was tested with 67 timed turns, and its performance was measured at p50 and p95 times. The p50 was 29 seconds, while the p95 was 182 seconds. This ratio became the deciding factor for the product. The product aimed to recover members who canceled, but their reasons for leaving were never recorded or addressed. The creator implemented an agent to provide a warm exit interview, condition matching, and a win-back draft.
The agent resolved 14 departures out of 54 captured judgements, showing that the keyword baseline was ineffective in linking related keywords. The architecture was redesigned to avoid synchronous fan-out and limit the live path, resulting in a more efficient system.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.