Comparing JEV and LLMs on Structured Decision Tasks
A benchmark comparing JEV with frontier LLMs reveals why decision accuracy alone isn't enough when structured scores and probabilities must agree.
The author began by asserting that accuracy is more important than speed when making decisions. They built a benchmark to test this theory against JEV and frontier language models (LLMs) on standardized decision tasks like intent routing, moderation classification, and search relevance.
Their initial hypothesis was that JEV's speed advantage would hold, as the best LLMs outperformed it in accuracy. However, the results were more nuanced. While JEV struggled with intent routing accuracy (81.0% vs. 85.6% to 86.1% for the LLMs), it excelled in moderation (98.2%) and search relevance tasks (47.4%). This suggests that JEV can be very fast without sacrificing meaningful accuracy, as the specifics of the decision contract matter.
The benchmark revealed an unexpected failure mode: models could return valid JSON, a plausible score, and a probability distribution that didn't align with the score. This showed that a model's output must be trustworthy - the score and probability distribution should agree, as downstream systems need to rely on one of these. The consistency of the score, probabilities, and structured fields was tested, and JEV achieved 100% consistency, compared to 99.8% for the best LLM.
In conclusion, the key takeaway is that trustworthiness, not just accuracy, is crucial for decision-making systems. The benchmark provided a reproducible framework to evaluate models based on quality (accuracy), consistency (agreement of score, probabilities, and structured fields), and operational performance (latency and token usage). The results challenge the notion that bounded decisions are too limiting, instead highlighting the hard problem of producing consistently trustworthy decisions.
Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.