Jev After Eight Days of Independent Tests: Level With Mid-Price LLMs, Behind the Frontier
This article provides a review of the independent evidence on TypeSafe's Jev, the open models built to replace it, and the prior art behind both, as of September 23, 2026. Every figure below is traced to a primary source, and re-scored from committed per-item outputs wherever the author published them. This is a snapshot eight days after launch, and the arXiv preprints it cites are days old and…
Here is a summary of the provided information about Jev, an open model from TypeSafe AI:
Jev is a hosted model that answers typed questions about a piece of text. It offers three main functions: choice (select one option from a list), score (place the text on an ordered scale), and noul (return the probability that a statement is true). It writes no text and is priced at $0.042 per million input tokens. Access is through a waitlist, OpenRouter, and Vercel's AI Gateway.
The model's accuracy is comparable to mid-price LLMs and lags behind the frontier by 6.5 to 11.5 points. Calibration error for Jev is 0.161. It performs well in binary and few-class decisions, such as spam detection and code monitoring for backdoors. Jev also excels in schema compliance, with zero invalid answers across 23,703 decision-model calls.
However, it is behind the best-performing LLMs in several tasks, including the social-science annotation suite with 7,977 human-labelled items and banking tasks. The model's performance is consistent across different studies and comparisons.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.