Urgent.News

What's breaking now, across thousands of outlets.

AI

Jev: New frontier model 40-400x cheaper and 20-200x faster

For years, models have been remarkable at generating human-like chat, leaving many wondering why automation hasn't progressed further. As a former OpenAI researcher, I witnessed the development of instruction-following and conversational language models that became the foundation for ChatGPT. However, I realized there was a significant gap that needed to be addressed.

After two years of secrecy and overcoming numerous technical hurdles, it's my great honor to announce the release of TypeSafe AI's first System One Model: a groundbreaking class of frontier models designed to facilitate rapid, structured decision-making that software can utilize directly. Our innovative stack, centered around automation, boasts a novel model architecture, a parallel sampler for unparalleled efficiency, and a groundbreaking training method called Reinforcement Learning for Calibrated Decisions (RLCD).

Our inaugural public model, Jev, is now accessible in early release. Operating with a speed and efficiency that outpaces existing large language models (LLMs), Jev excels at structured outputs, eschewing string generation while eliminating the risk of hallucinations. As skeptics ourselves, we invite further scrutiny and provide comprehensive evidence to support our claims.

Our side-by-side demonstration highlights a crucial distinction between our models and LLMs: Jev outputs all probabilities concurrently, contrasting with LLMs' autoregressive token-by-token generation. While strings possess versatility and broad applicability, their cost results in a trade-off. By choosing to "give up" strings, we unlock a host of new advantages.

For individuals with access to TypeSafe AI, I present a simplified query showcasing the model's capabilities. The accompanying dataset is deliberately succinct and human-readable, emphasizing our model's superior sampling methodology. Observant readers will note a single discrepancy between Jev and GPT-5.6 Terra's output regarding "Churn likelihood level."

Given the ambiguity surrounding the actual answer, we opted to use GPT-5.6 Terra with default reasoning for this example, as it represents a comparable level of intelligence to Jev in average scenarios. An intriguing tidbit: a comparable demonstration inspired our decision to pursue System One Models wholeheartedly. We devised an innovative evaluation method to gauge AI's effectiveness within code.

Rather than focusing on binary classification or allowing the harness and model to influence each other (potentially enabling overfitting through harness engineering), we assume the existence of a correct compute graph embedded in code. We then utilize the predictions of the most advanced, powerful, and costliest external models as reference probabilities, ensuring an equitable comparison.

Our findings reveal that Jev surpasses the Pareto frontier by nearly two orders of magnitude compared to its contemporaries. To further substantiate our claims, we employ a distinct set of production workloads, which prove to be significantly more intricate and representative of genuine business automation requirements than the side-by-side demonstration.

These workflows are far more complex than the initial example, as they more accurately reflect the nuanced, domain-specific engineering required for true business automation. For a comprehensive exploration of our findings, including examples, disagreements, full queries, and each workflow, visit our workflow evaluations site. The impressive performance metrics highlighted on our homepage, such as the purported 193.6x speed increase and 444.6x cost reduction, are drawn from these workflows and are, in our estimation, conservative estimates of the actual gains achievable.

In constructing these workflows, we took great care to avoid selecting or creating examples specifically to showcase our model's strengths, as they were curated by our model capabilities team. While some bias may be present, we believe the average performance of GPT-6 Astra and Fable 5.1 serves as a reasonable reference point, albeit one that inadvertently privileges OpenAI and Anthropic's models.

We acknowledge the potential for underestimation, as our model and DeepSeek's models have yet to be thoroughly evaluated against this benchmark. Our System One LLM wrapper constrains LLMs to output structured decisions compatible with our API, enhancing decision-making accuracy and reliability. However, this approach often results in slower and more expensive computations compared to generating decisions without accompanying probabilities.

Ultimately, hallucinations and guaranteed type-safety are inextricably linked, and we firmly believe that the latter is a prerequisite for automation. Existing models, despite their intelligence, continue to produce hallucinations, making them unsuitable for automation-driven applications.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at typesafe.ai →

More in AI

More from Tuesday 15 September →