{
  "id": 11276880,
  "title": "I benchmarked Cloudflare's new open decision model against the hosted API it's trying to replace",
  "url": "https://urgent.news/2026/10/01/i-benchmarked-cloudflares-new-open-decision-model-against-the-hosted",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-01T20:45:14.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/prodbymarcu/i-benchmarked-cloudflares-new-open-decision-model-against-the-hosted-api-its-trying-to-replace-2ded"
  },
  "original_language": "en",
  "account": "Cloudflare unveiled Clef, an open-weights decision model, designed to choose, rank, and gate decisions instead of generating text. This model comes in two variants: a 27B flagship and a 9B Flash, both constructed on Qwen backbones and equipped with a joint schema head that outputs probabilities over typed questions. The main selling point is that Clef serves as a drop-in replacement for TypeSafe's Jev API, which many agent builders currently pay per call for. The author tested Clef against Jev, a paid API, on their own system to assess performance and cost differences. Clef-Flash 9B ran on an RTX 3090, while Jev-1.13.0 was accessed via its hosted API. The author used 42 labeled decisions from an automated coding agent running overnight, categorized into three task families: computer-use action choices, subagent supervision calls, and message triage routes. Results showed that Jev achieved 71.4% accuracy, while Clef-Flash scored 66.7%. However, on computer-use action choices, both models performed perfectly, including handling safety traps. In terms of latency, Jev was faster with a p50 of 225ms, while Clef-Flash had a p50 of 315ms on the local machine. The author concluded that local deployment of Clef provides cost savings and privacy benefits, as data never leaves the user's device. Additionally, the 9B model offers a cost-effective upgrade path for those seeking better local decision-making capabilities.",
  "summary": "Cloudflare released Clef today: an open-weights \"decision model\" built to pick, rank, and gate instead of generating text. There's a 27B flagship and a 9B Flash variant, both built on Qwen backbones with a joint schema head that outputs calibrated probabilities over typed questions. The pitch is that it's a drop-in alternative to TypeSafe's Jev API, which is what a lot of agent builders currently…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}