{
  "id": 10681794,
  "title": "Jeeves. Reasoning improves Jev-like decision models",
  "url": "https://urgent.news/2026/09/29/jeeves-reasoning-improves-jev-like-decision-models",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-29T11:13:54.000Z",
  "source": {
    "name": "Hacker News",
    "slug": "hacker-news",
    "url": "https://github.com/PostHog/jeeves"
  },
  "original_language": "en",
  "account": "Jeeves introduces a reasoning model to enhance decision-making capabilities of Jev-like classifiers. These classifiers are trained with SFT and CISPO but often rely on a reasoning model when their decision probabilities are inaccurate. Jeeves trains a Jev-like Qwen3.5-9B model using CISPO, resulting in improved performance on out-of-domain tasks and surpassing Jev in JevBench hard (public).\n\nThe Jev-like model utilizes a diffusion drafter and can be trained with a 2,560-token cap. The Kev-9B and Jev columns represent numbers published by Kev, while the Kev-9B JevBench result is not available. All JevBench numbers are based on the public easy, standard, and hard tiers, which consist of 231 items. The sealed judge tier is not included, and Kev's numbers are restricted to the same public items.\n\nWithout the reasoning component, the same checkpoint achieves an accuracy of 0.804 on the test split (2,962 items) compared to 0.840 with it. The weights released can be downloaded and served using a standalone model or by fusing a trained checkpoint with a drafter. The server can be accessed on one H100 (FP8) with three questions processed in parallel.\n\nJeeves is a drop-in replacement for Jev's Python SDK (typesafe-sdk), requiring no API key and waiting up to 120 seconds for responses. The client connects to http://127.0.0.1:8009 by default or JEEVES_BASE_URL. Questions, states, and answers are loaded into the Qwen chat template, with the model rolling out a reasoning chain after the \"/think\" token. A pointer head scores each option using a scaled dot product between a query projection of the hidden state at \"decide\" and a key projection at the option's \"/opt\" token. This method maintains calibration and achieves high dev scores when stopping at step 402.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}