{
  "id": 11700576,
  "title": "span-01 vs mercury-decide: same score, opposite failures",
  "url": "https://urgent.news/2026/10/03/span-01-vs-mercury-decide-same-score-opposite-failures",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-03T14:24:12.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/sunnydachs/span-01-vs-mercury-decide-same-score-opposite-failures-1a25"
  },
  "original_language": "en",
  "account": "span-01 and mercury-decide achieved identical scores in a recent test assessing their ability to identify English words within Japanese narration, but their performance differed in critical ways. Both models evaluated the same 23-case boundary suite, which included examples like English words, sentences, and URLs mixed into various languages. The identical scores (F1 0.93) indicated comparable accuracy, yet the underlying performance varied significantly.\n\nMerkury-decide, developed by Inception, returned a binary decision (yes/no or score) with associated probabilities for each case. In contrast, span-01-lite generated probabilities ranging from near-zero for negatives to higher values for positives, exhibiting a broader probability distribution. This difference in probability output led to distinct reactions to instruction wording. Span-01 performed consistently across various instruction phrasings, while mercury-decide showed divergent behavior based on the clarity of the instructions—certain cases were dropped or missed depending on the level of detail provided in the instructions.\n\nPerhaps most intriguing was the day-to-day variability. Over a single day, mercury-decide's scores fluctuated markedly for certain cases. For instance, a katakana word it correctly identified as non-English the previous day was misclassified as positive the next day. Conversely, span-01 demonstrated remarkable stability, maintaining its verdicts across multiple measurements over several days. These discrepancies highlight the reliability of span-01 and the volatility of mercury-decide, suggesting that span-01 may be the more dependable choice for consistent performance.\n\nIn summary, while both models performed equally well in this controlled test, their underlying behaviors and responses to input variations diverged substantially. Span-01 provided a more consistent and reliable decision-making process, making it a preferable choice for applications requiring stable performance.",
  "summary": "span-01 vs mercury-decide: same score, opposite failures Last time I tested a \"decision model\" — a model that takes a plain-language question about a text and answers with a probability — as a gate for keeping Japanese narration free of English words. That article is here: Is regex enough? I tested span-01 on mixed-language text Code and measured data: sunnydachs / span01-eval Evaluating a…",
  "key_points": [
    "Both span-01 and mercury-decide scored equally (F1 0.93) in English word detection test.",
    "Span-01 provided consistent probabilities, mercury-decide varied based on instruction clarity.",
    "Span-01 showed stability over days, mercury-decide exhibited daily score fluctuations."
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}