Urgent.News

What's breaking now, across thousands of outlets.

AI

span-01 vs mercury-decide: same score, opposite failures

span-01 vs mercury-decide: same score, opposite failures Last time I tested a "decision model" — a model that takes a plain-language question about a text and answers with a probability — as a gate for keeping Japanese narration free of English words. That article is here: Is regex enough? I tested span-01 on mixed-language text Code and measured data: sunnydachs / span01-eval Evaluating a…

span-01 and mercury-decide achieved identical scores in a recent test assessing their ability to identify English words within Japanese narration, but their performance differed in critical ways. Both models evaluated the same 23-case boundary suite, which included examples like English words, sentences, and URLs mixed into various languages. The identical scores (F1 0.93) indicated comparable accuracy, yet the underlying performance varied significantly.

Merkury-decide, developed by Inception, returned a binary decision (yes/no or score) with associated probabilities for each case. In contrast, span-01-lite generated probabilities ranging from near-zero for negatives to higher values for positives, exhibiting a broader probability distribution. This difference in probability output led to distinct reactions to instruction wording.

Span-01 performed consistently across various instruction phrasings, while mercury-decide showed divergent behavior based on the clarity of the instructions—certain cases were dropped or missed depending on the level of detail provided in the instructions.

Perhaps most intriguing was the day-to-day variability. Over a single day, mercury-decide's scores fluctuated markedly for certain cases. For instance, a katakana word it correctly identified as non-English the previous day was misclassified as positive the next day. Conversely, span-01 demonstrated remarkable stability, maintaining its verdicts across multiple measurements over several days.

These discrepancies highlight the reliability of span-01 and the volatility of mercury-decide, suggesting that span-01 may be the more dependable choice for consistent performance.

In summary, while both models performed equally well in this controlled test, their underlying behaviors and responses to input variations diverged substantially. Span-01 provided a more consistent and reliable decision-making process, making it a preferable choice for applications requiring stable performance.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Attach an API spec to a Jira workflow validator: what the CogniRunner AI rule never reads

Attach an API spec to a Jira workflow validator: what the CogniRunner AI rule never reads Key takeaways END STATE: a validator that reads your own schema before deciding, with the verdict provably…

  • Attach JSON schema, API spec, or field-mapping table to Jira workflow validator via CogniRunner AI.
  • Confirm verdict changes by running input with and without added document.
  • Prompt ceiling limits substring to 30,000 characters, cutting mid-string.

Give a Jira workflow rule a memory that survives, without blowing the prompt budget

Give a Jira workflow rule a memory that survives, without blowing the prompt budget Key takeaways END STATE: a workflow rule that accumulates lessons about YOUR instance across runs, injected under a…

  • Stand up a local harness to run the module without Forge
  • Save lessons with unique Jira instance details
  • Measure Jaccard similarity to avoid merging similar lessons

More from Saturday 3 October →