Urgent.News

What's breaking now, across thousands of outlets.

AI

Why Did the Test Fail? A Triage Rulebook, Tested on Two LLMs

Why did the test fail: the product, the test, the data or the environment? An evidence-based triage rulebook with rules and 21 labelled cases.

Why Did the Test Fail? A Triage Rulebook, Tested on Two LLMs

A failing test indicates a problem, but not what the problem is. A new rulebook aims to sort test failures by determining which artifact is incorrect: the product, the test itself, the test data, or the environment. Two large language models (LLMs) applied this rulebook to 21 failures. Initially, one model assumed the API change caused a new test's expected value to be incorrect.

After applying the rulebook, that assumption was eliminated, and each failure was associated with the missing evidence needed to identify the root cause. The rulebook helps teams avoid chasing the wrong issue and clarifies the specific artifact that's responsible for the failure.

Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at hackernoon.com →

More in AI

Meshy AI 3D Generator Review 2026: 35 Credits, 69.18 MB Monkey

Disclosure: this is an independent hands-on test. Both tools were run on ordinary customer accounts; neither company supplied review access or saw this piece before publication.

  • Meshy generated 69.18 MB file, SupaVoxel generated 11.17 MB file
  • Meshy took 46 seconds, SupaVoxel took 7 seconds to generate
  • Meshy required 35 credits, SupaVoxel required 3 credits

Why I am building Threshold around replaceable agent sessions

About four months ago, I started using coding agents, beginning with Codex. Small tasks went well. I could describe a change, inspect the result, and move on. Longer projects felt different.

  • Threshold aims to maintain project continuity across different agent sessions.
  • Core concepts of Threshold include Project, Task, and Run.
  • A test tool using Threshold successfully completed tasks across multiple agent sessions.

More from Sunday 11 October →