Urgent.News

What's breaking now, across thousands of outlets.

AI

Tests are the only fixed point left when the implementation is disposable

Originally published at https://aicoding-guide.com . The more disposable the implementation becomes, the more the spec moves into the tests. That is the claim. Working with an AI coding tool, the same feature gets rewritten several times in a day. What gets rewritten is not the spec. The spec is whatever survives and can tell you that you are wrong: tests, types, and the checks you put in hooks.…

In the age of disposable implementations, tests have emerged as the single constant point. When working with AI coding tools, the same feature is rewritten repeatedly throughout the day. The spec, which dictates correctness, is not replaced by the rewritten code; instead, it is the tests, types, and checks in hooks that serve as the guiding force.

The guidelines for CLAUDE.md, sourced from the Claude Code documentation, emphasize that what will be learned is the importance of not treating instructions as a fixed point. Instructions cannot be relied upon as a constant reference point. The line between advisory and failing the build is clear when tests, types, or hooks in the form of hooks become the fixed point.

Instructions may shape what Claude tries to do, but they do not alter what Claude Code allows. A fixed point is anything that can fail, and it is the per-event hook behavior that makes this line explicit. When a machine can fail, that becomes the fixed point. Hooks such as Stop can be used to prevent Claude from stopping the conversation, allowing a test suite to run and exit with an exit code of 2 if the tests fail.

This turns the request "don't report done while tests fail" into a mechanism. The importance of creating verification targets cannot be overstated. By including test cases, screenshots, or defining expected output in the prompt, Claude can verify its own work and catch issues before they become a problem. This is not an endorsement for test-first development, but rather a way to avoid wasting tokens.

The demand is the same: fix the expected value first. This reasoning aligns with the case for plan mode, which aims to prevent expensive re-work when the initial direction is incorrect. The structure of the site itself enforces this principle, with every article requiring lint and build scripts to pass before publication. The promise you want kept must be shaped so it can fail.

Reviewing a pull request diff no longer holds the same weight, as each rewrite replaces the diff wholesale. Turning review output into something that fails becomes an essential part of the review process. However, the claim does have its limitations. When AI writes the tests themselves, the fixed point is limited to the human-decided part of the implementation.

Tests generated afterwards only serve as a regression net, not the spec of record. Certain domains, such as visual appearance, perceived performance, accessibility, and security posture, cannot be expressed through tests. Building a different fixed point, such as screenshot comparisons, benchmark thresholds, or an audit checklist, is necessary in these cases.

Lastly, the cost of implementing a fixed point must be considered. Tests in legacy codebases with no test infrastructure are expensive, and a throwaway prototype may never recover that investment. Making the decision to build one must be made deliberately and not overlooked.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

An LLM observability platform stores prompts, and prompts are the application

An LLM observability platform stores prompts, and prompts are the application A title query for Langfuse returns 346 matches in ZoomEye. The number is small and the contents are unusual.

  • LLM observability platform records prompts as application logic
  • Traces contain prompts, model parameters, completion details
  • Langfuse indexes sensitive traces, including potential secrets

Your LLM Types One Token at a Time. It Doesn't Have To.

Every token your LLM emits costs one full forward pass through the entire model. Seventy billion parameters loaded from memory, multiplied, discarded — for a single token. Then again. And again.

  • Speculative decoding drafts tokens with cheap model before big model verification
  • Acceptance rule accepts drafted token with probability min(1, q(d)/p(d))
  • EAGLE-3 achieves 2-3x speedup over vanilla decoding, but diminishing returns at high batch sizes

More from Thursday 8 October →