Urgent.News

What's breaking now, across thousands of outlets.

AI

Measure an AI Coding Harness from L0 to L4 with Harness Score

AI coding assistants can edit a repository quickly, but speed does not tell you whether the repository can catch a bad edit. A project with no instructions, tests, or CI may produce a plausible patch and still leave every important check to chance. This tutorial shows how to measure that surrounding system with Harness Score and use a small open-source lab to improve it step by step. The result…

Harness Score is an AI-driven tool designed to evaluate the maturity of a software repository, offering a deterministic checklist of repository evidence rather than a quality certificate. The process involves running the harness-score command and recording the initial level and score. The tutorial provides a step-by-step guide to improve the score by addressing each maturity layer sequentially.

To begin, create a disposable copy of the Harness Score tutorial using GitHub CLI, cloning the template into a new directory. The process requires Node.js 24 or newer, as well as an AI coding agent capable of editing the local repository. GitHub CLI is optional but recommended for creating a private or public copy of the template.

The first step is to establish a baseline score by running npx --yes harness-score . --json, which outputs a JSON report containing maturity level, earned points, dimensions, checks, evidence, and remediation links. Record the scanner version and baseline score, as they will only apply to a specific commit and scanner version.

Improving the repository's maturity involves several stages. First, provide the AI agent with context by creating a basic Meeting Cost CLI application. This includes defining the domain calculation separately from terminal argument handling and using Node.js built-ins only. After confirming that the command works, add a substantive root AGENTS.md file describing the application's domain, invariants, error handling, and security boundaries.

The next stage involves adding path-scoped rules, reusable skills, and explicit workflows to the project. This separation of context files, scoped rules, and skills helps the scanner detect which dimensions are blocking progress. The repository should then achieve a score of L1 - Harnessed.

Subsequent improvements include adding sensors and CI feedback, such as tests for valid and invalid calculations, formatters, linters, strict type checking, and GitHub Actions workflows. While Harness Score can detect the presence of these sensors, it is crucial to review whether the tests cover meaningful behavior.

Finally, the last stage involves implementing Cursor hooks to close the loop, such as gate hooks that deny dangerous shell commands and feedback hooks that format supported files after edits. Hooks execute at a boundary where prose can be ignored but should not be considered a universal security boundary.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

What GPT-6.1 Sol chose to build

A cutting layout can look convincing while leaving its cut order ambiguous. A document comparison can work while putting its navigation offscreen.

More from Thursday 1 October →