Urgent.News

What's breaking now, across thousands of outlets.

AI

Property Testing with Agent Swarms

Agents can construct the necessary tools to identify bugs within your codebase, utilizing those tools to generate and validate test cases without the need for excessive computational resources. This process entails creating reference implementations, generators, and assertions for each custom data structure and algorithm, which can be delegated to a swarm of agents.

Begin by selecting a complex repository at your place of work and discuss with knowledgeable team members the aspects they find particularly challenging. Provide a robust planning agent with the property-testing capability from a designated agent-skills repository along with these insights. The agent is then equipped with methods for selecting targets, constructing reference models and generators, and transforming errors into reviewable fixes.

This approach can be applied to custom data structures, query planners, and graph algorithms, allowing for the examination of complex behavior through a straightforward method of verifying correctness. Utilizing Codex as an example, OpenAI's frontier models, including yet-to-be-released versions, are employed. Thirteen separate correctness bugs were identified, with one rollback operation that was expected to have no effect inadvertently erased saved review history.

A comprehensive plan resulted in multiple wrong plans being replaced by examples hidden within citations. As of October 7, thirteen upstream reports have been submitted, each accompanied by twelve proposed fixes and two failing-test-only pull requests in a personal fork; two of these fixes address the same carriage-return framing bug in separate diff renderers.

The same playbook can be applied to other repositories such as jj, uv, Prometheus, and Babel during regular work hours. As of October 6, four uv fixes have been merged by upstream: activation of optional dependencies during export, merging compatibility tags from separate wheel metadata rows, twice caching a workspace root, and losing optional-dependency guards.

The latter could potentially include a dependency even when it was not explicitly requested. By October 7, Prometheus has incorporated two fixes: addressing BucketQuantile's panic on empty input despite its documented NaN result, and histograms losing counter-reset metadata upon schema reduction. Frank Noirot also utilized this method on KittyCAD's cloud-sync IndexedDB code, uncovering two bugs, both of which have been merged: waiting for transaction commit before acknowledging writes, and closing connections after aborted cursor transactions.

These changes have been credited to the post and skill. This approach is adaptable to various individuals, repositories, and programming languages. At Apollo GraphQL on Router, a Rust GraphQL router used in production by Intuit and Wayfair with over 30 correctness bugs found in edge cases of internal data structures and algorithms through agent-driven testing, the recipe proved effective.

Initial guidance and occasional adjustments were provided by the human operator, while agents wrote simple reference implementations and assertions to compare them with the actual implementations over generated operation sequences. Tools such as Astra for planning, Sol for orchestration, and Sols and Lunas for writing proptests in parallel worktrees were employed.

The planner was encouraged to audit beyond the initial leads, comparing cached metadata with recomputation and fresh traversals of incremental graph algorithms. Generated values were round-tripped through serializers, and the planner sought out opportunities across module boundaries. Failures, including violations of test assumptions, were thoroughly investigated, with regression tests and fixes generated for contract violations and ambiguities.

Bugs discovered through code reading also warranted regression tests. Proptests continued to run while additional agents wrote more tests; periodic checks ensured the search remained on track. A simple HashMap was modeled to represent a store with cached lookups, and generated sequences of Put(key, value), Get(key), and Remove(key) were subjected to both implementations, with read results compared.

The reference lacked a cache to invalidate. A small key pool, consisting of mostly unrelated inserts and missing-key lookups using fresh random keys, was used to overwrite values and expose stale results. If a target yielded no findings, the initial assumption was that coverage was insufficient; eviction of a cache invalidation was temporarily removed, and if the tests passed, the generators or assertions were improved.

Upon reverting the deliberate bug, proptest was used to shrink failing histories by removing operations and simplifying arguments while maintaining the failure. An agent was tasked with extracting a standalone regression test. Confidentiality was maintained by addressing potential security issues privately through the company's security process or the project's private reporting channel, ensuring that repros and fixes remained undisclosed in public issues, PRs, and discovery branches until cleared for disclosure.

Each bug was branched from the main branch, containing solely the regression test and fix, which was verified by ensuring the test failed before the fix and passing after. Descriptions of the triggering sequence and violated contract were provided, focusing on end-to-end application crashes rather than extensive details.

Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at recursion.wtf →

More in AI

More from Thursday 8 October →