Why I am building Threshold around replaceable agent sessions
About four months ago, I started using coding agents, beginning with Codex. Small tasks went well. I could describe a change, inspect the result, and move on. Longer projects felt different. As a conversation grew, it accumulated more than code: why I had chosen a direction, which alternatives I had rejected, what I wanted to leave alone, and the boundaries of the work. Some of that understanding…
In the past four months, I have been experimenting with coding agents, starting with Codex. While small tasks were manageable, longer projects proved more challenging. As the conversation grew, it encompassed more than just code; it included reasons for decisions, rejected alternatives, goals, and boundaries. I grew reluctant to replace this understanding each time a new agent session began. This led to the creation of Threshold, a system designed to allow a project to continue seamlessly when the agent session changes.
Initially, I tried to preserve comprehensive records and determine which records were accurate and authoritative. However, this approach became overly complicated as I added unnecessary machinery to the design. After a major refactor, I narrowed the question to what a fresh session needs to continue its work effectively. The essential elements are project files, Git state, executable tests, and checkpoints that explain previous agent actions, choices, and remaining tasks.
Threshold is structured around three core concepts: Project, Task, and Run. A Project represents ongoing work, a Task describes an action to be taken, and a Run is an agent session working on a Task. Runs are independent and can exist alongside each other without becoming a hierarchy of supervisors and subagents. A coordinating Run can also be replaced by a new one, with checkpoints, messages, and task state helping the next Run understand its context.
To test the concept, I used a tiny CSV summary tool built in a cloud environment with Threshold 0.2.0-alpha.7 and DeepSeek's deepseek-flash model. The tool's job was to read unquoted CSV rows, reject malformed input, sum amounts by category, and print JSON with sorted keys. The process involved a staged handoff where the first Run focused on parsing, summing, and writing unit tests.
The second Run, with a different session, inspected the handoff, finished the CLI, and added integration tests and documentation. The results showed that the second Run not only passed all initial tests but also successfully added new functionality.
While this demonstration provides a concrete example of Threshold's capabilities, it does not yet address the reliability of this approach on larger projects, especially when decisions change or checkpoints are insufficient. I am sharing Threshold to gather feedback on what makes people hesitant to replace long-running coding sessions and how Threshold can address these concerns. The project is available on GitHub at https://github.com/Key-of-door/Threshold.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.