Urgent.News

What's breaking now, across thousands of outlets.

AI

Lint your agent tasks before anyone claims them

More and more tasks on a team board now have two readers : The coding agent that does the work. The teammate who starts the agent, watches what it does, and decides if the result is good enough to ship. Most task descriptions are written for neither of them. They are too short for the agent, and they mix context with instructions, so a person can't scan them quickly during a run. The usual fix is…

When tasks on a team board are assigned to both an automated agent and a human teammate, the initial task descriptions often fall short. They are either too brief for the agent to understand, or they mix context with explicit instructions, making it difficult for a person to quickly review them during a work cycle. The standard approach is to expand the task description, but a more effective strategy is to create two distinct components: a detailed brief for the agent and a concise runbook for the person.

For the agent, a lengthy and precise brief is essential. The agent lacks context from previous standups and doesn't have prior knowledge of which files are sensitive or which parts of the code require exclusions. However, the agent will read whatever brief you provide, so the brief can be as extensive as necessary. It should answer key questions such as the desired outcome, relevant context like files, systems, earlier decisions, and links, what must not change, and what kind of evidence is required to validate the task.

The brief should also specify when the agent should pass the task back to a human, rather than making an assumption.

The person, already familiar with the project, requires a short checklist to guide them while the agent works. This runbook should list what judgments the agent is not permitted to make, what red flags warrant stopping the process and seeking human intervention, and provide a straightforward set of actions to verify the task's completion. The runbook should be concise, ideally not exceeding a few paragraphs, to ensure that it can be easily remembered by the person during the work cycle.

This dual approach has several benefits. Firstly, it keeps the runbook brief on purpose; if the human section becomes too lengthy, it indicates that the task may require a more comprehensive brief. Secondly, the evidence required should be something concrete that can be easily checked, such as passing a specific test, adding a new test, or providing a visual comparison.

Thirdly, both readers should have a clear stop condition: for the agent, it's about scope – ensuring that the fix doesn't extend beyond the areas specified; for the person, it's about the overall run – such as a failed build, unusual requests, or outcomes that don't align with the expected results. Lastly, sensitive information like credentials should never be included in the task, as it can inadvertently be shared with the agent, leading to security risks.

This method can be implemented using a simple issue tracker template with two distinct headings, allowing both halves of the task to exist separately. It's worth considering tools that store these instructions in separate fields, ensuring that the agent only receives the relevant instructions and the human receives a clean checklist tailored for their responsibilities.

However, it's important to acknowledge that this approach does not enhance a vague task; if the outcome is unclear, splitting the task into two components won't make it clear. The additional time spent on creating these separate components may not be justified for minor tasks like fixing a typo. Nevertheless, the benefits become evident when dealing with complex tasks that could easily deviate from the defined scope, or when human judgment is required to determine the acceptability of the work.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

How accurate is Bengaluru’s AI traffic enforcement? | Explained

Amid debates around the shortcomings of using AI in smart policing and rule enforcement, The Hindu looks at how the Bengaluru Traffic Police’s Intelligent Traffic Management System works, how accurate are the violations flagged, how are the challans generated and the means through which erroneous cases can be challenged.

Jev is the if statement of AI

You would never call an LLM to compare two numbers. Yet that is roughly what we all do with AI right now: we hand a model that could draft a legal opinion a question with two possible answers, and we…

  • Jev acts as an unspoken branch in AI decision-making
  • Unlike LLMs, Jev generates calibrated probabilities, not text
  • Jev is 194 times faster and 445 times cheaper than traditional models

AI Issue Assistant: Turn Past Issues into Useful Insights

AI Issue Assistant — Turning Past Issues into Useful Insights This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend .

  • AI Issue Assistant streamlines issue creation and management
  • Generates vector embeddings for issue representation
  • Identifies similar historical issues via semantic similarity

More from Friday 2 October →