Urgent.News

What's breaking now, across thousands of outlets.

AI

What happens when you reject an AI agent's work in Bees

Most agent tools give you two choices when the output is wrong. Start over, or edit the prompt and hope. We wanted something in between for Bees, the open source desktop app we build that runs a team of AI agents on your own computer. Every goal gets a reviewer A goal in Bees runs through three stages: Work, Review, Done. A work agent does the job. Then a fresh reviewer agent checks the result.…

When you reject an AI agent's work in Bees, the process goes beyond simply starting over or editing the prompt. Bees implements a middle ground in the form of a reviewer agent. This reviewer checks the work against the goal and the evidence recorded by Bees, without considering the worker's private reasoning. If the work fails to meet the requirements, the reviewer sends the task back to the work stage with detailed feedback, starting a new attempt.

This process is capped with a limit on retries before the goal transitions to a "Needs your attention" state.

In addition, Bees offers the option for users to directly review certain stages or goals by providing "Approve" or "Reject" choices. When rejecting, a reason is required. This reason is scoped, meaning its impact depends on the context: for a single goal, the feedback is given back to the agent for revision, while for scheduled jobs, the reason is added to a playbook for that specific job.

This playbook is managed by a specialist, a specific agent tailored for that job. When an agent starts handling a named schedule for the first time, a specialist is created, inheriting the base agent but adding job-specific notes. This ensures that changes are isolated to the relevant job and do not affect other aspects of the agent's behavior.

Notably, one-off goals do not influence future behavior, meaning there's no automatic learning based on rejection scores. The reviewer's decision doesn't alter the base agent for the entire organization, nor does it change the agent's base functionality due to a rejection. This distinction between denying an agent's tool action and rejecting the work itself is maintained. For instance, rejecting an email sending action is a permission decision for the moment and doesn't lead to a permanent rule.

Bees provides transparency into the review process. Each specialist and playbook can be inspected under Process Runs → Schedules, showing the playbook, its revision number, and the history of changes. These changes can be edited, undone, or reset to the base agent state. This approach prioritizes human oversight over automatic learning, recognizing that a poorly reasoned rule based on opaque criteria can be more detrimental than no rule at all.

Bees is open source, available under MIT or Apache 2.0 licenses, and supports various AI models like Codex, Claude Code, or local models. It is accessible on Mac, Windows, and Linux platforms at https://bees.bot/download/ and its source code is hosted at https://github.com/Bees-bot/bees-desktop.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

I asked 63 models the same 76 questions, with and without web search

The 2026 standard deduction for a single filer is $16,100. I asked 63 models what it was. Asked from memory, 8 of 61 got it right. 24 invented a number, three of them landing on $8,300.

  • Only 8 out of 63 models answered 2026 standard deduction correctly from memory
  • 24 models fabricated numbers, including $8,300, when given from memory
  • 14 out of 15 models with web search capability provided correct answers

10M Batch LLM Inference at $0 Cloud Cost: O(1) Memory Clamped Architecture

High cloud API costs and Out-Of-Memory (OOM) failures in large-scale data pipelines are architectural defects, not hardware constraints.

  • 10 million-record batch LLM inference executed on local workstation
  • Memory usage clamped between 6.72GB and 9.4GB, O(1) space complexity
  • No cloud compute costs incurred, all processing local

More from Thursday 8 October →