The app worked, the product didn’t: Can we install judgement into AI agents?
Our app worked. That was the problem. My team had spent roughly half a year working with our developer and using AI to build an in-house learning management app. In our testing environment, every function appeared to work. Zoom links could be updated. Calendars were connected. The automated checks reported that the system worked. When […] The post The app worked, the product didn’t: Can we…
The AI-built app appeared functional during testing, but real-world use revealed critical flaws. The team's lack of testing from the user's perspective led to missed calls, overlooked bookings, and an overall poor client experience. The software's interface was difficult to navigate, and the coaches did not have the necessary reminders and information to monitor student bookings.
This failure highlighted the importance of evaluating the product through multiple lenses, including client-facing, administrative, and internal-user perspectives.
The author learned that autonomy should be paired with clear completion criteria and regular feedback points. Pausing the launch allowed the team to improve the interface and define what "done" meant. This experience demonstrates the need for AI agents to assess not only functional correctness but also the consequences of their actions through human judgment.
While AI can generate test cases, it is essential for people in each role to test the real work and provide feedback. By incorporating human judgment into AI agents, we can create a system that evaluates consequences, maintains accountability, and prevents irreversible actions.
Written by urgent.news from e27's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.