Urgent.News

What's breaking now, across thousands of outlets.

AI

8 of my AI agent's 30 test calls failed. Every one was my fault.

The record said there was a gas leak. It did not say where. { "urgency" : "emergency" , "emergency_type" : "gas" , "callback_number" : null , "service_address" : { "street" : null , "city" : null , "postal_code" : null }, "outcome" : "escalated" } My AI intake agent wrote that after a caller said their furnace was out and the basement "kind of smells like eggs." It told them to leave the house…

Eight out of 30 test calls conducted by an AI intake agent for home service contractors in the US failed, all due to human error. The AI intake agent is designed to handle HVAC, plumbing, and roofing emergencies, and it answers phone calls (or SMS, or web forms), determines urgency, collects necessary information, and writes a JSON record for the contractor's system.

A total of 30 scripted test calls were created, with ten per trade. The agent's responses were scored on nine criteria, four of which were checked by plain code, while the other five required a secondary model for assessment. Four cases were found to be impossible to pass, nearly a quarter of the test suite, due to the lack of mandatory information like phone numbers and addresses that were never provided in the scripted caller lines.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Tuesday 22 September →