Urgent.News

What's breaking now, across thousands of outlets.

AI

Questions for a chatbot

Today I have a file open with an empty table. At the top are the numbers from two weeks ago: out of 68 answers, one passed. At the bottom, the hypotheses I wrote so I wouldn't cheat myself when measuring again. The table in the middle, the one that would say whether the chatbot got better, has been empty for 16 days. What I ask it The chatbot answers questions about pasture growth rates with a…

I have an empty table with today's numbers from two weeks ago: out of 68 answers, just one passed. The hypotheses I wrote to prevent cheating when measuring again sit at the bottom. The audit stage, which costs about 28 dollars, was halted due to an expired API credit on September 15. I know what changes I made, but I don't know if the chatbot's performance has improved.

When you evaluate something you've created, does it give you a score or a list of what needs fixing? And do you write down what you expect to change before conducting the evaluation?

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Thursday 1 October →