I asked 63 models the same 76 questions, with and without web search
The 2026 standard deduction for a single filer is $16,100. I asked 63 models what it was. Asked from memory, 8 of 61 got it right. 24 invented a number, three of them landing on $8,300. 19 gave a figure from an earlier year. 9 refused to answer. Of the 15 seats that could search the web, 14 got it and the fifteenth errored out. That is one question. This is what happened across 76 of them, and…
This investigative report delves into the testing of 63 artificial intelligence models on 76 questions, some with and some without web search capabilities. The study aimed to determine the accuracy and reliability of the models when given a short, unambiguous answer and a primary source to back it up.
Out of the 63 models, only 8 got the 2026 standard deduction for a single filer correct when given from memory. However, 24 models fabricated numbers, including $8,300, while 19 gave figures from previous years and 9 refused to answer. When allowed to search the web, 14 out of the 15 models that had the ability to do so provided the correct answer, with one model encountering an error.
The study analyzed 76 questions, involving 63 models, resulting in a total of 5,776 graded answers. The process involved a single call to the provider before every question and a panel of two judges to grade the answers that did not match the verified answer or a recorded earlier value. Of the 2,910 graded answers that went to the panel, 2,867 agreed on a verdict, resulting in a 98.5% agreement rate among the judges.
The study also examined the cost of using paid models and native web search, revealing that the paid models charge per call, while the web search cost is passed through and charged at request finish. This led to an over- recording of $38.22, compared to the provider's metered cost of $31.74. The report concludes by highlighting the importance of precise grading criteria, accurate data collection, and the need to avoid relying on the models themselves to grade their own answers.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.