Urgent.News

What's breaking now, across thousands of outlets.

AI

I asked 63 models the same 76 questions, with and without web search

The 2026 standard deduction for a single filer is $16,100. I asked 63 models what it was. Asked from memory, 8 of 61 got it right. 24 invented a number, three of them landing on $8,300. 19 gave a figure from an earlier year. 9 refused to answer. Of the 15 seats that could search the web, 14 got it and the fifteenth errored out. That is one question. This is what happened across 76 of them, and…

This investigative report delves into the testing of 63 artificial intelligence models on 76 questions, some with and some without web search capabilities. The study aimed to determine the accuracy and reliability of the models when given a short, unambiguous answer and a primary source to back it up.

Out of the 63 models, only 8 got the 2026 standard deduction for a single filer correct when given from memory. However, 24 models fabricated numbers, including $8,300, while 19 gave figures from previous years and 9 refused to answer. When allowed to search the web, 14 out of the 15 models that had the ability to do so provided the correct answer, with one model encountering an error.

The study analyzed 76 questions, involving 63 models, resulting in a total of 5,776 graded answers. The process involved a single call to the provider before every question and a panel of two judges to grade the answers that did not match the verified answer or a recorded earlier value. Of the 2,910 graded answers that went to the panel, 2,867 agreed on a verdict, resulting in a 98.5% agreement rate among the judges.

The study also examined the cost of using paid models and native web search, revealing that the paid models charge per call, while the web search cost is passed through and charged at request finish. This led to an over- recording of $38.22, compared to the provider's metered cost of $31.74. The report concludes by highlighting the importance of precise grading criteria, accurate data collection, and the need to avoid relying on the models themselves to grade their own answers.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Countering Developer Burnout with Agentic Test Execution

The QA Fatigue Epidemic: Recent surveys reveal developers feel reduced to "human meat proxies," burning out from manually debugging "almost right" AI code.

  • Developer burnout stems from reduced agency in AI coding workflow.
  • Junior engineers suffer most from overreliance on AI snippets.
  • Agentic test execution tools automate verification to combat burnout.

What happens when you reject an AI agent's work in Bees

Most agent tools give you two choices when the output is wrong. Start over, or edit the prompt and hope. We wanted something in between for Bees, the open source desktop app we build that runs a team…

  • Bees uses a reviewer agent to check AI work against goals and evidence.
  • Rejecting AI work requires a reason, scoped by context (single goal vs scheduled job).
  • Reviewer's decision doesn't alter base agent for organization or its functionality.

10M Batch LLM Inference at $0 Cloud Cost: O(1) Memory Clamped Architecture

High cloud API costs and Out-Of-Memory (OOM) failures in large-scale data pipelines are architectural defects, not hardware constraints.

  • 10 million-record batch LLM inference executed on local workstation
  • Memory usage clamped between 6.72GB and 9.4GB, O(1) space complexity
  • No cloud compute costs incurred, all processing local

More from Thursday 8 October →