Urgent.News

What's breaking now, across thousands of outlets.

AI

My laptop AI got 0 of 88 contest deadlines right. Sanity got it to 70. Time tools got it to 87.

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content "Why build this? You can just look it up." That was my first reaction too, so I tested it. I took 22 real contests and asked an AI model on my laptop when each one closes, in four cities. That's 88 questions, and every one asks for the date on a final answer line. Here's what came back: No lookup…

This report examines the results of a contest involving an AI model attempting to accurately determine the closing dates of various contests across four cities. Using a dataset of 25 real contests, the AI was tested and achieved varying levels of accuracy. In one instance, without utilizing any lookup tools, the AI correctly answered none of the 88 questions.

However, when fed the official contest records, the AI's accuracy improved significantly, correctly answering 70 out of 88 questions. The final result was even more impressive, with the AI achieving 87 correct answers when combined with the use of time tools. The primary challenge was determining the correct local time for each contest, as the AI made several mistakes in its time calculations.

Despite this, the use of time tools ultimately led to the highest accuracy rate. The report also outlines the specific rules and conditions under which the AI's performance was evaluated, including the use of a specific AI model, dataset version, and context length.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

AI Dev Weekly #28: GPT-6.1 Sol, Claude Sonnet 5.5, Dots and NVIDIA OpenShell

AI Dev Weekly is a Thursday series where I cover the week's most important AI developer news, with my take as someone who actually uses these tools daily.

  • OpenAI launched GPT-6.1 Sol with 1,050,000-token context window on September 29
  • Anthropic released Claude Sonnet 5.5 for Opus-like performance at lower costs
  • NVIDIA introduced Dots agents for always-on functionality in ChatGPT

More from Thursday 1 October →