Urgent.News

What's breaking now, across thousands of outlets.

More in AI

The Model Was Never the Problem

The shortlist post ended with a promise: stage two of the eval, selection within the shortlist, measured with a live model, scored so that a recall miss can never masquerade as a selection miss.

  • Study focused on tool-selection, not model performance
  • gpt-5.4-mini model used, 90-97% selection accuracy
  • Emphasizes recall importance over model size

More from Thursday 24 September →