Urgent.News

What's breaking now, across thousands of outlets.

AI

アンソロピックのAI、殺人事件の虚偽情報を通報 勝手にビザ申請も

A dedicated AI reviewer, such as Open Code Review, is tested against a universal AI agent in a practical experiment. The test aims to determine whether the specialized reviewer outperforms the general-purpose agent in reviewing a large change of 60 files. Both reviewers are run under matched conditions, and six metrics are scored: precision, recall, file coverage, anchor validity, wall-clock time and token count.

The test plants known defects into a large diff and runs both reviewers to score their performance. The article provides a protocol, scoring script and results template, but does not provide a verdict. It emphasizes the importance of running the same model, temperature and input for both reviewers to ensure a fair comparison.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at asahi.com →

More in AI

The AI Risk Gap Is a Release Gap

Originally published on the Dromeas blog . Gartner's latest quarterly has a line I keep coming back to: to protect the value you get from AI, plan to spend at least twice as much on derisking it as…

  • CEOs and CFOs should allocate twice the budget to derisking AI compared to AI tools.
  • Derisking AI starts at the release stage with decisions to say no or proceed.
  • Gartner recommends certifying each AI release with a verdict and evidence.

More from Saturday 10 October →