Urgent.News

What's breaking now, across thousands of outlets.

AI

Copilot tops GitHub’s own AI code review benchmark. An independent one tells a different story.

If there’s one thing the AI software engineering world doesn’t lack, it’s benchmarks. Want to know whether an agent can The post Copilot tops GitHub’s own AI code review benchmark. An independent one tells a different story. appeared first on The New Stack .

Copilot tops GitHub’s own AI code review benchmark. An independent one tells a different story.

GitHub has unveiled ReviewBench, an open benchmark created in collaboration with Microsoft, designed to assess the effectiveness of AI code review agents in identifying useful issues within pull requests. The benchmark, announced on Monday, swiftly positioned GitHub Copilot code review atop the inaugural leaderboard, a result that is, unsurprisingly, not unexpected.

In October 2024, GitHub introduced Copilot code review, making it widely available to paid subscribers in April of the following year. Since then, the company has expanded its capabilities, integrating it into an agentic architecture, making it billable through GitHub Actions minutes on private repositories, enabling it to approve pull requests, and setting the "Balanced" review mode as the default in September of the same year.

ReviewBench evaluates 219 public pull requests from 187 repositories, encompassing 19 programming languages. It assembles a reference set from human review comments, subsequent changes by authors, static-analysis tools, and other LLM reviewers. The benchmark's methodology identified 47 findings initially classified as true positives, which were later manually corrected, resulting in a 96.6% agreement between human reviewers and the classifier on the classification of true and false positives.

However, there are notable limitations to this benchmark. GitHub generated the initial entries using its publicly available versions of each product, without conducting or verifying those tests. Additionally, the products were tested on different dates, leading to varying ages of results. For instance, Copilot was tested on October 1st, while Cubic and Greptile were tested in June.

ReviewBench cautions that these products may have changed since their testing, and the performance of its corpus might not accurately translate to a specific company's codebase.

Despite these caveats, GitHub has made its dataset, methodology, and judging setup publicly available, allowing other vendors to submit their own runs. Alejandro Carderera de Diego, a staff applied engineer at GitHub, explained that the company uses ReviewBench to enhance Copilot code review and that its offline results have consistently predicted the direction of subsequent production experiments.

In response, Martian, an AI research company, published its own benchmark, the Code Review Bench, in February. Martian's benchmark combines an offline test with an online tracker based on developers' responses to review comments across real open-source repositories. As of October 6th, Martian’s online leaderboard ranks Cubic first with a 64.9% F1 score, followed by Greptile and CodeRabbit, while GitHub Copilot sits at fourth place with a 60.9% F1 score.

F1 is a standard metric that combines precision and recall equally. Martian's approach has diverged from what it observed in real-world use, making its online benchmark the headline metric.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at thenewstack.io →

More in AI

check this

She used Claude as a diary. The terms of service are now part of the charge. Sam LABBE Sam LABBE Sam LABBE Follow Oct 6 She used Claude as a diary. The terms of service are now part of the charge.

Detour: My Open-Source AI Entry for the Touch Grass Challenge (Challenge Week 1: Touch Grass)

**Detour: An AI Agent Designed to Kick You Off Your Screen 🌿 This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What I Built I built Detour , an interactive AI…

  • Detour is an interactive AI web app designed to break screen-time habits.
  • Users answer a 3-step questionnaire to receive a personalized real-world mission.
  • Detour encourages offline activities and user privacy through open-weight models.

More from Tuesday 6 October →