Urgent.News

What's breaking now, across thousands of outlets.

AI

GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?

The post compares the performance of GPT-5.6 Luna and GPT-6 Astra, two AI models used to review pull requests for code correctness and security bugs. The key findings are:

1. Price difference: Luna costs $0.20 per million input/output tokens, while Astra costs $10 for input and $50 for output, totaling $5.66 per review.

2. Bug detection: Astra found 92 verified bugs across 50 pull requests, while Luna found 69. Astra found more security bugs (19 vs. Luna's 9).

3. Cost efficiency: Luna achieved 75% of Astra's bug detection for 3.6% of the cost. Astra's cost was 20x higher per verified bug.

4. Correctness issue: Luna had a 28x higher error rate, with 24 of its 93 findings failing verification, compared to Astra's 4 of 96. However, Luna found 9 security bugs that Astra missed.

5. Bug types: Luna had a higher rate of data and logic bugs (39 vs. Astra's 47). Astra excelled in concurrency issues (13 vs. Luna's 10). Luna found 9 security bugs that Astra did not.

6. Cost comparison: Running both models on every PR would have detected 117 of the 143 verified bugs (82%) for a total cost of $5.86, compared to Luna's $0.20 and Astra's $5.66 for a combined $6.86.

7. Training data: Both models were trained on code that was not recent enough to include fixes to the benchmark bugs, giving them an advantage.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at entelligence.ai →

More in AI

More from Monday 14 September →