GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
The post compares the performance of GPT-5.6 Luna and GPT-6 Astra, two AI models used to review pull requests for code correctness and security bugs. The key findings are:
1. Price difference: Luna costs $0.20 per million input/output tokens, while Astra costs $10 for input and $50 for output, totaling $5.66 per review.
2. Bug detection: Astra found 92 verified bugs across 50 pull requests, while Luna found 69. Astra found more security bugs (19 vs. Luna's 9).
3. Cost efficiency: Luna achieved 75% of Astra's bug detection for 3.6% of the cost. Astra's cost was 20x higher per verified bug.
4. Correctness issue: Luna had a 28x higher error rate, with 24 of its 93 findings failing verification, compared to Astra's 4 of 96. However, Luna found 9 security bugs that Astra missed.
5. Bug types: Luna had a higher rate of data and logic bugs (39 vs. Astra's 47). Astra excelled in concurrency issues (13 vs. Luna's 10). Luna found 9 security bugs that Astra did not.
6. Cost comparison: Running both models on every PR would have detected 117 of the 143 verified bugs (82%) for a total cost of $5.86, compared to Luna's $0.20 and Astra's $5.66 for a combined $6.86.
7. Training data: Both models were trained on code that was not recent enough to include fixes to the benchmark bugs, giving them an advantage.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.