Benchmarking AI for Code Optimization: Which LLM Actually Writes the Most Efficient C/Rust Code?
This is a submission for the Kaggle Benchmarking Challenge ( https://dev.to ) ## What I Benchmarked As a third-year Computer Science and Data Science student, I frequently work with lower-level system operations where runtime efficiency is everything. While many public leaderboards evaluate LLMs on whether their code simply functions or passes a basic test suite, they rarely measure the…
This is a report on a study comparing how well different AI language models create efficient C and Rust code. As a third-year Computer Science and Data Science student, this researcher frequently works with low-level system operations where performance matters greatly. While most evaluations of AI models focus on whether the code runs at all, this study measured key performance metrics such as runtime efficiency, compilation overhead, and algorithmic complexity.
The study tested four AI models on a set of 15 complex tasks involving things like custom memory management, OS kernel loops, and data sorting arrays in C and Rust. These models were DeepSeek-Coder-V2-Instruct, known for being highly specialized in code logic; Claude 3.5 Sonnet, a top-tier model for complex reasoning and multi-file software construction; GPT-4o, used as a benchmark for general logic capabilities; and Gemini 1.5 Pro, chosen for its large context window that allows prompt constraints to include detailed hardware architecture specifications.
The results showed that while all four models successfully compiled their output 90% of the time, their performance varied significantly. DeepSeek-Coder-V2 generated code that ran up to 18% faster than GPT-4o on complex C pointer manipulation tasks because it avoided unnecessary stack-to-heap data transfers. When creating Rust code, Claude 3.5 Sonnet consistently chose the most efficient zero-cost abstractions that satisfied the Rust borrow checker without needing unsafe code blocks.
Gemini 1.5 Pro, despite producing structurally sound code, often resorted to unsafe blocks to force compilation on complex data ownership scenarios.
One surprising finding was that the model specializing in general code reasoning (Claude) slightly outperformed the highly specialized model (DeepSeek) in terms of code readability and helpful comments, even though it didn't match DeepSeek's micro-optimizations. The researcher suggests that future benchmarking should examine how well these models optimize code for specific architectures, like ARM versus x86 assembly.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.