It's Not the Model, It's You
Every time a new frontier model comes out, we go through the same rituals: benchmarks get reported, leaderboards get reshuffled, and the social feeds fill up with a heady cocktail of vibes, hot takes, and everything we’ve found the new model can “one-shot.” Anyone watching from the sidelines could be forgiven for thinking that model performance is the single biggest factor driving AI…
Every time a new frontier model arrives, a familiar routine begins: performance benchmarks are published, leaderboards are updated, and the social media feeds buzz with excitement over the new model's capabilities. Many observers might assume that model performance is the most significant factor driving AI transformation today. However, upon closer examination of how these new models impact real-world engineering teams, the reality is more nuanced than the initial hype suggests.
For instance, this summer, Cursor modified its default settings after SpaceX acquired the platform. Engineers started using its in-house Grok model by default, with free users compelled to do the same. Within months, the usage of Grok surged, accounting for nearly half of all Cursor requests, while Claude, GPT, and other models dropped to under a quarter.
This shift wasn't limited to free users; among paying customers, Grok usage reached nearly 40%, with a third of engineers who had been using Claude or GPT primarily using Grok by September.
If models truly mattered as much as we claim, we might expect engineers to see noticeable improvements in their coding output. Yet, there was no discernible difference in output between engineers who heavily adopted Grok and those who used it minimally (over 12,000 engineers at 222 companies). Deeper analysis revealed that this phenomenon isn't unique to a single model or vendor.
When comparing engineers who frequently switch between major frontier models and families, no measurable difference in their output could be found.
So, what does impact engineering teams' productivity? Our large-scale analyses across the data indicate that the model and its associated tools contribute less than 1%. Approximately 9% of the difference is attributed to the volume of AI the engineers utilize. Around 12% is influenced by their company, team, and organizational culture.
However, more than 75% of the variance is determined by individual engineers - their role, experience, habits, and other personal attributes that distinguish seasoned AI engineers from those who are still learning.
While models undoubtedly play a role in AI transformation, their impact is relatively limited when considering the entire ecosystem. Better models do raise the ceiling for everyone, and the frontier models are a significant part of this advancement. However, the specific model engineers choose to use among the available options explains only a small fraction of the difference in their output.
In essence, for the top performers, the choice of model might be akin to an amateur athlete choosing between two different carbon-plated racing shoes - a small distinction compared to a seasoned marathoner who has dedicated countless hours to training and refining their physique. Therefore, for engineering leaders and practitioners, the news is ultimately positive, albeit somewhat disappointing.
The model, while essential, is the most accessible lever at their disposal, as it can be changed with a simple dropdown, router setting, or procurement decision (though often beyond their control, as in the Cursor example).
The more significant factors that matter are the slower and more challenging aspects of building AI proficiency within teams. This includes fostering a culture that encourages deep engagement with AI, establishing practices that enable engineers to work effectively with AI tools, and investing in the professional development of the individuals. While the pace of model progression is largely beyond most organizations' control, these other factors are within their grasp.
Therefore, the next time a new model drops, it's worth trying it out and sharing your experiences. However, if you aim to understand why some teams achieve remarkable results with AI while others struggle, the answer is more likely to lie within your organization's culture, practices, and individual talent rather than the model picker.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.