Choosing an AI model: one prompt, 11 models, different results
Article URL: https://www.netlify.com/blog/one-prompt-11-models-very-different-results/ Comments URL: https://news.ycombinator.com/item?id=49285327 Points: 210 # Comments: 91
When we ran the same prompt across 11 different AI models on Netlify, the results varied significantly. OpenRouter's new Agent Runners feature, which utilizes a full coding agent, allowed us to compare how each model performed when generating a simple one-page site for a coffee shop. The test focused on functionality rather than design, checking if the generated site utilized the appropriate tools, such as Netlify Database or Identity, and whether it was over-engineered.
The results showed a wide range in credit usage for each model. One model, GPT 5.6 Sol, was specifically used with a lower-effort setting, offering a more cost-effective alternative to more expensive models. The Claude Opus model, in particular, had a notably higher credit usage due to one of its three runs consuming 1,055 credits, which is nearly four times the average for the other models tested. This model, however, produced a detailed, visually appealing site with a custom map and dark mode functionality.
Other models, such as OpenAI Codex and Gemini CLI, also demonstrated impressive results, albeit with slightly lower credit usage. Overall, the test highlighted the trade-offs between cost and performance when choosing an AI model. While some models produced more budget-friendly results, others provided a more detailed and visually impressive site. It's clear that each model has its strengths, and the best choice will depend on the specific needs and priorities of the user.
Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.