The Model That Costs 3x More Won by Exactly One Question
I gave the same 29-question order-reading exam to two models. A cheap one (Haiku 4.5) and one that costs about three times as much (Sonnet 5). The result Cheap model 28 of 29 questions clean Expensive model all 28 executed questions clean Fatal errors zero for both (One of the expensive model's questions never ran — rate limit.) The 3x model really was better. By exactly one question. That one…
Two models, Haiku 4.5 and Sonnet 5, were tested on the same 29-question order-reading exam. The cheap model answered 28 questions correctly, while the expensive model got all 29 correct. However, the expensive model had a fatal error on one question - the 300-neck lid size was determined without the need for confirmation. The cheap model, on the other hand, asked for confirmation, which led to a harmless question being asked instead of an automatic response.
Both models had zero fatal errors, but the cheaper model scored better overall. The difference between the two models was one question, which was resolved by the expensive model by asking for confirmation instead of automatically providing a response. The takeaway is to compare models based on severity of errors, not just the overall score.
Additionally, temperature settings can affect the model's output, so pinning the temperature or avoiding a single run for critical tasks is recommended.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.