Urgent.News

What's breaking now, across thousands of outlets.

AI

Kinetic-4B vs Claude Haiku 4.5: The 4B Model Wins Tools

Kinetic-4B wins on tool calling. On a 300-sample Composio evaluation, the 4-billion-parameter model from Bengaluru lab Conscious Engines scored 82.33% accuracy at 1.61s p95 latency, against 80.0% and 4.02s for Anthropic's Claude Haiku 4.5, and 76.33% at 7.99s for OpenAI's GPT-OSS-120B ( Conscious Engines, 1 April 2026 ). That is 2.5x lower tail latency at slightly higher accuracy. Haiku remains…

The Kinetic-4B model, developed by Bengaluru-based Conscious Engines, outperformed Anthropic's Claude Haiku 4.5 in a recent tool calling benchmark. Kinetic-4B achieved 82.33% accuracy, 95.33% tool-name accuracy, 4.67% failed calls, and a p95 latency of 1.61 seconds, compared to Claude Haiku 4.5's 80.0% accuracy, 90.33% tool-name accuracy, 9.67% failed calls, and a p95 latency of 4.02 seconds.

In contrast, OpenAI's GPT-OSS-120B trailed with 76.33% accuracy and a p95 latency of 7.99 seconds. The smaller Kinetic-4B model, trained on a single rented GPU for about 4.5 hours, demonstrates that fine-tuned, smaller models are more cost-effective and faster for specific tasks like tool calling, while larger models excel at open-ended reasoning and code generation.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

DeepSeek V4.1-Flash Wins on Coding Agent Cost

Verdict: For coding agents, DeepSeek V4.1-Flash is the best open source LLM right now, because it lands level with Claude Opus 5 on software engineering while costing a fraction as much per million…

  • DeepSeek V4.1-Flash leads in coding agent performance, scoring 74.2 in DeepSWE v1.1
  • V4.1-Flash offers cost-effective pricing at $0.003 per 1M cached input tokens
  • V4.1-Flash struggles with complex reasoning tasks compared to Claude Opus 5

Agentic AI vs Generative AI: The 2026 Verdict

Generative AI wins for anyone producing content — drafts, images, code snippets, summaries — because it is cheaper, mature and easy to review.

  • Generative AI excels at creating content like drafts, images, and summaries at low cost and ease.
  • Agentic AI is ideal for multi-step workflows with verifiable results, such as fixing test suites.
  • Only 40% of agentic AI projects are expected to succeed by 2027 due to costs and risks.

More from Wednesday 23 September →