Urgent.News

What's breaking now, across thousands of outlets.

AI

LLM Food Recognition in Production: What Shipping soba Taught Me

TL;DR I spent the last few months building soba , an iOS app that photographs a meal and returns carbs, glycemic index, and portion weights for people who count carbohydrates. The recognition backend went through one model migration, one full prompt rewrite, and a stack of validation code. Three things carried almost all of the improvement: picking the model with a benchmark instead of vibes ,…

The story of developing soba, an iOS app that identifies food items in photos and provides nutritional information, is one of trial and error. To start, I selected the initial model based on benchmark performance rather than marketing hype. Gemini-3-Pro emerged as the top choice for both dish recognition and accurate nutrition estimation, with a 24.45% mean absolute percentage error (MAPE) compared to GPT-5's 32.17%.

After launching soba on Gemini-3.1-pro-preview, I conducted an A/B test comparing it to Gemini-3.6-flash on the live recognition path. Despite the same prompt and photos, the flash model significantly reduced costs by 58% and cut latency by a third while maintaining comparable accuracy. However, the models still misjudged portion weights by 25-35%, which proved to be the most significant hurdle.

Focusing on prompt design and implementing a weight-editing user interface proved more effective than continually chasing new models. Another cost-saving decision was to treat certain subtasks, like estimating glycemic index from barcode-scanned products, as independent lookups using a faster model at a lower temperature, thereby minimizing unnecessary processing.

To ensure robustness and maintain control over the output, all API requests were routed through OpenRouter, providing failover capabilities in case of provider outages. The final API schema enforced strict JSON structure, guaranteeing correct parsing without additional token usage.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Every "done" needs a receipt

When an AI agent tells you it's done, how do you know it's true? That question sits under everything I've built this past month.

Why your AI coding agent forgets team decisions (and what to store instead)

Why your AI coding agent forgets team decisions (and what to store instead) Teams running Cursor and Claude Code side by side hit the same wall. CLAUDE.md works for one seat.

  • AI coding agents forget team decisions when used together.
  • Shared memory stores are often missing or divergent.
  • Persist key information like decisions, reasons, scope, and provenance in structured formats.

More from Monday 5 October →