{
  "id": 13713372,
  "title": "We Wanted a Fast Offline AI. First, We Had to Fix Our Benchmark.",
  "url": "https://urgent.news/2026/10/11/we-wanted-a-fast-offline-ai-first-we-had-to-fix-our-benchmark",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-11T12:33:36.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/lastbrowser/we-wanted-a-fast-offline-ai-first-we-had-to-fix-our-benchmark-2n9k"
  },
  "original_language": "en",
  "account": "Our team sought to create a browser assistant capable of performing tasks like reading invoices, comparing offers, and finding page buttons, all without requiring an internet connection. After rigorous testing with various model types and configurations, we determined that speed, usefulness, and reliability were paramount rather than model size. However, we soon realized that our initial benchmarking did not account for the specific characteristics of the machines we were testing on. The first iteration of our benchmarking took place on a powerful machine, while the actual target was an older, less capable CPU. This discrepancy meant that our initial results did not accurately reflect the performance of our models on the intended target machine. Furthermore, we found that our evaluation process had shortcomings. For instance, our coding assistant had provided premature assessments on the thoroughness of our reviews, and our first question about a specific CPU model led us down a path that obscured the true focus of our experiment. To truly gauge the performance of our models in a browser assistant context, we needed to establish a consistent benchmarking framework that accurately reflected the conditions in which these models would be used. This involved documenting the target machine specifications before running any tests and ensuring that all measurements were consistent across different hardware configurations. In essence, we learned that a well-defined target machine and a comprehensive evaluation process are crucial for accurately assessing the capabilities of local AI models intended for real-world applications.",
  "summary": "A browser assistant should be able to read an invoice, compare two offers, or find the right button on a page. Ideally, it should still work when the internet disappears. And it should answer before we start wondering whether the application has frozen. That was the starting point of our local AI experiment for Lastbrowser. We were looking for a practical minimum: a model small enough to run…",
  "key_points": [
    "Team prioritized speed, usefulness, and reliability over model size",
    "Initial benchmarking on powerful machine, not target CPU",
    "Established consistent benchmarking framework for real-world AI use"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}