{
  "id": 11047506,
  "title": "In-browser machine translation: I measured 30 phrases, found 30% wrong in meaning, and changed the engine",
  "url": "https://urgent.news/2026/09/30/in-browser-machine-translation-i-measured-30-phrases-found-30-wrong",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-30T22:30:05.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/convertilo/in-browser-machine-translation-i-measured-30-phrases-found-30-wrong-in-meaning-and-changed-the-1146"
  },
  "original_language": "en",
  "account": "Our free translator ran a small machine-translation model directly in the browser using Bergamot / Opus-MT style code compiled to WebAssembly. The main appeal is privacy, as the text never leaves the device. However, the downside is quality. To measure the model's accuracy, I tested 30 phrases in both English to Russian and Russian to English directions. Results showed 43% correct translations, 26% with minor errors, and a concerning 30% where the meaning was completely altered.\n\nConnected prose and business text translated well in both directions. Technical text, on the other hand, had issues such as inventing words - for example, \"driver\" was translated as \"водитель,\" which refers to a vehicle driver rather than software driver. Idioms and casual speech failed completely; phrases like \"I'll hit you up\" were translated as \"I will hit you,\" which can be interpreted as a threat. Ambiguities were also a problem; a legal sentence like \"the bank refused the loan\" was translated with the prohibition reversed. These issues are particularly concerning for legal documents.\n\nTo further test the model, I ran the same 30 phrases through eight different models via an API. The benchmark cost approximately $0.013. The hardest phrases for the in-browser model to translate were tested on the other models. The results showed that the cheapest model, Qwen3 30B, was the least accurate, translating phrases like \"Klöße\" as \"cutlets\" and \"Vorspeisen\" as \"first courses.\" More expensive models with larger parameters did not necessarily perform better for this specific task.\n\nReasoning models proved to be a trap for translation tasks. One of the models returned empty text for all 30 phrases because it spent the entire token budget on thinking and never reached a conclusion. Simply increasing the token limit would mean paying for more expensive tokens. Ultimately, I chose a model that fixed the problematic translations, with a cost of roughly $0.16 per million source characters. This model successfully handled cases like the threatening \"hit you\" and the inverted legal sentence.\n\nWhile moving translation to a server model means that the text no longer remains on the device, it does provide a higher quality translation for everyday text. However, this comes at the cost of privacy. If privacy is the main concern, the small local model remains the best option. If quality on everyday text is more important, the in-browser model may not be sufficient. Users can compare the performance of different translation models for themselves by accessing the browser translator and an English to Russian page.",
  "summary": "Our free translator ran a small machine-translation model directly in the browser (a Bergamot / Opus-MT style model compiled to WebAssembly). The appeal is privacy: the text never leaves the device. The downside is quality. I finally stopped guessing and measured it on the live site with 30 phrases. The test 30 phrases in both directions between English and Russian, grouped by type. I marked each…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}