{
  "id": 10547190,
  "title": "Building a text checker: Unicode counts, source offsets, and stale results",
  "url": "https://urgent.news/2026/09/28/building-a-text-checker-unicode-counts-source-offsets-and-stale",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-28T22:22:50.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/mylistenapp/building-a-text-checker-unicode-counts-source-offsets-and-stale-results-34ao"
  },
  "original_language": "en",
  "account": "Developing a text checker involves navigating various design considerations. One challenge is understanding how JavaScript strings and textarea selections function. JavaScript uses UTF-16 code units, which means an emoji can occupy two code units even though it represents a single Unicode code point. For instance, the code sample 😀你好 would be represented as four UTF-16 code units, but only three Unicode code points. Therefore, when calculating displayed character counts, the checker uses Array.from(text).length instead of relying solely on JavaScript string lengths. This distinction is crucial, as using the displayed count as a selection offset could lead to bugs.\n\nAnother consideration is handling line endings consistently. The tool calculates line starts using a specific split operation that accounts for different newline characters (\\r\\n, \\r, and \\n). This approach ensures that the separator length is preserved rather than assuming every newline occupies one code unit. The line-start calculation is essential for selecting flagged lines accurately. For example, to select a flagged line, the function uses starts[i] through starts[i] + lines[i].length, which accounts for the actual separator length.\n\nThe checker also needs to invalidate results whenever the source changes. This is particularly important for tools that run checks on user input triggered by a button click. By hiding the old result panel immediately after a change and prompting for a fresh check, the tool ensures that each displayed diagnostic corresponds to one exact version of the input. For a background or asynchronous checker, additional mechanisms like a revision number or guard would be necessary to prevent older analyses from overwriting newer results. The current implementation is synchronous, so this extra layer of protection was not implemented.\n\nTo prevent unexpected document edits, the tool caps its work at 100,000 code points. If a longer document is pasted, the checker displays a limit message and leaves the entire textarea value intact without truncation or slicing. This decision is both a performance consideration and a product decision, ensuring that users are not surprised by unexpected document modifications. Additionally, the tool caps the rendered suggestions at 80, while still displaying the total number detected and explaining that longer documents can be checked in sections. This approach bounds the amount of result UI while preserving the integrity of the source text.",
  "summary": "A small textarea tool can fail in surprisingly visible ways: an emoji shifts the highlighted line, an edit makes the suggestions point at old text, or an input limit silently removes the end of a document. We ran into these design questions while building a browser-only preparation tool for text people intend to listen to. It flags a few patterns worth reviewing: long paragraphs, URLs, Markdown…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}