{
  "id": 13360094,
  "title": "Split your OCR error rate in two before you touch preprocessing",
  "url": "https://urgent.news/2026/10/10/split-your-ocr-error-rate-in-two-before-you-touch-preprocessing",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-10T08:01:49.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/pm_cheng_3f36acecfb9c59f5/split-your-ocr-error-rate-in-two-before-you-touch-preprocessing-5b3f"
  },
  "original_language": "en",
  "account": "In September, a test was conducted to compare the performance of 14 different OCR preprocessing methods on a single crop containing vertical Chinese text. The text achieved a character error rate (CER) between 89.7% and 94.9% across all versions, while the original 1x crop scored 92.3% on all three OCR tiers. This inconsistency indicated that pixel quality may not be the primary factor affecting OCR accuracy.\n\nThe OCR samples were taken from a custom web page featuring community service hour information, with additional crops from a holiday notice template obtained online. The recognition was performed using ImgIng, Fast OCR, and Ultimate OCR on an Apple M4 device with a Chromium 149 build, employing the WebGPU backend. Tesseract.js 5 (default parameters) served as the control for comparison.\n\nUpon analyzing the results, it became apparent that the order of characters played a significant role in determining the OCR accuracy. A new metric was introduced to account for this, treating both the reference and hypothesis texts as bags of characters, and counting the shared characters while calculating the remaining errors. High CER with low order-free rates indicated that characters were correct but in the wrong sequence, while high values for both suggested that text was lost or misread.\n\nAfter identifying the importance of character order, the test proceeded by evaluating the impact of rotation and tilt on recognition accuracy. Vertical text was rotated left by 90°, resulting in zero errors for all three OCR tiers on the 2x crop, while only Fast and Professional OCR made mistakes on the 1x crop (2 punctuation marks error for Fast and no errors for Ultimate). Tesseract.js continued to perform poorly on both rules, regardless of the applied preprocessing.",
  "summary": "On September 30 I ran an OCR preprocessing benchmark with 14 versions of each crop: the original, the built-in Scan enhancement, 2x upscaling, grayscale, four fixed thresholds, Otsu, adaptive thresholding, two denoisers and two JPEG qualities. One row barely moved. A small block of vertical Chinese text sat between 89.7% and 94.9% character error rate in every version on the Professional tier,…",
  "key_points": [
    "Vertical Chinese text shows CER between 89.7% and 94.9% across OCR methods",
    "Order of characters significantly impacts OCR accuracy, introducing new metric",
    "Rotating vertical text left by 90° yields zero errors on 2x crop"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}