I tried 13 OCR preprocessing tricks on screenshots. None helped
Almost every OCR tip I found said the same thing. Before recognizing a screenshot, clean it up: go grayscale, binarize it, run a denoise filter, maybe upscale it 2x. I had never checked whether any of that actually helps, so I took one small set of screenshots and ran every trick on the list, one at a time, against text I already knew. On clean screenshots, not a single step beat the untouched…
Thirteen different OCR preprocessing techniques were tested on screenshots, but none of them improved the results. The screenshots consisted of various text sizes and colors, including a made-up Chinese notice about community service hours and a holiday notice template. The text was captured at 1x and 2x device pixel ratios. ImgIng was used as the OCR engine, with three tiers: Professional, Fast, and Ultimate.
The benchmark engine was tesseract.js 5 with default parameters and chi_sim+eng. All OCR runs were performed locally using a Chromium 149 open-source build on an M4 Mac. The character error rate (CER) was calculated as the edit distance divided by the length of the answer, summed over 10 crops. The untouched original screenshot achieved a CER of 0.1% on the Professional OCR tier.
Several preprocessing techniques, such as grayscale conversion, upscaling, and various binarization methods, either matched or exceeded the original CER or even worsened the results. For example, a 3×3 median filter caused a CER of 9.7%, while a threshold of 160 resulted in a CER of 9.4%. Unexpectedly, the grayscale conversion, despite being a "gentle" technique, transformed lighter gray text into white, leading to an empty result for that strip.
The median filter's negative impact was more severe, increasing the CER from 0% to 32.4% on the Professional OCR tier. The tier differences were minimal, with Fast, Professional, and Ultimate OCR tiers scoring 0.8%, 0.1%, and 0.7%, respectively, on the original screenshots. Threshold 128 proved detrimental to several OCR variants, increasing their CER by approximately 25 points.
Vertical text, a challenging scenario, was not helped by any of the preprocessing techniques. The rotate-left-90° function in the image adjustments was the only solution that improved recognition accuracy from 92.3% to 0% on the Professional OCR tier. Overall, the report suggests that recognizing the untouched file first, keeping it as a baseline, and measuring the darkest gray text before applying thresholds are crucial steps.
When dealing with vertically oriented text, rotating the image before applying any preprocessing is recommended.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.