Urgent.News

What's breaking now, across thousands of outlets.

Tech

Measure ink and lighting before you binarize an image for OCR

My team is looking at client-side OCR for an internal tool where people attach screenshots and photos of printed notices. The first preprocessing snippet anyone pastes into that kind of pipeline is grayscale plus a fixed threshold at 128. It's one line of OpenCV, so before it got near our upload flow I spent an evening measuring what it does. It fails in two quiet ways. It deletes text lighter…

Before employing any thresholding method in client-side OCR pipelines for internal tools that handle screenshots and photos of printed notices, it is crucial to measure and analyze two key factors: the darkness of the ink and the variations in paper brightness across the image.

The most common approach involves converting the image to grayscale and applying a fixed threshold at a value of 128. However, this method has two significant flaws. Firstly, it discards text that is lighter than the cutoff, effectively losing important information. Secondly, under uneven lighting conditions, the global threshold applies a uniform value across the entire image, resulting in whole regions being painted black.

These issues do not generate any errors; instead, the OCR call returns fewer characters, potentially compromising the accuracy of the recognition process.

To illustrate these problems, the author tested various scenarios using a web page they created (a made-up Chinese notice) and sampled screenshots at different resolutions, including an 11 px footer in #999 gray on #f5f5f5. The author employed ImgIng (https://imging.ai/) with its default Professional OCR tier, Fast OCR, and Ultimate OCR, comparing their performance against tesseract.js 5 with default parameters on an Apple M4 Mac.

The results showed that even with higher thresholds, the text was either completely lost or significantly degraded, particularly when dealing with colored backgrounds and uneven lighting.

To mitigate these issues, the author proposes measuring the darkness of the ink and analyzing the paper brightness drift before applying any threshold. This approach involves reading the grayscale image, defining a trimming factor to exclude the desk border, and applying local dilation to estimate the paper brightness. The code then reports the minimum ink density, the lowest paper brightness percentile, and the Otsu threshold value, along with a verdict indicating whether the fixed threshold will erase the text or if a global threshold will blacken the paper instead.

In conclusion, before implementing any thresholding technique in client-side OCR pipelines for internal tools processing screenshots and photos of printed notices, it is essential to perform a thorough audit of the ink darkness and paper brightness variation. By doing so, developers can avoid discarding crucial text and prevent recognition errors caused by uneven lighting or other artifacts.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

CF7 Real Estate Lead Capture to CRM: A Complete Implementation Guide

Real estate websites live and die by lead response time. When a potential buyer submits an inquiry about a property, every minute of delay reduces the chance of conversion.

  • CF7 form captures real estate inquiries with structured fields
  • CF7 API plugin automates data transfer to CRM system
  • Property ID and inquiry type crucial for personalized follow-up

Split your OCR error rate in two before you touch preprocessing

On September 30 I ran an OCR preprocessing benchmark with 14 versions of each crop: the original, the built-in Scan enhancement, 2x upscaling, grayscale, four fixed thresholds, Otsu, adaptive…

  • Vertical Chinese text shows CER between 89.7% and 94.9% across OCR methods
  • Order of characters significantly impacts OCR accuracy, introducing new metric
  • Rotating vertical text left by 90° yields zero errors on 2x crop

CrowdSec alert from my own IP: the 'attack' was my phone's photo app

The worry was simple: CrowdSec was "being hammered". The numbers looked alarming, and I wanted to know who was attacking my homelab. The answer, after a proper look on 21 September, was me.

  • CrowdSec flagged author's IP for suspicious thumbnail requests
  • Photo-backup app on author's phone triggered alerts
  • Author built tool to push bans to Cloudflare's edge

Launch day is the worst day to judge a product

Most launches I've watched, my own included, have the same shape. Day one is a screenshot, a one-liner and a link. And the product on that day is the least it will ever be: half the settings are…

  • First day is worst for judging product quality
  • Features incomplete, onboarding vague
  • By third week, product significantly improves

I tried 13 OCR preprocessing tricks on screenshots. None helped

Almost every OCR tip I found said the same thing. Before recognizing a screenshot, clean it up: go grayscale, binarize it, run a denoise filter, maybe upscale it 2x.

  • Thirteen OCR preprocessing techniques tested on screenshots, none improved results
  • Grayscale conversion worsened CER from 0.1% to 9.7% in Professional OCR tier
  • Vertical text recognition not helped by any preprocessing techniques

TrailWing: An Offline Bird Call Identifier That Works Where Your Signal Doesn't

What I Built Problem statement: "A bird call identifier that works on the trail with no signal." TrailWing is an open-source, offline-first birding companion.

  • TrailWing is an offline birding app for trails without cell service.
  • Users select a trail, download a trail pack of bird species, and identify birds by sound.
  • App provides species identification, confidence level, and visual confirmation tips.

More from Saturday 10 October →