Urgent.News

What's breaking now, across thousands of outlets.

Tech

Guarding OCR Costs With Node.js Hash Dedupe for Supplier Images

TL;DR: In a Node.js and Express catalogue importer, dedupe supplier images by using normalized metadata as a quick filter and a SHA-256 hash of the original bytes as the final decision. Store one immutable source object per digest, cache OCR output by that digest plus an extractor version, and record why each upload was accepted or reused. This keeps repeated storage and processing out of the…

In a Node.js and Express catalogue importer, optimize supplier image processing by deduplicating images using normalized metadata and SHA-256 hashes. Store a single immutable source object per unique hash, cache OCR output linked to the hash and extractor version, and document the rationale for accepting or reusing each upload. This approach limits unnecessary storage and processing while distinguishing true duplicates from visually similar files.

Before: Each import row generates a new image object and an OCR job.

After: Rapid metadata filtering narrows the search, a precise hash confirms identity, and only a distinct hash triggers further processing.

Suppliers often reuse names like "front.jpg" and rename the same file. Two distinct product photos may share the same MIME type, byte length, dimensions, and file type, making metadata an effective initial filter but insufficient for definitive identity verification.

MIME type declarations are unreliable inputs. Inspect the file signature directly before decoding, then determine dimensions using a pixel and memory-efficient image parser. Consult MDN's image format guide for common formats, but ultimately, the importer determines which formats it supports and can decode.

When hashing, use the raw upload bytes without resizing, stripping metadata, or transcoding. Derivative processing should have separate identities: sourceDigest + transformVersion for processed pixels, and sourceDigest + extractorVersion for OCR text. Consider a single supplier row with identical dimensions but different filenames, or a newly compressed version of the same package.

Metadata can quickly group the first two as candidates, but only their hashes can determine reuse. The third version remains distinct, preserving the original input for later OCR analysis.

Implement a two-stage intake process with a compact Express example using in-memory bodies for clarity. Limit the sample policy to 20 MiB, but for higher limits, stream the file to a quarantined temporary file during hashing, inspect, and promote it atomically.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Tuesday 6 October →