OCR Uploaded Scans and Store Extracted Text in 4 Stages (With Validation)
An e-commerce document service should accept a scan, persist the private original, enqueue OCR, store extracted text under the document ID, and redact a derived copy before anybody shares it. The least complex defensible design is an asynchronous four-stage pipeline with one immutable evidence record per transition. Sign that record, not a mutable dashboard row. TL;DR: validate the upload at…
An e-commerce document service should follow a four-stage asynchronous pipeline to process an uploaded scan. The stages are: accepting the scan, persisting the original scan, performing OCR, and storing the extracted text under the document ID. Each stage should maintain an immutable evidence record. The service should validate the upload before creating a work item, rejecting empty bodies, unsupported media types, and payloads exceeding the configured limit.
The content-type is a hint, not a requirement; production validation should inspect the file signature and parse the document. The service computes a SHA-256 digest while writing the scan to private storage, then creates the document record and publishes the document ID. The original scan remains available to allow re-extraction if needed.
The OCR service is called via a verified route, with retries for 429 responses using exponential backoff. The extracted text is stored as an artifact under the document ID, with parsing into normalized text handled by a schema-specific adapter.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.