{
  "id": 12103482,
  "title": "Keep the original text when a PDF page falls back to OCR",
  "url": "https://urgent.news/2026/10/05/keep-the-original-text-when-a-pdf-page-falls-back-to-ocr",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-05T08:03:22.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/iterandum/keep-the-original-text-when-a-pdf-page-falls-back-to-ocr-4fkg"
  },
  "original_language": "en",
  "account": "When a PDF extraction fails and falls back to OCR, it can be difficult to determine which page required assistance and where the replacement originated. To address this issue, a local experiment was conducted using controlled PDF samples. Each result was kept attached to its page and source, including those that appeared incorrect. This approach was tested on a notice generated with ReportLab, with separate versions created by altering the Unicode mapping and incorrect mapping. Despite visual similarity, the extracted text from these altered versions differed from the original.\n\nBy keeping each result separate, it became possible to examine the extraction result without altering the visible wording. The missing Unicode mapping removed all 56 Chinese characters from the extracted output, while the pypdf version retained 152 characters. The basic import check would have passed a nonempty string, and counting replacement characters would not have helped either since the output contained no U+FFFD.\n\nCombining a normal page and a damaged page into one PDF made the issue with document-level results obvious. The complete extraction contained Chinese text due to the first page being fine, while the second page still needed investigation. When both outputs were flattened into a single string, distinguishing between them became more challenging. However, keeping separate page records preserved this distinction without requiring an OCR scheduler or new document-processing service.\n\nEach record in the experiment included the file name, page number, raw checks, font observations, and reference comparison. The reference used was the original generated wording, not another extractor's output. Pages with obvious character anomalies received a \"no_signal\" designation, which did not imply correctness. Pages matching their complete known reference were marked as \"sample_checked,\" while pages without a complete reference remained unverified.\n\nThe distinction between verified and unverified pages mattered in a subsequent failure. Changing one mapping entry caused two occurrences of a Chinese character to be extracted as another valid character. Despite the output containing 56 Chinese characters and no replacement symbol, PDF.js, pypdf, and PyMuPDF all returned the same wrong substitutions. Basic character checks missed these mismatches, but comparing against an independent source text revealed the discrepancy. Majority voting as ground truth would have hidden this issue, so missing mappings were recorded as observations while keeping the text evidence and final decision separate.\n\nOut of the eight pages checked in the experiment, three needed review, three were checked against their known source, and two remained unverified. These unverified pages were from an existing copy of RAG-Safety-Bench, with manual inspection conducted but no complete reference available. The experiment also tested ImgIng in the browser on its Chinese site, imging.cn. Enabling and disabling image-text recognition in PDF content extraction produced identical text for the damaged PDF. Converting the file to HTML resulted in damaged Chinese text, but both outcomes were documented in the record.\n\nA successful image-recognition path existed, but the PNG came from an external renderer, and the PDF-to-image conversion attempt on this sample failed due to an SVG rendering error. Feeding the clear PNG into a separate image OCR tool successfully recovered the notice's five lines, although the spaces inside the Chinese text differed from the original PDF. The interface count of 134 characters did not represent an accuracy measurement, and the screenshot shows the image OCR result with an English interface. The notice itself remained Chinese, as it was the tested input, and this experiment documented a different extraction source rather than a repair to the original PDF's mapping.\n\nIf a fallback integration were implemented, it would retain the original page record and attach the OCR output as an additional candidate with its own source and comparison scope. This integration is a proposal and has not been implemented, as the experiment focused on testing and understanding the issue rather than providing a solution.",
  "summary": "A PDF extraction failure becomes harder to investigate when the fallback result replaces the first output. You eventually have readable text, but cannot tell which page needed help or where the replacement came from. I tested a small recording approach using controlled PDF samples. The useful change was keeping each result attached to its page and source, including results that looked wrong. This…",
  "key_points": [
    "Each result kept separate to examine extraction without altering visible wording",
    "Missing Unicode mapping removed 56 Chinese characters from extraction output",
    "Keeping separate page records preserved distinction without OCR scheduler"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}