{
  "id": 12103472,
  "title": "When I would stop trying PDF extractors and ask for the source",
  "url": "https://urgent.news/2026/10/05/when-i-would-stop-trying-pdf-extractors-and-ask-for-the-source",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-05T08:06:16.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/shu_jing_915fa287b22539ad/when-i-would-stop-trying-pdf-extractors-and-ask-for-the-source-1o8b"
  },
  "original_language": "en",
  "account": "After repeatedly trying PDF extractors on a damaged notice with Chinese text, I still did not obtain the original Chinese title. Even when the outputs appeared different, they all lacked the critical Chinese characters I needed. This realization prompted a controlled test to determine if another conversion method could retrieve the missing information. Using a synthetic notice with five lines of both Chinese and English, I preserved the original, made one copy without the Chinese font s ToUnicode mapping, and created another copy with an incorrect mapping. Having the source text allowed me to directly compare the extractor outputs and identify when they deviated from the accurate information. All three extractors, including PyMuPDF 1.26.5, pypdf 6.10.2, and PDF.js 5.6.205, recovered the 56 Han characters from the original notice. However, when the Chinese font mapping was removed or altered, none of the extractors managed to recover the original Chinese characters. Some unwanted characters appeared in the outputs, varying between implementations, but they never restored the essential words. The altered copy demonstrated that the extractors agreed with an incorrect mapping, which was still incorrect despite their uniform agreement. This discrepancy highlighted the importance of having the source content to verify the actual document text. For a small project, requesting a confirmed source passage instead of a vague statement like \"the PDF opens correctly\" would be more informative. A short, correctly identified source passage provides a concrete basis for comparison and helps determine whether an import would be usable. The experiment also revealed that image OCR could recover readable content from a PNG rendering of the original notice, but the spacing differed from the source, requiring further review. ImgIng s PDF-to-image operation failed on this test, and the application could not be considered a complete chain for recovery. In a similar document-import scenario, I would inquire whether an editable source or a confirmed text export was available to ensure the semantic information needed is present. I did not measure support time or monetary savings but focused on obtaining the source content to verify the exact fields required by the requester. Asking for source text, such as the page number, a readable passage, and the extracted result, would be sufficient to evaluate the current problem without overloading the provider with an entire project archive. If the provider sent a recently exported PDF, I would repeat the same comparison to ensure consistency. If the source is no longer available, I would describe the deliverable as reviewed OCR text from selected page images, acknowledging that additional work remains, such as obtaining a correct image, recognizing it, and checking pertinent fields. This approach maintains transparency about the work still required and avoids claiming the original PDF has been repaired. Keeping the failed direct extraction alongside the reviewed result allows for tracing any later corrections back to the route that produced them. The key takeaway from this test is to employ a simple \"stopping rule\": if a second extraction route reproduces the same critical error, pause and determine what new evidence another attempt could provide. If none is available, requesting source content or agreeing on a reviewed image-based route would be the prudent course of action.",
  "summary": "After trying three PDF extractors on the same damaged notice, I still did not have the Chinese text I needed. The outputs looked different, but none contained the original Chinese title. That distinction matters when I am deciding whether another dependency is worth adding to a small tool. A different-looking failure is not yet a better deliverable. Before trying another converter, I want a…",
  "key_points": [
    "Tried PDF extractors on damaged Chinese text notice with no success",
    "Controlled test with synthetic notice showed extractors need source text",
    "Requesting source content is key to verifying document accuracy"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}