Urgent.News

What's breaking now, across thousands of outlets.

Tech

When I would stop trying PDF extractors and ask for the source

After trying three PDF extractors on the same damaged notice, I still did not have the Chinese text I needed. The outputs looked different, but none contained the original Chinese title. That distinction matters when I am deciding whether another dependency is worth adding to a small tool. A different-looking failure is not yet a better deliverable. Before trying another converter, I want a…

After repeatedly trying PDF extractors on a damaged notice with Chinese text, I still did not obtain the original Chinese title. Even when the outputs appeared different, they all lacked the critical Chinese characters I needed. This realization prompted a controlled test to determine if another conversion method could retrieve the missing information.

Using a synthetic notice with five lines of both Chinese and English, I preserved the original, made one copy without the Chinese font s ToUnicode mapping, and created another copy with an incorrect mapping. Having the source text allowed me to directly compare the extractor outputs and identify when they deviated from the accurate information.

All three extractors, including PyMuPDF 1.26.5, pypdf 6.10.2, and PDF.js 5.6.205, recovered the 56 Han characters from the original notice. However, when the Chinese font mapping was removed or altered, none of the extractors managed to recover the original Chinese characters. Some unwanted characters appeared in the outputs, varying between implementations, but they never restored the essential words.

The altered copy demonstrated that the extractors agreed with an incorrect mapping, which was still incorrect despite their uniform agreement. This discrepancy highlighted the importance of having the source content to verify the actual document text. For a small project, requesting a confirmed source passage instead of a vague statement like "the PDF opens correctly" would be more informative.

A short, correctly identified source passage provides a concrete basis for comparison and helps determine whether an import would be usable. The experiment also revealed that image OCR could recover readable content from a PNG rendering of the original notice, but the spacing differed from the source, requiring further review. ImgIng s PDF-to-image operation failed on this test, and the application could not be considered a complete chain for recovery.

In a similar document-import scenario, I would inquire whether an editable source or a confirmed text export was available to ensure the semantic information needed is present. I did not measure support time or monetary savings but focused on obtaining the source content to verify the exact fields required by the requester. Asking for source text, such as the page number, a readable passage, and the extracted result, would be sufficient to evaluate the current problem without overloading the provider with an entire project archive.

If the provider sent a recently exported PDF, I would repeat the same comparison to ensure consistency. If the source is no longer available, I would describe the deliverable as reviewed OCR text from selected page images, acknowledging that additional work remains, such as obtaining a correct image, recognizing it, and checking pertinent fields.

This approach maintains transparency about the work still required and avoids claiming the original PDF has been repaired. Keeping the failed direct extraction alongside the reviewed result allows for tracing any later corrections back to the route that produced them. The key takeaway from this test is to employ a simple "stopping rule": if a second extraction route reproduces the same critical error, pause and determine what new evidence another attempt could provide.

If none is available, requesting source content or agreeing on a reviewed image-based route would be the prudent course of action.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Building a Password Generator in Python

📌 Quick Info Topic: A beginner project using Python's built-in libraries Target Audience: Beginners who want a second project Goal: Build something useful in under 40 lines 1.

  • Python tutorial builds password generator in under 40 lines
  • CLI tool creates robust, random passwords in under a minute
  • Uses random, string modules without additional installation

Feature Flag SDK Design for Multi-Language Consistency and Performance

You see inconsistent experiment numbers, customers who get different behavior on mobile vs server, and alerts that point to "the flag" — but not which SDK made the wrong call.

  • Enforce deterministic evaluation with canonical JSON and SHA-256 hash function
  • Optimize initialization with non-blocking default path and blocking option
  • Implement reliable updates via streaming with SSE and resilient reconnection

The best engineer on my team ships the least code.

The best engineer on my team almost got a bad review last quarter for shipping too little code. I'm not being cute. I pulled the numbers before his review because I knew the numbers were going to be…

  • Engineers evaluated based on code output, not true value
  • Cheap code production masks scarce review skills
  • Review process fails to capture prevention-focused engineers

More from Monday 5 October →