Urgent.News

What's breaking now, across thousands of outlets.

Tech

Why PDF converters mangle Cyrillic (and how to pick one that doesn't)

If you work with Russian, Ukrainian, Bulgarian or any Cyrillic-language documents, you know the special failure: the layout survives perfectly, and the text inside reads as squares, mojibake, or question marks. The document is useless — you cannot search it, spell-check it, or copy a single correct name out of it. Why it happens A PDF does not have to store text at all — it can store drawings of…

When working with documents in Cyrillic languages, you may have experienced the frustrating issue of layout remaining intact, but the text displaying as squares, mojibake, or question marks. This renders the document useless for searching, spell-checking, or copying accurate information. The reason behind this lies in the way PDFs store text.

Unlike storing actual text, PDFs can store drawings of letters, which are then mapped to character codes using an internal encoding table. When converting PDFs to other formats, good converters reconstruct this table, resulting in proper Unicode text. However, poor converters assume a single Western encoding, causing every Cyrillic letter to transform into the corresponding Latin character, leading to the classic mojibake.

The layout engine itself does not impact this issue, as even if the text mapping fails, a well-designed paragraph layout will still produce garbage. Another contributing factor is the usage of fonts lacking proper Unicode tables within the PDF, which can be attributed to older scanners or documents exported from Word during the 2000s. Lastly, OCR pipelines that recognize only Latin scripts can lead to this problem, as Cyrillic text is treated as the closest Latin shape, resulting in plausible-looking but incorrect words.

Before entrusting any converter with important work, perform a quick test using a sample document. Copy a sentence, paste it into a search box, and verify if the known word appears in the search results. If it doesn't, the text layer is broken, regardless of how aesthetically pleasing the output may appear. Additionally, try a mixed-alphabet test using a line containing both Cyrillic and Latin characters, and finally, examine a table with Cyrillic headers to ensure the converter maintains the integrity of tables and headers. This test can save you valuable time and prevent headaches before it's too late.

Our own PDF to Word converter is designed to maintain Cyrillic and Latin characters intact, preserve tables, and process documents locally on your machine. However, this test method works with any converter, including free options, to ensure the tool respects your documents before it affects your deadlines. This article was written by Yodsira, who developed a plugin specifically to address these issues.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Getting Google Trends data in Python without fighting 429 errors

I built this actor; it's a paid tool on Apify with a free trial credit. If you've used pytrends , you've probably seen 429 Too Many Requests .

  • Actor developed to bypass Google Trends rate limits in Python and Node
  • Handles Google's throttling with unique cookies per session
  • Saves results per search, preventing blockage of other searches

Check Markdown examples before someone copies them

A project can have a green test suite while its README example is stale. The docs are often the first thing a new user runs, and one missing colon or moved setup link can waste the first few minutes.

  • Snipverify checks Markdown syntax in code blocks and links
  • Reports file and line for invalid syntax, skips external URLs
  • JSON mode for CI integration with file, line, status, message

More from Saturday 3 October →