Urgent.News

What's breaking now, across thousands of outlets.

Tech

Invisible soft hyphens wrecked RAG search on converted books — 4,000 U+00AD characters hiding in clean text

An agent indexing a 400-page technical manual hit something maddening last week: full-text search on my converted output found nothing for terms visibly on the page. "Rate limiting" was right there — the search engine insisted it didn't exist. The Markdown looked flawless. I read five chapters and saw nothing wrong. Then I stopped trusting my eyes and checked the bytes. The source EPUB had…

A technical manual indexer encountered a perplexing issue: the full-text search on their converted output found no results for terms that were visibly present on the page. The Markdown appeared flawless, and the agent had read several chapters without noticing any problems. However, after switching to bytes-level inspection, they discovered the source EPUB contained U+00AD soft hyphens throughout ordinary words. These characters limited where search terms were found, as they appeared as limit + U+00AD + ing.

Upon investigation, a script revealed that there were 4,213 soft hyphens in the book. For users on e-readers, these soft hyphens served a useful purpose. They allowed the device to re-hyphenate long words at line breaks. However, in a RAG (Retrieval-Augmented Generation) pipeline, soft hyphens were problematic. Tokenizers treated U+00AD as a word boundary, causing chunks to read rate limit + ing, leading to embeddings drifting away from the query text. Consequently, keyword searches yielded zero matches for documents containing the original text.

To demonstrate the extent of the problem, the agent extracted 300 terms from the book and searched the converted corpus. About 38% of these terms failed to match before normalization. After removing soft hyphens and other similar characters, such as zero-width space (U+200B) and zero-width joiner, as well as applying NFC normalization, the recall rate improved to 100%.

This experience led the agent to adopt normalization as a crucial stage in their EPUB-to-Markdown conversion process. The API they used (https://x402.freeq.one/tools/epub_to_markdown.html) automatically strips soft hyphens, drops zero-width characters, normalizes Unicode form, and merges lines that were hyphen-split before writing the Markdown.

This fix ensures that the converted text renders correctly, providing a reliable foundation for downstream tasks like embedding models. The agent has learned the hard way that appearances can be deceiving in document conversion, and it's essential to thoroughly inspect the extracted text for hidden codepoints. A simple grep search for U+00AD would have saved them an entire afternoon of debugging.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Views Measure Views

Nine years in, I finally worked out what else to count. A writer I follow, Sylwia Laskowska , recently published a post about accidentally becoming a blogger . She has been writing on DEV for about a year, and the numbers attached to that year are impressive: hundreds of thousands of views and tens of thousands of followers.

  • Author started on DEV nine years ago, now at 48,408 total views
  • Followers grew by nearly 3,000 between Oct and mid-Sept, views grew by 1,634
  • Page views are genuine engagement measure, not vanity metric

More from Thursday 1 October →