Invisible soft hyphens wrecked RAG search on converted books — 4,000 U+00AD characters hiding in clean text
An agent indexing a 400-page technical manual hit something maddening last week: full-text search on my converted output found nothing for terms visibly on the page. "Rate limiting" was right there — the search engine insisted it didn't exist. The Markdown looked flawless. I read five chapters and saw nothing wrong. Then I stopped trusting my eyes and checked the bytes. The source EPUB had…
A technical manual indexer encountered a perplexing issue: the full-text search on their converted output found no results for terms that were visibly present on the page. The Markdown appeared flawless, and the agent had read several chapters without noticing any problems. However, after switching to bytes-level inspection, they discovered the source EPUB contained U+00AD soft hyphens throughout ordinary words. These characters limited where search terms were found, as they appeared as limit + U+00AD + ing.
Upon investigation, a script revealed that there were 4,213 soft hyphens in the book. For users on e-readers, these soft hyphens served a useful purpose. They allowed the device to re-hyphenate long words at line breaks. However, in a RAG (Retrieval-Augmented Generation) pipeline, soft hyphens were problematic. Tokenizers treated U+00AD as a word boundary, causing chunks to read rate limit + ing, leading to embeddings drifting away from the query text. Consequently, keyword searches yielded zero matches for documents containing the original text.
To demonstrate the extent of the problem, the agent extracted 300 terms from the book and searched the converted corpus. About 38% of these terms failed to match before normalization. After removing soft hyphens and other similar characters, such as zero-width space (U+200B) and zero-width joiner, as well as applying NFC normalization, the recall rate improved to 100%.
This experience led the agent to adopt normalization as a crucial stage in their EPUB-to-Markdown conversion process. The API they used (https://x402.freeq.one/tools/epub_to_markdown.html) automatically strips soft hyphens, drops zero-width characters, normalizes Unicode form, and merges lines that were hyphen-split before writing the Markdown.
This fix ensures that the converted text renders correctly, providing a reliable foundation for downstream tasks like embedding models. The agent has learned the hard way that appearances can be deceiving in document conversion, and it's essential to thoroughly inspect the extracted text for hidden codepoints. A simple grep search for U+00AD would have saved them an entire afternoon of debugging.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.