Urgent.News

What's breaking now, across thousands of outlets.

Tech

My URL-to-Markdown extractor returned a cookie banner as the article body — text density has no taste

Most of my document converters take a file you already have. The URL-to-Markdown one is different: it fetches the page itself and has to decide what the "main content" even is before converting anything. That decision taught me my most humbling lesson so far. My first extraction heuristic was pure text density: score every candidate block by how much text it contains, pick the winner, convert it.…

Most of my document converters require a file you already possess. However, the URL-to-Markdown converter is unique in that it retrieves the web page itself, determining the main content before any conversion takes place. This challenge taught me a crucial lesson. Initially, I employed a simple heuristic based on text density, scoring each candidate block by its text content and selecting the highest-scoring one for conversion.

This method worked well on blogs, documents, and news pages. Yet, when I processed a batch of real-world URLs through it, a news site returned a 600-word GDPR consent dialog instead of the article. The actual story consisted of around 300 words in a narrow column, while the consent banner contained more text than the article itself.

It became evident that text density alone cannot distinguish between articles and legal boilerplate, as both appear as large blocks of text. To enhance accuracy, I introduced two additional signals. First, I analyzed link density; navigation headers, footers, and cookie banners are often filled with anchor links, while genuine article text primarily consists of prose.

Penalizing blocks where anchor text exceeds approximately 30% of the total text helped eliminate many false positives. Second, I considered semantic anchors; when a single "article" or "main" element exists, it outperforms any heuristic score. Additionally, I maintain a limited denylist of id/class substrings—such as "consent," "cookie," "banner," and "paywall"—that automatically disqualify candidates.

This approach proved effective. Furthermore, I recognized a third important aspect: single-page applications. Fetching a client-rendered page via a simple HTTP request yields a nearly empty div, making it impossible for any scoring method to extract meaningful text when the server never provided it in the first place. In such cases, I now return an explicit "no extractable content" flag instead of an empty Markdown document, notifying callers to retry with a rendering step.

For pages whose static HTML primarily consists of script tags, this feature is particularly useful. As a result, I packaged the fetch-and-clean functionality into a small API: https://x402.freeq.one/tools/markdown.html. However, the heuristic details I've mentioned above constitute the essential transferable part. The same scoring techniques can be applied to any HTML that requires cleaning before being indexed.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Sunday 4 October →