My URL-to-Markdown extractor returned 340 characters for a 2,000-word article — the page rendered itself in JavaScript
A few weeks into running my conversion pipeline, I hit a strange class of failures: some articles came through beautifully, others returned a stub of navigation text and nothing else. The worst case was a 2,000-word engineering blog post that extracted to 340 characters — mostly "Subscribe" and a footer. Fetching the raw page told me everything: the article body simply wasn't in the HTML. There…
In the process of converting a 2,000-word engineering blog post into markdown format, I encountered a peculiar issue: the extractor only returned 340 characters, mainly consisting of "Subscribe" and a footer. Upon inspecting the raw page, I discovered that the article body was absent — the entire content was rendered using JavaScript.
The page itself had an empty div with the id "root" and was entirely powered by JavaScript. My text-density heuristic relies on the presence of blocks of text; when the body is missing, it defaults to the largest text available, typically a cookie notice or footer. This led me to introduce a fallback mechanism: I calculated the ratio of extracted characters to the total plain-text characters in the source HTML.
If the ratio fell below ~5%, it indicated that the site was rendered client-side. This single ratio proved to be far more effective at distinguishing between normal pages and single-page applications (SPAs) than relying on URL patterns, which often resulted in false positives. To address this, I implemented a fallback strategy: for low-yield extraction, the service would retry with a rendered fetch using a headless browser, allowing the extractor to process the rendered DOM and continue with the normal pipeline.
While this approach improved the output quality, it came with trade-offs. The rendered pass took significantly longer, transitioning from hundreds of milliseconds to several seconds. To inform callers about the serving path, a "rendered: true" flag was added to each response. Sites requiring authentication or facing aggressive bot checks still faced challenges, and it was preferable to return an honest error message rather than a stub.
I integrated both paths into the URL-to-Markdown API, ensuring a consistent interface for callers. The key takeaway applies to any fetch pipeline feeding a model: monitor extraction yield, and when it appears suspiciously low, assume the content was likely built using JavaScript until further evidence contradicts this assumption.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.