{
  "id": 12536689,
  "title": "My URL-to-Markdown extractor returned 340 characters for a 2,000-word article — the page rendered itself in JavaScript",
  "url": "https://urgent.news/2026/10/07/my-url-to-markdown-extractor-returned-340-characters-for-a-2-000-word",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-07T03:39:05.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/imapphelp/my-url-to-markdown-extractor-returned-340-characters-for-a-2000-word-article-the-page-rendered-4if5"
  },
  "original_language": "en",
  "account": "In the process of converting a 2,000-word engineering blog post into markdown format, I encountered a peculiar issue: the extractor only returned 340 characters, mainly consisting of \"Subscribe\" and a footer. Upon inspecting the raw page, I discovered that the article body was absent — the entire content was rendered using JavaScript. The page itself had an empty div with the id \"root\" and was entirely powered by JavaScript. My text-density heuristic relies on the presence of blocks of text; when the body is missing, it defaults to the largest text available, typically a cookie notice or footer. This led me to introduce a fallback mechanism: I calculated the ratio of extracted characters to the total plain-text characters in the source HTML. If the ratio fell below ~5%, it indicated that the site was rendered client-side. This single ratio proved to be far more effective at distinguishing between normal pages and single-page applications (SPAs) than relying on URL patterns, which often resulted in false positives. To address this, I implemented a fallback strategy: for low-yield extraction, the service would retry with a rendered fetch using a headless browser, allowing the extractor to process the rendered DOM and continue with the normal pipeline. While this approach improved the output quality, it came with trade-offs. The rendered pass took significantly longer, transitioning from hundreds of milliseconds to several seconds. To inform callers about the serving path, a \"rendered: true\" flag was added to each response. Sites requiring authentication or facing aggressive bot checks still faced challenges, and it was preferable to return an honest error message rather than a stub. I integrated both paths into the URL-to-Markdown API, ensuring a consistent interface for callers. The key takeaway applies to any fetch pipeline feeding a model: monitor extraction yield, and when it appears suspiciously low, assume the content was likely built using JavaScript until further evidence contradicts this assumption.",
  "summary": "A few weeks into running my conversion pipeline, I hit a strange class of failures: some articles came through beautifully, others returned a stub of navigation text and nothing else. The worst case was a 2,000-word engineering blog post that extracted to 340 characters — mostly \"Subscribe\" and a footer. Fetching the raw page told me everything: the article body simply wasn't in the HTML. There…",
  "key_points": [
    "Extractor returned only 340 characters from 2,000-word article",
    "Article body rendered using JavaScript, not present in raw page",
    "Fallback mechanism added to detect client-side rendering"
  ],
  "editors_take": "The author's tweak to detect and handle JavaScript-heavy sites improves extraction accuracy but adds latency, forcing a trade-off between speed and reliability in URL-to-Markdown processing.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}