Urgent.News

What's breaking now, across thousands of outlets.

Tech

My URL-to-Markdown extractor returned 340 characters for a 2,000-word article — the page rendered itself in JavaScript

A few weeks into running my conversion pipeline, I hit a strange class of failures: some articles came through beautifully, others returned a stub of navigation text and nothing else. The worst case was a 2,000-word engineering blog post that extracted to 340 characters — mostly "Subscribe" and a footer. Fetching the raw page told me everything: the article body simply wasn't in the HTML. There…

In the process of converting a 2,000-word engineering blog post into markdown format, I encountered a peculiar issue: the extractor only returned 340 characters, mainly consisting of "Subscribe" and a footer. Upon inspecting the raw page, I discovered that the article body was absent — the entire content was rendered using JavaScript.

The page itself had an empty div with the id "root" and was entirely powered by JavaScript. My text-density heuristic relies on the presence of blocks of text; when the body is missing, it defaults to the largest text available, typically a cookie notice or footer. This led me to introduce a fallback mechanism: I calculated the ratio of extracted characters to the total plain-text characters in the source HTML.

If the ratio fell below ~5%, it indicated that the site was rendered client-side. This single ratio proved to be far more effective at distinguishing between normal pages and single-page applications (SPAs) than relying on URL patterns, which often resulted in false positives. To address this, I implemented a fallback strategy: for low-yield extraction, the service would retry with a rendered fetch using a headless browser, allowing the extractor to process the rendered DOM and continue with the normal pipeline.

While this approach improved the output quality, it came with trade-offs. The rendered pass took significantly longer, transitioning from hundreds of milliseconds to several seconds. To inform callers about the serving path, a "rendered: true" flag was added to each response. Sites requiring authentication or facing aggressive bot checks still faced challenges, and it was preferable to return an honest error message rather than a stub.

I integrated both paths into the URL-to-Markdown API, ensuring a consistent interface for callers. The key takeaway applies to any fetch pipeline feeding a model: monitor extraction yield, and when it appears suspiciously low, assume the content was likely built using JavaScript until further evidence contradicts this assumption.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

🚀 Deep Dive into Maliklang V4: Space Optimization & Self-Healing Architecture

Hello DEV Community! 👋 I want to share the architectural breakthrough behind the latest version of my open-source project: Maliklang V4.

  • Maliklang V4 focuses on space optimization and self-healing in distributed ledger systems.
  • Implements autonomous resilience layer to detect and correct runtime anomalies.
  • Open-source project with GitHub repository for community collaboration.

Docker permission denied: why it happens and how to fix it properly

Docker permission denied: what the error is actually telling you A client messages me about this error almost every week. They run a command, see "permission denied", and assume Docker is broken.

  • "Docker permission denied" has two meanings: before container start and inside running container
  • Permission errors occur due to mismatch between process UID/GID and file/socket ownership
  • Fix errors by adjusting file permissions, ownership, or Docker daemon configuration

More from Wednesday 7 October →