Urgent.News

What's breaking now, across thousands of outlets.

Tech

I benchmarked my website-to-Markdown crawler against Apify's Website Content Crawler on 5 real sites

Disclosure first: I built one of the two tools in this test, Website Markdown Crawler . The other one is Apify's official Website Content Crawler (WCC), which about 12,000 people use every month and which has 234 reviews averaging 4.6 stars. I have a conflict of interest, so I'm publishing the method and every per-run number with this post ( benchmark page ). You can check the numbers and tell me…

This report outlines the results of a benchmark test comparing the Website Markdown Crawler (MMC) and Apify's Website Content Crawler (WCC) on five different websites. Both crawlers were run with the same start URL and path scope, and were capped at 25 pages each.

The five websites tested were a Python tutorial, Sphinx static HTML site, Docusaurus static site, Next.js React site, and Intercom Help Center. Each site was designed differently to test the crawlers on various content structures.

The testing compared several metrics, including the number of headings, paragraphs, and code blocks extracted, as well as the occurrence of template lines (repeating lines such as menus and footers). The test also measured the time taken for each crawler to complete and the associated cost.

One key finding was that WCC demonstrated superior content quality by removing page chrome, such as menus and footers, resulting in higher quality extracted content. In contrast, MMC left more template lines on certain pages, particularly the Intercom Help Center site.

Another significant issue was that WCC struggled with Next.js documentation, often outputting only the version picker and failing to extract code blocks. MMC, on the other hand, was able to extract full text for Next.js but had issues with fenced code blocks and escaping.

After identifying these shortcomings, the author made improvements to MMC, such as focusing on main content extraction, handling code blocks, tables and links more robustly, and adjusting the crawl order to optimize for different site structures. These changes resulted in significant improvements in extracted content quality and quantity for the majority of sites tested.

In summary, while both crawlers performed adequately on the Python tutorial and Docusaurus sites, MMC demonstrated superior content quality and handled Next.js documentation more effectively after optimization. WCC still showed advantages in certain areas, particularly with Intercom Help Center and Cloudflare Blog. The benchmarks highlight the importance of tailored crawler configurations for different website types to achieve optimal results.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Klaviyo vs Shopify Email: When to Switch

Stay on Shopify Email, now sold as Shopify Messaging, while 10,000 free emails a month covers your sends and your automations fit its templates (Shopify help centre, 4 October 2026).

  • Shopify Email offers free email marketing up to 10,000 emails/month.
  • Klaviyo provides native Shopify integration and advanced automation features.
  • Klaviyo's Starter plan starts at $60/month for 250 profiles and 500 emails/month.

When Skipping an API Error Can Delete Valid Data: A Cartography Fix

TL;DR: I fixed a BigQuery sync failure in Cartography, then narrowed the error tolerance after review showed that skipping failed enumeration could feed false deletions into graph cleanup.

  • Skipping API errors can delete valid Cartography data
  • Fixed error tolerance policy to prevent data loss
  • Adjusted handlers for table details, datasets, and lists

More from Monday 5 October →