I benchmarked my website-to-Markdown crawler against Apify's Website Content Crawler on 5 real sites
Disclosure first: I built one of the two tools in this test, Website Markdown Crawler . The other one is Apify's official Website Content Crawler (WCC), which about 12,000 people use every month and which has 234 reviews averaging 4.6 stars. I have a conflict of interest, so I'm publishing the method and every per-run number with this post ( benchmark page ). You can check the numbers and tell me…
This report outlines the results of a benchmark test comparing the Website Markdown Crawler (MMC) and Apify's Website Content Crawler (WCC) on five different websites. Both crawlers were run with the same start URL and path scope, and were capped at 25 pages each.
The five websites tested were a Python tutorial, Sphinx static HTML site, Docusaurus static site, Next.js React site, and Intercom Help Center. Each site was designed differently to test the crawlers on various content structures.
The testing compared several metrics, including the number of headings, paragraphs, and code blocks extracted, as well as the occurrence of template lines (repeating lines such as menus and footers). The test also measured the time taken for each crawler to complete and the associated cost.
One key finding was that WCC demonstrated superior content quality by removing page chrome, such as menus and footers, resulting in higher quality extracted content. In contrast, MMC left more template lines on certain pages, particularly the Intercom Help Center site.
Another significant issue was that WCC struggled with Next.js documentation, often outputting only the version picker and failing to extract code blocks. MMC, on the other hand, was able to extract full text for Next.js but had issues with fenced code blocks and escaping.
After identifying these shortcomings, the author made improvements to MMC, such as focusing on main content extraction, handling code blocks, tables and links more robustly, and adjusting the crawl order to optimize for different site structures. These changes resulted in significant improvements in extracted content quality and quantity for the majority of sites tested.
In summary, while both crawlers performed adequately on the Python tutorial and Docusaurus sites, MMC demonstrated superior content quality and handled Next.js documentation more effectively after optimization. WCC still showed advantages in certain areas, particularly with Intercom Help Center and Cloudflare Blog. The benchmarks highlight the importance of tailored crawler configurations for different website types to achieve optimal results.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.