{
  "id": 12130924,
  "title": "I benchmarked my website-to-Markdown crawler against Apify's Website Content Crawler on 5 real sites",
  "url": "https://urgent.news/2026/10/05/i-benchmarked-my-website-to-markdown-crawler-against-apifys-website",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-05T10:45:36.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/tidytools/i-benchmarked-my-website-to-markdown-crawler-against-apifys-website-content-crawler-on-5-real-sites-oee"
  },
  "original_language": "en",
  "account": "This report outlines the results of a benchmark test comparing the Website Markdown Crawler (MMC) and Apify's Website Content Crawler (WCC) on five different websites. Both crawlers were run with the same start URL and path scope, and were capped at 25 pages each.\n\nThe five websites tested were a Python tutorial, Sphinx static HTML site, Docusaurus static site, Next.js React site, and Intercom Help Center. Each site was designed differently to test the crawlers on various content structures.\n\nThe testing compared several metrics, including the number of headings, paragraphs, and code blocks extracted, as well as the occurrence of template lines (repeating lines such as menus and footers). The test also measured the time taken for each crawler to complete and the associated cost.\n\nOne key finding was that WCC demonstrated superior content quality by removing page chrome, such as menus and footers, resulting in higher quality extracted content. In contrast, MMC left more template lines on certain pages, particularly the Intercom Help Center site.\n\nAnother significant issue was that WCC struggled with Next.js documentation, often outputting only the version picker and failing to extract code blocks. MMC, on the other hand, was able to extract full text for Next.js but had issues with fenced code blocks and escaping.\n\nAfter identifying these shortcomings, the author made improvements to MMC, such as focusing on main content extraction, handling code blocks, tables and links more robustly, and adjusting the crawl order to optimize for different site structures. These changes resulted in significant improvements in extracted content quality and quantity for the majority of sites tested.\n\nIn summary, while both crawlers performed adequately on the Python tutorial and Docusaurus sites, MMC demonstrated superior content quality and handled Next.js documentation more effectively after optimization. WCC still showed advantages in certain areas, particularly with Intercom Help Center and Cloudflare Blog. The benchmarks highlight the importance of tailored crawler configurations for different website types to achieve optimal results.",
  "summary": "Disclosure first: I built one of the two tools in this test, Website Markdown Crawler . The other one is Apify's official Website Content Crawler (WCC), which about 12,000 people use every month and which has 234 reviews averaging 4.6 stars. I have a conflict of interest, so I'm publishing the method and every per-run number with this post ( benchmark page ). You can check the numbers and tell me…",
  "key_points": [
    "WCC showed superior content quality by removing page chrome",
    "MMC struggled with Next.js documentation and code blocks",
    "MMC improved content extraction after targeted optimizations"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}