{
  "id": 11836244,
  "title": "My HTML-to-Markdown output passed every test I wrote — then flattened in a strict parser",
  "url": "https://urgent.news/2026/10/04/my-html-to-markdown-output-passed-every-test-i-wrote-then-flattened",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-04T03:34:49.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/imapphelp/my-html-to-markdown-output-passed-every-test-i-wrote-then-flattened-in-a-strict-parser-cbo"
  },
  "original_language": "en",
  "account": "A developer discovered an issue with their HTML-to-Markdown converter after running a series of tests. The converter passed all the tests, but the output flattened nested list items into flat lists when fed into a strict parser. The problem stemmed from the converter indents nested list items using two spaces, which is accepted by lenient parsers and most browser previews but not by strict CommonMark parsers.\n\nThe developer's test suite only checked the output against one parser, a lenient one, and this proved insufficient. The fix was simple, but the diagnosis was difficult. To ensure compatibility with different parsers, the developer now runs a small harness in CI that converts HTML to Markdown, re-parses the result with three different Markdown implementations, and compares the resulting trees. If there is any disagreement, the build fails.\n\nThis experience highlighted that Markdown is not a single format but a family of dialects with varying interpretations of indentation, heading style, emphasis, and tables. The developer's converter, available at https://x402.freeq.one/tools/html_to_markdown.html, provides options for heading style, bullet marker, and code fence to accommodate the various parsers that might be encountered. The takeaway is that when converting HTML to Markdown for use in RAG pipelines, it is crucial to test the round trip with the specific parser that will be used at the end of the pipeline, not just the one used for initial testing.",
  "summary": "An agent piped a wiki export through my HTML-to-Markdown converter into a RAG pipeline and reported that hierarchy was disappearing. Every nested task list came out flat: sub-items rendered as siblings of their parents. Sections got chunked together that had nothing to do with each other, and retrieval started returning paragraphs attributed to the wrong topic. Here's the part that stung: the…",
  "key_points": [
    "Converter passed all tests in lenient parser",
    "Nested lists flattened into flat lists in strict parser",
    "Developer added CI harness for multi-parser testing"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}