MarkItDown in Python: Convert PDF, DOCX, XLSX and More to Markdown (and the One Case It Fails Silently)
MarkItDown turns PDF, Word, Excel, PowerPoint, HTML, EPUB and more into Markdown with one Python call, and it is designed for LLM pipelines rather than pixel-perfect documents. I ran it against 14 fixtures to find its edges: the sharpest one is that scanned PDFs return a single newline with exit code 0 — a silent failure you have to check for. TL;DR pip install 'markitdown[all]' and markitdown…
MarkItDown is a Python utility for converting various file formats, including PDF, DOCX, XLSX, PowerPoint, HTML, EPUB, and more, into Markdown, specifically designed for use with LLM pipelines. It preserves headings, lists, tables, and links while converting text-based files cleanly. However, scanned PDFs return a single newline with an exit code of 0, which is a silent failure that must be checked for.
The tool is MIT-licensed and requires Python 3.10–3.14. Installation can be done using a virtual environment and pip install markitdown[all] to include all formats. The Python API is straightforward, requiring just three lines of code to convert a file or a folder. While MarkItDown does an excellent job with text-based files, tables often need light cleanup, and scanned PDFs produce empty output without raising an error.
The tool was tested on 14 fixtures, with most conversions producing readable Markdown output, but scanned PDFs occasionally resulted in empty files.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.