Microsoft Markitdown converts any document to Markdown for LLM input - 3,600 GitHub stars
Microsoft's open-source Markitdown tool converts documents in any format (PDF, DOCX, PPTX, HTML, images) to clean Markdown for LLM consumption. The project gained 3,600+ stars in its first two weeks of April activity.
Microsoft released Markitdown as an open-source Python tool that converts virtually any document format into clean Markdown suitable for LLM input. The project quickly gained over 3,600 GitHub stars, reflecting strong demand for this kind of utility.
The tool handles PDFs, Word documents (.docx), PowerPoint files (.pptx), Excel spreadsheets (.xlsx), HTML pages, images (via OCR), and several other formats. Output is clean Markdown with preserved structure: headings, lists, tables, and code blocks all come through correctly. The conversion maintains the logical hierarchy of the document rather than dumping raw text.
Installation is a single pip command: pip install markitdown. Usage is equally simple. The command-line tool takes a file path and outputs Markdown to stdout. The Python API lets you integrate it into scripts and pipelines with a few lines of code. There is also a batch mode for processing entire directories of documents at once.
For AI agent builders, the value is in preprocessing. Most LLMs work best with text input, but the real world runs on PDFs, Word documents, and slide decks. Markitdown bridges that gap without requiring separate libraries for each format. Before Markitdown, you might have needed PyMuPDF for PDFs, python-docx for Word files, python-pptx for slides, and BeautifulSoup for HTML. Markitdown handles all of these through a single interface.
The tool is particularly useful for RAG (Retrieval-Augmented Generation) pipelines. Instead of building custom parsers for each document type, you run everything through Markitdown first and get consistent Markdown output that chunks well for embedding databases. The structural preservation means headings and sections are maintained, which improves retrieval quality.
Microsoft released it under the MIT license, so there are no usage restrictions. The project accepts contributions and has an active issue tracker with responsive maintainers. Quality of extraction varies by document type: simple text documents convert perfectly, complex layouts with images and charts may lose some visual formatting, and scanned PDFs depend on OCR quality.
For anyone who has written custom document parsing code for their AI workflows, Markitdown is worth evaluating as a replacement. It handles the common cases well, and the consistent output format simplifies everything downstream in your pipeline.