Convert PDF to Markdown, free
Why convert a PDF to Markdown at all
A PDF is a description of a printed page: where the ink goes, not what the document means. That is fine for printing and awkward for everything else. Converting PDF to Markdown turns it back into structure — headings that are headings, lists that are lists, tables with rows — which is the form the next thing along actually wants.
- Feeding a model: Markdown costs far fewer tokens than raw PDF text with its repeated headers, and the headings tell the model what is a section and what is an aside. It is the format to use when you want to give a report to ChatGPT, Claude or your own pipeline.
- Search and RAG: chunking works on structure. A Markdown heading is a natural chunk boundary; a page break is not.
- Version control: a Markdown file diffs. Two PDFs of the same report tell you only that the bytes differ.
- Re-use: Markdown pastes into a wiki, a ticket, a doc site or a static site without a second conversion.
What comes through, and what cannot
| In the PDF | In the Markdown |
|---|---|
| Headings and body text | Kept, in reading order |
| Bullet and numbered lists | Kept as lists |
| Tables with a text layer | Kept as Markdown tables |
| Page numbers, running headers and footers | Removed |
| Images, charts and diagrams | Not extracted — the text around them is |
| A scanned page with no text layer | Nothing, and the tool says so rather than returning an empty file |
If you get “no text layer”
It means the PDF is a picture of a document rather than a document. Government notices, older filings and anything that came off a scanner are usually like this. Open it and try to select a sentence: if the cursor will not take the text, no converter can read it without OCR first. Run it through an OCR tool and convert the result.
A whole library of documents, on a schedule
One PDF is a tool. A policy library, a set of regulatory notices or a manufacturer’s manuals is a crawl: MeshArc follows every PDF, Word file, spreadsheet and slide deck a site links to — including the ones behind a viewer page that never exposes a .pdf link — converts each one, and tells you when a document is replaced or withdrawn.
Frequently asked questions
How do I convert a PDF to Markdown?
Paste a link to the PDF above and the converter returns Markdown in a second or two. It reads the file’s text layer, keeps headings, lists and tables, and drops the page furniture that repeats on every page.
Does the converter keep tables?
Tables in a PDF’s text layer come through as Markdown tables. A table that is an image of a table does not, because there is no text in it to read.
Can it convert a scanned PDF?
No. A scan has no text layer, and this does not do OCR. The tool says so rather than handing back an empty file that looks like a conversion.
What is the best way to give a PDF to an AI model?
Clean Markdown of the text, with the headings kept. It costs a fraction of the tokens that raw PDF extraction does, and the structure tells the model what belongs to what.
Is there a limit on size or pages?
The free tool reads one document per run and returns up to about 60,000 characters. A crawl on an account has no such ceiling and follows every document a site links to.
Convert one, or watch a library
The tool is free. The crawl that follows every document on a site is one API call.