Convert a scanned PDF to Markdown

A scanned page is a picture of text, so an ordinary converter finds nothing to convert. This tool recognises the words in the image instead. Unlike everything else on this site, it works by sending the pages to our server — you are told exactly what will be sent and have to confirm it before anything leaves your device.

5.0 out of 5 from 23 ratings

Drop your PDF or Word file here

or choose a file from your computer

PDF or Word (.docx) · up to 50 MB

How do I convert a scanned PDF to Markdown?

Drop the scan onto this page. It is checked in your browser first, and if it has no text layer you are offered recognition. Confirm, and the pages are read on our server and returned as Markdown.

Does my scanned file get uploaded?

Yes, and only this tool does that. The pages go to our own server, are held in memory while they are read, and are destroyed when the job ends. Nothing is written to disk and no other company receives them.

How to convert a scanned PDF to Markdown

  1. 1

    Add your scanned PDF

    Drag the file onto the box above, or pick it from your device. It is inspected in your browser first, so a PDF that already has a text layer is converted here without being sent anywhere.

  2. 2

    Choose the language

    Pick the language the document is written in. Recognition accuracy depends heavily on this, and the wrong choice is the most common cause of poor output.

  3. 3

    Confirm the upload

    You are shown what will be sent, where it goes, and how long it lives. Nothing is transmitted until you press the button, and declining leaves you exactly where you were.

  4. 4

    Check the result

    The Markdown appears in the editor with a note about recognition confidence. Read names, numbers and dates against the original before you rely on them.

What a scanned PDF actually is

PDFs come in two kinds that look identical on screen. One stores the characters, so text can be selected, searched and copied. The other stores a photograph of a page — from a flatbed scanner, a phone camera or a fax — and contains no characters at all.

Our ordinary converter reads the first kind directly in your browser and never sends the file anywhere. Faced with the second kind it correctly finds nothing, which is why scans need a different tool rather than a better version of the same one.

You can tell which you have without any software: open the PDF and try to select a line of text with your cursor. If nothing highlights, it is a scan.

Why this page uploads and the others do not

Optical character recognition needs to rasterise every page at high resolution and run a recognition engine over the image. Doing that in a browser is possible but slow, heavy on memory, and noticeably less accurate — on a mid-range phone, which is where most photographed documents come from, it is not a realistic option.

So this one tool sends the pages to a server. We built it to keep as much of the original promise as possible: recognition runs on our own hardware with our own engine, so no third party ever receives your document, and the service has no database, no object storage and no writable disk to keep it on. The pages exist in memory for the length of the job and then they are gone.

Consent is asked for each file, every time, and is never remembered. If you decline, nothing is sent and the rest of the site behaves exactly as it always does.

What recognition does well, and what it does not

A clean 300 DPI scan of printed text recognises very accurately. A photograph taken at an angle, in poor light, or of a faded fax is materially worse — recognition is a statistical estimate, not a transcription.

Two failure modes are worth knowing about because you cannot see them in the output. Characters are misread, most often in names, reference codes and long numbers where there is no surrounding word to correct against. And words the engine cannot read at all are dropped, so text can be missing rather than wrong. The result tells you the confidence score and names any pages that scored poorly.

  • Printed text in the supported languages — reliable on a clean scan
  • Headings, paragraphs and lists — rebuilt from the position and size of the recognised words
  • Tables — often recovered, but check the column alignment
  • Handwriting — not supported
  • Mathematical formulae — not supported
  • Images and diagrams — not extracted

Getting a better result

Scan at 300 DPI rather than 150 if you have the choice; it is the single biggest difference between a good result and a mediocre one. Photograph pages flat and evenly lit rather than at an angle. And set the language correctly — an English engine reading a German document will produce confident nonsense.

Frequently asked questions

Is my scanned document stored on your server?
No. It is held in memory while it is being read and destroyed when the job finishes or fails. The recognition service has no database and no file storage, so there is nowhere for it to persist.
Do you send my document to Google or another OCR service?
No. Recognition runs on our own server using our own engine. Using this tool does not add any company to the list of processors in our Privacy Policy.
Which languages can it recognise?
English, German, Spanish, French, Portuguese, Hindi, Japanese, Korean, Simplified Chinese and Arabic. Choose the document's language before confirming, as accuracy depends on it.
Can it read handwriting?
No. The engine is trained on printed text. Handwritten notes, signatures and annotations will not be recognised, and we would rather say so than let you find out from the output.
How many pages can I recognise for free?
There is a daily page allowance and a per-document limit, both shown before you confirm. They exist because recognition costs real processing time, unlike the in-browser converter.
Can I convert a scan without uploading it at all?
Not on the web. If your PDF does have a text layer the ordinary converter handles it in your browser; a true scan needs recognition on the server.

Related converters

Further reading