VSThiran

How to extract text from a scanned PDF

Quick answer

Upload the scanned PDF to the OCR tool and choose PDF to Text. Each page is read in turn, and any page that already has real text is used directly rather than re-read - only pages that are genuinely scans get OCR.

  1. 1Upload your scanned PDF.
  2. 2Choose PDF to Text.
  3. 3Watch the per-page progress.
  4. 4Copy or download the extracted text.

Free, no sign-up, and your file is read on your own device rather than uploaded.

A scanned PDF is a picture of a page saved inside a PDF file - not real text, no matter how much it looks like a normal document on screen. This is exactly why PDF to Word or PDF to Excel comes back empty on a scan: there is no text in the file for them to recover, only an image.

OCR is the actual fix: each page is rendered and read for its text, one page at a time, so even a long scanned document can be processed without needing to load the whole thing into memory at once.

Step by step

  1. Upload your scanned PDF

    Read directly in your browser.

    OCR Text Extractor
  2. Choose PDF to Text

    Extracts the text from every page.

    OCR Text Extractor
  3. Watch the progress

    "Processing page 4 of 12" and similar, so a long document does not look frozen.

    OCR Text Extractor
  4. Copy or download the result

    The combined text from every page, ready to copy or save as a .txt file.

    OCR Text Extractor

Tips

  • A mixed document - some pages typed, some scanned - is handled correctly either way: pages that already have real text are used directly and page that are genuinely scans get OCR, so nothing is OCR'd unnecessarily.
  • Check with the PDF Inspector first if you are not sure whether a document is a scan - it tells you directly whether a page has selectable text.
  • A rotated or sideways page will not OCR correctly - straighten it first with Rotate PDF, then run OCR on the corrected file.

Common problems

PDF to Word or PDF to Excel came back empty before you found this page.

That confirms the document is a scan - those converters can only recover text that already exists in the file, and a scan has none. OCR reads the image instead.

A long scanned PDF is taking a while.

Expected - each page is rendered and read individually. The page-by-page progress indicator shows exactly where it is, and browser OCR on the largest documents genuinely does take longer than a quick glance suggests.

One page in the middle of the document looks wrong in the result.

Check whether that specific page is rotated, unusually low-quality, or a different kind of content (a photo rather than a scanned page) - the rest of the document is unaffected either way.

Frequently asked questions

Why did PDF to Word not work on my scanned PDF?
A scan is a picture of a page, not real text - PDF to Word can only recover text that already exists in the file. OCR reads the image and produces the text that was never there to extract.
Does this work on a PDF with some scanned and some typed pages?
Yes - each page is checked, and pages that already have real text are used directly rather than being OCR'd unnecessarily.
Is there a page limit?
Yes - OCR is memory-intensive and runs on your own device, so very long documents are refused with a clear explanation rather than risking a frozen tab.
Is my PDF uploaded to a server?
No. Every page is rendered and read entirely in your browser.

Doing several things to this file?

"Editing a PDF" is usually one of five different jobs. Open the file once and you can do all of them in sequence - reorder the pages, mark it, number it, sign it and shrink it - then download once at the end.

How to edit a PDF without Acrobat

Related OCR tools

Related articles

Ready to do it?

Free, no sign-up, and nothing is uploaded to a server.