VSThiran

OCR & Documents

Extract text from an image or PDF

Free to useNo sign-up requiredNo watermarkRuns in your browser

An image or a scanned PDF has no real text in it - to a computer, it is just a picture of text, which is why you cannot select, search or copy from it the way you can a typed document. OCR (optical character recognition) reads the shapes in the picture and turns them back into real text.

This runs the recognition itself in your browser: the file is read on your own device, nothing is uploaded to a VSThiran server, and the extracted text is yours to copy or download as soon as it is ready.

How this tool works

  1. Upload your image or PDF

    JPG, PNG or PDF, read directly in your browser.

  2. Choose what you need

    Image to Text or PDF to Text extracts the words. Scanned PDF to Searchable PDF (PDFs only) keeps the original look and adds a text layer you can select and search.

  3. Wait for recognition

    A progress indicator shows what is happening - initializing, recognizing text, and which page of a multi-page PDF is currently being read.

  4. Copy or download the result

    Copy the text directly, download it as a .txt file, or download the searchable PDF.

How it works

Recognition runs on Tesseract, an open-source OCR engine compiled to WebAssembly so it can run directly in your browser rather than on a server. VSThiran did not write the OCR engine itself - it uses this mature, widely-used one, which is also what tools like OCRmyPDF are built on.

For a PDF, each page is rendered as an image first (the same rendering VSThiran's PDF workspace already uses for previews), then OCR runs on that image, one page at a time, so a multi-page document does not need to load entirely into memory at once.

For "Scanned PDF to Searchable PDF", the OCR engine itself produces each page as a small PDF containing the original page image with an invisible, precisely positioned text layer on top - the same technique other serious OCR tools use. VSThiran only combines those pages into one file; it does not fabricate the text layer itself.

The OCR engine and its language data are hosted on vsthiran.com, not fetched from a third-party CDN - so no request for OCR leaves VSThiran's own domain.

What works well, and what does not yet

OCR is genuinely good at typed or printed text, even from a fairly rough scan or a heavily compressed photo - and reliably tells you when there is nothing to read rather than inventing text from a blank page.

  • Clear typed or printed text: reliable, including scans and screenshots.
  • Rotated or sideways pages: not currently supported - rotate the page first with Rotate PDF, then run OCR on the corrected file.
  • Handwriting: not reliable with this engine, which is built for printed text.
  • A PDF that already has real, selectable text: use PDF to Word or PDF to Excel instead - OCR is for when there is no text to extract yet.

Extract, or make searchable - which do you need?

Text extraction gets you the words themselves, to copy, edit or paste elsewhere. Searchable PDF keeps the document looking exactly as it did - the same scan, same layout - but adds an invisible layer underneath so you can select, search and copy from it in any PDF viewer, the same as a typed document.

Worked examples

A photographed page of notes

Upload the photo, choose Image to Text, and get the words back as real, copyable text - even if the photo has some glare or is slightly blurry.

A scanned contract you need to search

Upload the scanned PDF, choose Scanned PDF to Searchable PDF, and download a PDF that looks identical but now supports Ctrl+F and text selection.

Frequently asked questions

Is my file uploaded to a server?

No. Recognition runs entirely in your browser, on your own device, using a self-hosted copy of the OCR engine rather than a third-party service.

What file types are supported?

JPG, PNG and PDF. A PDF can be scanned (image-only) or a mix of scanned and typed pages.

Can it read handwriting?

Not reliably. This engine is built and trained for printed and typed text, not handwriting.

Can it read a page that is upside down or sideways?

Not currently - rotate the page to the correct orientation first (Rotate PDF) and then run OCR on the corrected file.

Does "Scanned PDF to Searchable PDF" change how the document looks?

No - the original page images are kept exactly as they were. Only an invisible text layer is added underneath, which does not affect the visible appearance at all.

What languages are supported?

English, to start. Support for additional languages is planned once each one has been tested for accuracy.

Is there a file size or page limit?

Yes - OCR is memory-intensive, and it runs on your own device rather than a server, so very large files and very long PDFs are refused with a clear explanation rather than being allowed to freeze the tab.

Guides for this tool