VSThiran

PDF & Files

Free PDF Inspector

Free to useNo sign-up requiredNo watermarkRuns in your browser

This free inspector tells you what is actually inside a PDF: how many pages, what size each one is, which are rotated, how much text can be extracted, how many images are drawn, and whether the document has form fields or links.

The most useful thing it does is tell you whether a document is a scan. A PDF of photographed pages looks identical to a text PDF from the outside, and it is the reason PDF to Word sometimes produces an empty file for no apparent reason. If no text comes back, this says so and explains what that means.

How this tool works

  1. Open your PDF

    Drag it in or browse. It is read entirely on your device.

  2. Read the summary

    Page count, file size, extractable text, images, links and form fields, all at a glance.

  3. Check the page sizes

    Distinct sizes are grouped, so a document that mixes A4 and Letter is obvious immediately.

  4. Open the per-page detail

    Every page with its exact dimensions in points, the paper size it matches, its orientation and its rotation.

How it works

Two libraries answer different halves of the question, because neither can answer both. pdf-lib reads the structure — page boxes, rotation, the document dictionary — and pdf.js reads the content, extracting the text and counting the image-painting operations. Both run in your browser.

Page sizes are reported in points, the PDF's own unit, at 72 to the inch. A size is named as A4 or Letter only when it is within three points of the standard, so a nearly-A4 page is reported by its measurements rather than mislabelled.

"Extractable text" counts the characters pdf.js can pull out. A page with a picture of text contains none, which is what makes scan detection possible.

The image count is how many times the pages paint an image, not how many distinct images the file holds — a logo on every page of a twenty-page document counts twenty times.

What a scan means for the other tools

  • PDF to Word and PDF to Excel read text. From a scan they produce an empty or nearly empty result, because there is no text to read.
  • Compression works well on scans, because they are images and images are what the compressor re-encodes.
  • Searching inside the document will not find anything, in this or any other reader.
  • Getting text out of a scan needs OCR — optical character recognition. VSThiran's OCR workspace does this directly in your browser, once you know that is what the document needs.

Frequently asked questions

Is my PDF uploaded anywhere?

No. The file is opened and changed by this page on your own device, using your browser. It is never sent to VSThiran and no copy is stored on any server. It is held only in the tab you have open, which is also why reloading the page clears it.

Why does it say my PDF has no text when I can see the words?

Because they are a picture of words. A scanner or a photograph produces an image of the page, and the letters you are reading are pixels rather than characters. Nothing can select, search or copy them without OCR first.

Why do the page sizes not match what my printer says?

PDF pages are measured in points, at 72 to the inch — A4 is 595 × 842 and Letter is 612 × 792. A printer usually reports millimetres or inches, and may also be describing the paper rather than the page.

Does this change my document?

No. The inspector only reads. There is no download button here because there is nothing new to download — the document you opened is untouched.