arrow_backDashboard

OCR PDF

Extract text from scanned PDF with OCR.

document_scanner

Select Scanned PDF

Or drag the file here

Read the text out of a scanned PDF, on your own machine

Run optical character recognition over a scanned document and get the recognised text back. English and Spanish; nothing is uploaded.

A scanned PDF looks like a document and behaves like a photograph. You can read it, but your computer cannot: search finds nothing, you cannot select a sentence, and every conversion tool comes back empty. The pages contain pictures of words, not words.

Optical character recognition closes that gap by looking at the shapes in the image and working out which characters they are. This tool runs a recognition engine inside your browser tab, page by page, and gives you the text it found.

Be clear about the output, because it is the thing most often misunderstood: you get the recognised text, as a plain text file you can download or copy. It does not produce a searchable PDF — the original document is not modified and no invisible text layer is added to it. If what you need is the words, this gives you the words.

Recognition happens locally

The engine and its language model run in your browser. Scanned contracts, records and correspondence are never uploaded.

No page limits or quotas

OCR is expensive to run, which is why cloud services meter it. Here it uses your processor, so there is nothing to ration.

Copy or download the text

Take the result straight to the clipboard, or download it as a .txt file with the page boundaries marked.

How to OCR a scanned PDF

1

Select the scanned PDF.

2

Choose whether the document is in English or Spanish.

3

Each page is rendered and recognised in your browser.

4

Copy the text or download it as a .txt file.

Your scans are never uploaded

OCR is one of the operations most often outsourced to a third-party AI service, which means handing over the whole document. Here the recognition engine runs inside your browser, so a scanned medical record or contract stays exactly where it was.

When people reach for this

Quoting from a scanned contract

You need three clauses out of a signed agreement that exists only as a scan. Recognising the text is considerably faster than retyping, and less prone to introducing an error into a quotation that matters.

Making an archive searchable

Boxes of scanned correspondence are unsearchable by definition. Extracting the text gives you something you can grep, index or paste into a notes system, even though the PDFs themselves stay as they are.

Getting data off a photographed document

A receipt, an invoice or a form captured with a phone. Recognition turns it into figures you can put in a spreadsheet rather than transcribe by hand.

Reading a document that resists copying

Some PDFs are scans precisely so that text cannot be lifted from them. If you are entitled to the content, recognition is a legitimate way to work with it.

What actually happens to your file

Each page of your PDF is first rendered to an offscreen canvas by pdf.js at twice its natural size. That 2× scale is deliberate: recognition accuracy depends heavily on how many pixels each character occupies, and rendering at native size leaves small text too coarse to identify reliably.

The rendered page is then handed to Tesseract — a mature open-source OCR engine, compiled to WebAssembly so it runs in your tab rather than on a server. The first time you use it, the language model for your chosen language downloads and is then cached.

The engine analyses the image, segments it into lines and characters, and matches the shapes against its trained model. The recognised text for each page is collected and separated by a page marker, so you can see which page each passage came from.

The result is assembled as plain text. Nothing is written back into your PDF: the original file is untouched, and what you download is a .txt file containing what the engine read.

Where it falls short, honestly

  • It does not produce a searchable PDF. The output is text, not a modified document with an invisible layer behind the image. If you need a genuinely searchable PDF, this tool does not make one.
  • Only English and Spanish are supported. Documents in other languages will be recognised as though they were in whichever of the two you picked, which produces nonsense.
  • Accuracy depends almost entirely on the scan. Clean, straight pages at 300 DPI or better read very well. Low-resolution, skewed, shadowed or creased pages read poorly, and no amount of processing fixes a bad original.
  • Handwriting is not recognised. The engine is trained on printed type.
  • Layout is not preserved. Columns, tables and figure captions come back as a stream of lines, so a multi-column page will interleave.
  • It is slow and memory-hungry compared with every other tool here, because each page has to be rendered and then analysed. A long document takes minutes, not seconds — that cost is real, and it is the same cost cloud services charge you for.

Compared with a cloud OCR service

This toolTypical cloud OCR
Where your file goesNowhere. It stays in the browser tab.Uploaded, often to a third-party AI service.
Suitable for confidential scansYes — no disclosure occurs.Depends on their terms and subprocessors.
Page quotasNone.Almost always metered.
LanguagesEnglish and Spanish.Often dozens.
OutputPlain text.Sometimes a searchable PDF.
SpeedSlower — it uses your processor.Faster, on their hardware.

Frequently Asked Questions

Do I get a searchable PDF back?expand_more
No, and this is the most important thing to know before you start. The output is the recognised text as a .txt file. Your PDF is not modified and no invisible text layer is added to it. If a searchable PDF is what you need, this tool will not produce one.
Which languages does it recognise?expand_more
English and Spanish. You pick one before running it, and the corresponding language model downloads and is cached. A document in another language will be misread, because the engine will try to fit it to the model you chose.
Why is the recognition inaccurate on my document?expand_more
Almost always the scan itself. Recognition needs enough pixels per character and reasonably straight, evenly lit pages. A 150 DPI scan, a photograph taken at an angle, a creased or shadowed page — all of these degrade the result sharply. Rescanning at 300 DPI usually helps more than anything else.
Can it read handwriting?expand_more
No. The engine is trained on printed type; handwritten text will produce garbage.
Why is it so slow?expand_more
Every page has to be rendered to an image and then analysed character by character, which is genuinely expensive work. It runs on your processor rather than a server farm, so a long document takes minutes. That is the honest cost of not uploading it.
Are my tables and columns preserved?expand_more
No. The text comes back as a stream of lines with page markers, so a multi-column layout will interleave and a table loses its structure. Some manual tidying is normal.
Does my document leave the browser?expand_more
No. Both the rendering and the recognition happen in the tab. The language model is downloaded to your browser, not your file uploaded to ours.

Related guides

Tools that pair with this one