OCR PDF
Extract text from scanned PDF with OCR.
Select Scanned PDF
Or drag the file here
Read the text out of a scanned PDF, on your own machine
Run optical character recognition over a scanned document and get the recognised text back. English and Spanish; nothing is uploaded.
A scanned PDF looks like a document and behaves like a photograph. You can read it, but your computer cannot: search finds nothing, you cannot select a sentence, and every conversion tool comes back empty. The pages contain pictures of words, not words.
Optical character recognition closes that gap by looking at the shapes in the image and working out which characters they are. This tool runs a recognition engine inside your browser tab, page by page, and gives you the text it found.
Be clear about the output, because it is the thing most often misunderstood: you get the recognised text, as a plain text file you can download or copy. It does not produce a searchable PDF — the original document is not modified and no invisible text layer is added to it. If what you need is the words, this gives you the words.
Recognition happens locally
The engine and its language model run in your browser. Scanned contracts, records and correspondence are never uploaded.
No page limits or quotas
OCR is expensive to run, which is why cloud services meter it. Here it uses your processor, so there is nothing to ration.
Copy or download the text
Take the result straight to the clipboard, or download it as a .txt file with the page boundaries marked.
How to OCR a scanned PDF
Select the scanned PDF.
Choose whether the document is in English or Spanish.
Each page is rendered and recognised in your browser.
Copy the text or download it as a .txt file.
Your scans are never uploaded
OCR is one of the operations most often outsourced to a third-party AI service, which means handing over the whole document. Here the recognition engine runs inside your browser, so a scanned medical record or contract stays exactly where it was.
When people reach for this
Quoting from a scanned contract
You need three clauses out of a signed agreement that exists only as a scan. Recognising the text is considerably faster than retyping, and less prone to introducing an error into a quotation that matters.
Making an archive searchable
Boxes of scanned correspondence are unsearchable by definition. Extracting the text gives you something you can grep, index or paste into a notes system, even though the PDFs themselves stay as they are.
Getting data off a photographed document
A receipt, an invoice or a form captured with a phone. Recognition turns it into figures you can put in a spreadsheet rather than transcribe by hand.
Reading a document that resists copying
Some PDFs are scans precisely so that text cannot be lifted from them. If you are entitled to the content, recognition is a legitimate way to work with it.
What actually happens to your file
Each page of your PDF is first rendered to an offscreen canvas by pdf.js at twice its natural size. That 2× scale is deliberate: recognition accuracy depends heavily on how many pixels each character occupies, and rendering at native size leaves small text too coarse to identify reliably.
The rendered page is then handed to Tesseract — a mature open-source OCR engine, compiled to WebAssembly so it runs in your tab rather than on a server. The first time you use it, the language model for your chosen language downloads and is then cached.
The engine analyses the image, segments it into lines and characters, and matches the shapes against its trained model. The recognised text for each page is collected and separated by a page marker, so you can see which page each passage came from.
The result is assembled as plain text. Nothing is written back into your PDF: the original file is untouched, and what you download is a .txt file containing what the engine read.
Where it falls short, honestly
- It does not produce a searchable PDF. The output is text, not a modified document with an invisible layer behind the image. If you need a genuinely searchable PDF, this tool does not make one.
- Only English and Spanish are supported. Documents in other languages will be recognised as though they were in whichever of the two you picked, which produces nonsense.
- Accuracy depends almost entirely on the scan. Clean, straight pages at 300 DPI or better read very well. Low-resolution, skewed, shadowed or creased pages read poorly, and no amount of processing fixes a bad original.
- Handwriting is not recognised. The engine is trained on printed type.
- Layout is not preserved. Columns, tables and figure captions come back as a stream of lines, so a multi-column page will interleave.
- It is slow and memory-hungry compared with every other tool here, because each page has to be rendered and then analysed. A long document takes minutes, not seconds — that cost is real, and it is the same cost cloud services charge you for.
Compared with a cloud OCR service
| This tool | Typical cloud OCR | |
|---|---|---|
| Where your file goes | Nowhere. It stays in the browser tab. | Uploaded, often to a third-party AI service. |
| Suitable for confidential scans | Yes — no disclosure occurs. | Depends on their terms and subprocessors. |
| Page quotas | None. | Almost always metered. |
| Languages | English and Spanish. | Often dozens. |
| Output | Plain text. | Sometimes a searchable PDF. |
| Speed | Slower — it uses your processor. | Faster, on their hardware. |