PDF to Text
Convert a PDF to a text file quickly and easily.
Select PDF
or drag the file here
Extract the text from a PDF as plain text
Pull the words out of a PDF into clean text you can copy or download. Instant, exact, and processed entirely in your browser.
Copying text out of a PDF viewer is famously irritating. You drag across a paragraph and get line breaks in the middle of sentences, hyphens where words were split across lines, and a header repeated every page. What you wanted was the prose; what you got was the layout.
This tool reads the text the PDF actually stores and gives it to you as a single block, page by page. Because it reads the document's own text objects rather than scraping a rendered image, it is exact — every character is the character the document contains, not a guess.
That exactness is also the boundary: it only works when the PDF genuinely contains text. A scan does not, and this tool will come back empty on one. Try it first anyway, because it takes a second and tells you immediately which kind of document you are holding.
Nothing is uploaded
Extraction happens in your browser tab, so the document you are pulling text out of never reaches a server.
Exact, not recognised
The text comes from the document's own objects, so there is no recognition step and no chance of a misread character.
Copy or download
Take the result to your clipboard, or save it as a .txt file ready for any editor, script or notes app.
How to extract text from a PDF
Select the PDF you want the text from.
The text objects are read page by page in your browser.
Review the extracted text on screen.
Copy it or download it as a .txt file.
The document stays in your browser
Text extraction is exactly the operation people reach for with documents they need the contents of but should not be circulating. Here it runs on your own machine and the PDF is never transmitted.
When people reach for this
Quoting accurately from a long document
A regulatory filing or a judgment where the wording has to be exact. Extracting the text gives you something you can search and copy from cleanly, without the stray line breaks a viewer's selection introduces.
Feeding a document into another tool
Translation software, a summarisation workflow, a word counter or a script all want plain text. This is the shortest path from a PDF to something they can read.
Checking what a PDF really contains
Running extraction is the fastest way to find out whether a document has a text layer at all. Empty output means you are holding a scan, which changes what every other tool can do with it.
Rescuing content from a document you cannot edit
The source file is long gone and you need the wording, not the layout. Extraction is faster than any conversion and gives you exactly the words.
What actually happens to your file
The PDF is opened with pdf.js and, for each page, its text content is requested. What comes back is not a paragraph of prose but a list of text items — the individual runs of characters the document draws, each with its own position on the page.
Those items are joined in the order the document stores them and collected page by page, so what you get is the text as the file itself holds it.
This is a fundamentally different operation from OCR. Nothing is being recognised or guessed: the characters are read directly out of the document's own structures. That is why it is instantaneous even on a long file, and why it is exact rather than approximate.
The whole extraction runs in the tab, and the result is offered to you as text you can copy or download. Your PDF is neither modified nor transmitted.
Where it falls short, honestly
- A scanned PDF produces nothing. If the pages are images of paper, there are no text objects to read. That is not a failure — it is the tool correctly telling you the document has no text layer. Use the OCR tool for those.
- Layout is not preserved. Columns, tables, headers and footers all come through as part of the same stream, so a two-column page will read across rather than down.
- Reading order follows the document, not the eye. Text is emitted in the order the file stores it, which is usually sensible but can be surprising in complex layouts.
- Formatting is lost entirely. Bold, italics, headings and font sizes are not represented in plain text.
- Images, charts and any words that are part of a picture are not extracted. Only real text objects are read.
- Encrypted documents must be unlocked first, since their content cannot be read while encrypted.
Compared with copying from a PDF reader
| This tool | Selecting and copying by hand | |
|---|---|---|
| Whole document at once | Yes, every page. | Page by page, or a fight with scrolling. |
| Stray line breaks | Far fewer — items are joined per line. | Common, and tedious to clean up. |
| Works on a long file | Yes, in one step. | Impractical past a few pages. |
| Output as a file | Yes, downloadable .txt. | Clipboard only. |
| Where your file goes | Nowhere — it stays in the tab. | Nowhere. |
| Works on scans | No — use OCR. | No. |