arrow_backDashboard

PDF to Text

Convert a PDF to a text file quickly and easily.

upload_file

Select PDF

or drag the file here

Extract the text from a PDF as plain text

Pull the words out of a PDF into clean text you can copy or download. Instant, exact, and processed entirely in your browser.

Copying text out of a PDF viewer is famously irritating. You drag across a paragraph and get line breaks in the middle of sentences, hyphens where words were split across lines, and a header repeated every page. What you wanted was the prose; what you got was the layout.

This tool reads the text the PDF actually stores and gives it to you as a single block, page by page. Because it reads the document's own text objects rather than scraping a rendered image, it is exact — every character is the character the document contains, not a guess.

That exactness is also the boundary: it only works when the PDF genuinely contains text. A scan does not, and this tool will come back empty on one. Try it first anyway, because it takes a second and tells you immediately which kind of document you are holding.

Nothing is uploaded

Extraction happens in your browser tab, so the document you are pulling text out of never reaches a server.

Exact, not recognised

The text comes from the document's own objects, so there is no recognition step and no chance of a misread character.

Copy or download

Take the result to your clipboard, or save it as a .txt file ready for any editor, script or notes app.

How to extract text from a PDF

1

Select the PDF you want the text from.

2

The text objects are read page by page in your browser.

3

Review the extracted text on screen.

4

Copy it or download it as a .txt file.

The document stays in your browser

Text extraction is exactly the operation people reach for with documents they need the contents of but should not be circulating. Here it runs on your own machine and the PDF is never transmitted.

When people reach for this

Quoting accurately from a long document

A regulatory filing or a judgment where the wording has to be exact. Extracting the text gives you something you can search and copy from cleanly, without the stray line breaks a viewer's selection introduces.

Feeding a document into another tool

Translation software, a summarisation workflow, a word counter or a script all want plain text. This is the shortest path from a PDF to something they can read.

Checking what a PDF really contains

Running extraction is the fastest way to find out whether a document has a text layer at all. Empty output means you are holding a scan, which changes what every other tool can do with it.

Rescuing content from a document you cannot edit

The source file is long gone and you need the wording, not the layout. Extraction is faster than any conversion and gives you exactly the words.

What actually happens to your file

The PDF is opened with pdf.js and, for each page, its text content is requested. What comes back is not a paragraph of prose but a list of text items — the individual runs of characters the document draws, each with its own position on the page.

Those items are joined in the order the document stores them and collected page by page, so what you get is the text as the file itself holds it.

This is a fundamentally different operation from OCR. Nothing is being recognised or guessed: the characters are read directly out of the document's own structures. That is why it is instantaneous even on a long file, and why it is exact rather than approximate.

The whole extraction runs in the tab, and the result is offered to you as text you can copy or download. Your PDF is neither modified nor transmitted.

Where it falls short, honestly

  • A scanned PDF produces nothing. If the pages are images of paper, there are no text objects to read. That is not a failure — it is the tool correctly telling you the document has no text layer. Use the OCR tool for those.
  • Layout is not preserved. Columns, tables, headers and footers all come through as part of the same stream, so a two-column page will read across rather than down.
  • Reading order follows the document, not the eye. Text is emitted in the order the file stores it, which is usually sensible but can be surprising in complex layouts.
  • Formatting is lost entirely. Bold, italics, headings and font sizes are not represented in plain text.
  • Images, charts and any words that are part of a picture are not extracted. Only real text objects are read.
  • Encrypted documents must be unlocked first, since their content cannot be read while encrypted.

Compared with copying from a PDF reader

This toolSelecting and copying by hand
Whole document at onceYes, every page.Page by page, or a fight with scrolling.
Stray line breaksFar fewer — items are joined per line.Common, and tedious to clean up.
Works on a long fileYes, in one step.Impractical past a few pages.
Output as a fileYes, downloadable .txt.Clipboard only.
Where your file goesNowhere — it stays in the tab.Nowhere.
Works on scansNo — use OCR.No.

Frequently Asked Questions

My extracted text is empty. What does that mean?expand_more
It means your PDF has no text layer — the pages are images, so it is a scan. That is useful information rather than an error: it tells you immediately that conversion tools will also come back empty. Use the OCR tool, which recognises characters in the image.
How is this different from OCR?expand_more
This reads the characters the document actually stores, so it is exact and instant. OCR looks at a picture of a page and works out what the characters probably are, which is slower and can be wrong. Always try extraction first; only use OCR when it comes back empty.
Why is my two-column document jumbled?expand_more
Text is emitted in the order the file stores it, with no understanding of columns. In a two-column layout that often means reading across both columns rather than down each one. There is no reliable way to infer column structure from the text stream alone.
Are tables preserved?expand_more
No. A table becomes lines of text, since the structure lives in the positioning rather than the text itself. If you need tabular data, the PDF to Excel tool groups the same text by column position.
Does it keep bold and headings?expand_more
No. The output is plain text, so all styling is dropped. If you need formatting, the PDF to Word tool at least gives you paragraphs in an editable document.
Is there a page limit?expand_more
No. Extraction is cheap — it reads structures rather than rendering pages — so even very long documents are handled in seconds.
Can I extract text from a protected PDF?expand_more
Not while it is encrypted. Use the unlock tool with the password first, then extract from the decrypted copy.

Related guides

Tools that pair with this one