arrow_backDashboard

PDF to Excel — Extract Tables from PDF

Extract PDF tables to an Excel (.xlsx) spreadsheet locally.

upload_file

Select PDF File

or drag the PDF file here

Pull tables out of a PDF into Excel, on your own machine

Extract the text of a PDF into an .xlsx spreadsheet, one sheet per page, with cells laid out by their position on the page. Nothing is uploaded.

Financial statements, price lists, lab results and exported reports all tend to arrive as PDFs with the numbers locked into a fixed layout. You can see the table; you just cannot sum a column of it. Getting that data into a spreadsheet by hand is the kind of task that eats an afternoon and introduces typos.

This tool reads the text of each page and reassembles it into a grid: items sharing a vertical position become a row, and within that row they are ordered left to right by their horizontal position. The result is an Excel workbook with one worksheet per PDF page.

It works on the text layer, in your browser. A bank statement or a payroll export never reaches a server, which is generally the point when the numbers belong to someone else.

Financial data stays local

Statements, invoices and payroll exports are processed in the tab. No upload, no third-party copy, no retention policy to read.

One sheet per page

Each PDF page becomes its own named worksheet, so a multi-page report stays navigable instead of collapsing into one endless sheet.

Standard .xlsx output

Opens in Excel, Google Sheets, LibreOffice Calc and Numbers. Values arrive as ordinary cells you can sort, filter and total.

How to get a PDF's tables into Excel

1

Select the PDF containing the tables you need.

2

Text is extracted page by page with its position on the page.

3

Items are grouped into rows and ordered into columns.

4

Download the .xlsx and clean up the columns as needed.

Your spreadsheet data never leaves the tab

Bank statements, payroll exports and client ledgers are exactly the documents you should not be uploading to a free web service. Here the extraction runs on your own processor and the PDF is never transmitted.

When people reach for this

Reconciling a bank or card statement

The bank gives you a PDF, your accounting workflow wants rows. Extracting the transactions into a sheet lets you sort, filter and total them instead of reading down the page with a ruler.

Comparing supplier price lists

Three suppliers send catalogues as PDFs in three different layouts. Pulling each into a sheet gets them into a form where a lookup can compare them line by line.

Recovering data from a report with no source file

A published study or an annual report has the figures you need in a table, and nobody is going to send you the underlying spreadsheet. Extraction is faster and less error-prone than retyping.

Handling records you must not upload

Patient lists, salary data and client ledgers cannot be sent to an online converter without creating a disclosure. Processing locally avoids the question entirely.

What actually happens to your file

Each page is parsed with pdf.js, which returns the page's text items along with a transform matrix for each one. Two numbers from that matrix matter here: the vertical position and the horizontal position of the item on the page.

Items are bucketed by rounded vertical position, so everything printed at the same height lands in the same bucket — that bucket becomes a spreadsheet row. Within it, the items are sorted by their horizontal coordinate, which puts them in left-to-right reading order and turns each one into a cell.

The rows are then ordered from the top of the page downwards and handed to SheetJS, which builds a worksheet from that array of arrays. Each page gets its own sheet, named after the page number, and the workbook is written out as .xlsx in the browser.

Note what this does and does not use: position on the page, and nothing else. There is no table-detection model and no ruling-line analysis. Column structure is inferred purely from where things were printed.

Where it falls short, honestly

  • Columns are inferred, not read. Because cells are ordered by horizontal position rather than assigned to detected columns, a row with a blank cell shifts everything after it one place to the left. Expect to tidy up alignment.
  • Nothing outside the table comes through as structure. Headers, footers, page numbers and paragraphs of prose all land in the grid as ordinary rows.
  • Formatting, merged cells and formulas are not produced. You get values as text and numbers, in plain cells.
  • Ruled lines and cell borders are ignored entirely. A table drawn with visible gridlines is treated exactly like columns of text with wide spacing.
  • A scanned PDF yields nothing. Without a text layer there are no coordinates to position anything by. The OCR tool can read those pages, but it returns plain text rather than a positioned layer, so the column structure will not survive.
  • Very wide tables can wrap awkwardly if the source itself splits a table across pages, because each page becomes a separate sheet.

Compared with an online converter

This toolTypical cloud converter
Where your file goesNowhere. It stays in the browser tab.Uploaded to a server you do not control.
Suitable for financial recordsYes — no disclosure occurs.Depends on their terms and jurisdiction.
Page limitsBounded only by your device's memory.Commonly capped on free tiers.
Table detectionPosition-based; alignment may need tidying.Sometimes uses trained models, often better.
Works offlineYes, once the page has loaded.No.
CostFree, no account.Usually metered or subscription.

Frequently Asked Questions

My columns are misaligned. How do I fix that?expand_more
This is the most common result and it comes from how the grid is built: cells are placed in left-to-right order rather than into detected columns, so an empty cell in a row shifts the rest across. Sorting it out is usually a matter of inserting a cell or two in the affected rows. Tables with a value in every cell come through cleanly.
Can it detect tables automatically?expand_more
Not in the sense of recognising a table as an object. It reconstructs a grid from where text was printed on the page. For well-formed tables that amounts to the same thing; for irregular ones it does not.
Why is each page a separate sheet?expand_more
Page boundaries are the only structural division the PDF reliably gives us. Keeping them separate means you can see which page each row came from and merge them yourself if you want a single continuous table.
The output is empty or nearly empty.expand_more
Your PDF is almost certainly a scan, with no text layer to extract. The OCR tool will recognise the words, but it outputs plain text rather than a PDF with positioned text, so it cannot feed this converter — and a table read that way loses its column alignment. For scanned tables, expect to do some manual work.
Do formulas from the original spreadsheet come back?expand_more
No, and they cannot. When a spreadsheet is exported to PDF the formulas are gone — only their computed results were printed. Extraction recovers those results as values.
Is there a row or file size limit?expand_more
No fixed limit. The whole document is processed in the browser tab, so the ceiling is your device's memory rather than a plan tier.
Can I convert a password-protected PDF?expand_more
The text layer of an encrypted PDF cannot be read until it is decrypted. Run the file through the unlock tool with its password first, then extract from the result.

Related guides

Tools that pair with this one