PDF to Excel — Extract Tables from PDF
Extract PDF tables to an Excel (.xlsx) spreadsheet locally.
Select PDF File
or drag the PDF file here
Pull tables out of a PDF into Excel, on your own machine
Extract the text of a PDF into an .xlsx spreadsheet, one sheet per page, with cells laid out by their position on the page. Nothing is uploaded.
Financial statements, price lists, lab results and exported reports all tend to arrive as PDFs with the numbers locked into a fixed layout. You can see the table; you just cannot sum a column of it. Getting that data into a spreadsheet by hand is the kind of task that eats an afternoon and introduces typos.
This tool reads the text of each page and reassembles it into a grid: items sharing a vertical position become a row, and within that row they are ordered left to right by their horizontal position. The result is an Excel workbook with one worksheet per PDF page.
It works on the text layer, in your browser. A bank statement or a payroll export never reaches a server, which is generally the point when the numbers belong to someone else.
Financial data stays local
Statements, invoices and payroll exports are processed in the tab. No upload, no third-party copy, no retention policy to read.
One sheet per page
Each PDF page becomes its own named worksheet, so a multi-page report stays navigable instead of collapsing into one endless sheet.
Standard .xlsx output
Opens in Excel, Google Sheets, LibreOffice Calc and Numbers. Values arrive as ordinary cells you can sort, filter and total.
How to get a PDF's tables into Excel
Select the PDF containing the tables you need.
Text is extracted page by page with its position on the page.
Items are grouped into rows and ordered into columns.
Download the .xlsx and clean up the columns as needed.
Your spreadsheet data never leaves the tab
Bank statements, payroll exports and client ledgers are exactly the documents you should not be uploading to a free web service. Here the extraction runs on your own processor and the PDF is never transmitted.
When people reach for this
Reconciling a bank or card statement
The bank gives you a PDF, your accounting workflow wants rows. Extracting the transactions into a sheet lets you sort, filter and total them instead of reading down the page with a ruler.
Comparing supplier price lists
Three suppliers send catalogues as PDFs in three different layouts. Pulling each into a sheet gets them into a form where a lookup can compare them line by line.
Recovering data from a report with no source file
A published study or an annual report has the figures you need in a table, and nobody is going to send you the underlying spreadsheet. Extraction is faster and less error-prone than retyping.
Handling records you must not upload
Patient lists, salary data and client ledgers cannot be sent to an online converter without creating a disclosure. Processing locally avoids the question entirely.
What actually happens to your file
Each page is parsed with pdf.js, which returns the page's text items along with a transform matrix for each one. Two numbers from that matrix matter here: the vertical position and the horizontal position of the item on the page.
Items are bucketed by rounded vertical position, so everything printed at the same height lands in the same bucket — that bucket becomes a spreadsheet row. Within it, the items are sorted by their horizontal coordinate, which puts them in left-to-right reading order and turns each one into a cell.
The rows are then ordered from the top of the page downwards and handed to SheetJS, which builds a worksheet from that array of arrays. Each page gets its own sheet, named after the page number, and the workbook is written out as .xlsx in the browser.
Note what this does and does not use: position on the page, and nothing else. There is no table-detection model and no ruling-line analysis. Column structure is inferred purely from where things were printed.
Where it falls short, honestly
- Columns are inferred, not read. Because cells are ordered by horizontal position rather than assigned to detected columns, a row with a blank cell shifts everything after it one place to the left. Expect to tidy up alignment.
- Nothing outside the table comes through as structure. Headers, footers, page numbers and paragraphs of prose all land in the grid as ordinary rows.
- Formatting, merged cells and formulas are not produced. You get values as text and numbers, in plain cells.
- Ruled lines and cell borders are ignored entirely. A table drawn with visible gridlines is treated exactly like columns of text with wide spacing.
- A scanned PDF yields nothing. Without a text layer there are no coordinates to position anything by. The OCR tool can read those pages, but it returns plain text rather than a positioned layer, so the column structure will not survive.
- Very wide tables can wrap awkwardly if the source itself splits a table across pages, because each page becomes a separate sheet.
Compared with an online converter
| This tool | Typical cloud converter | |
|---|---|---|
| Where your file goes | Nowhere. It stays in the browser tab. | Uploaded to a server you do not control. |
| Suitable for financial records | Yes — no disclosure occurs. | Depends on their terms and jurisdiction. |
| Page limits | Bounded only by your device's memory. | Commonly capped on free tiers. |
| Table detection | Position-based; alignment may need tidying. | Sometimes uses trained models, often better. |
| Works offline | Yes, once the page has loaded. | No. |
| Cost | Free, no account. | Usually metered or subscription. |