Why tables from a PDF break in Excel: there are no tables in a PDF
You export a bank statement or a price list from PDF to Excel and get something that looks almost right: the numbers are there, but a column is shifted by one, a long product name has spilled into the next cell, two header rows have merged into one, and the totals no longer sit under their columns. It feels like a bug. It is, in fact, the honest result of asking a converter to recover information that was thrown away when the PDF was made.
What a PDF actually stores
Inside a PDF, a page is a list of drawing instructions: set this font, move to coordinate (72, 700), show the string "Invoice", move to (300, 700), show "2026-03-14", draw a line from here to there. There is no instruction that says "begin table" or "this is column 3". The table you see is an emergent property of alignment: numbers happen to share an x-coordinate, rules happen to be drawn between them. The application that produced the PDF knew it was a table; the PDF does not.
There is one exception. Tagged PDFs — files with a logical structure tree, as required for accessibility and for PDF/A-1a — can mark a region as a Table with TR and TD elements, exactly like HTML. When a converter finds those tags, it can rebuild the table precisely. Most PDFs are not tagged, and many that are have tags generated automatically and wrongly. In practice a converter has to fall back to geometry.
How a converter guesses the grid
The basic algorithm, which is what our PDF-to-Excel tool does, is straightforward. Extract every text fragment with its x and y position. Group fragments that share the same y into a row. Within a row, sort fragments by x and treat each as a cell. Emit the rows top to bottom, one sheet per page. On a simple table — one line per row, one value per column, consistent alignment — this reproduces the grid exactly, because the geometry really does encode it.
More elaborate converters add heuristics: they look at ruling lines to infer column boundaries, cluster x-positions across the whole page so a value that is slightly offset still lands in the right column, and try to detect header rows by font weight. These help, and they also introduce their own failure modes. None of them recovers information that is not in the geometry.
The five things that break it
1. Wrapped text
A cell whose contents span two lines is, geometrically, two text fragments at two different y-coordinates. A row-by-y converter sees two rows: the first with all the columns, the second with a single orphan fragment. In the spreadsheet, the continuation appears as a row of its own, shifting nothing but confusing everything below it if you sort. The fix is post hoc: join the orphan row to the one above.
2. Merged and spanning cells
A header that spans three columns ("Q1 2026") sits at a single x-position. The converter puts it in one cell, and which one depends on whether it was left-aligned or centred. The two columns it was supposed to cover get nothing. The same goes for a "Total" label that spans the first two columns of a footer row.
3. Right-aligned numbers
Numbers in a financial table are usually right-aligned, so their starting x-coordinate changes with their length: "1,250.00" starts further left than "50.00". A converter that assigns columns by starting position may put them in different columns; one that uses the ending position handles this case but then breaks on left-aligned text. Good converters cluster by the centre or by ruling lines; simple ones do not.
4. Empty cells
A row with a blank in column 4 has fragments only for columns 1, 2, 3 and 5. Sorted by x, they become columns 1, 2, 3 and 4 — the blank vanishes and everything after it shifts left. This is the cause of the "one column off" effect that makes a converted table look right at the top and wrong further down, where the sparse rows are. Only a converter that knows the column boundaries from other rows or from lines can leave the gap.
5. Scanned pages
If the PDF is a scan, there are no text fragments at all — just an image. A converter that reads text will produce an empty sheet. Running OCR first produces fragments, but their positions are approximate and their content may be wrong, so the geometry is noisier than in a born-digital PDF. It works for clean, well-aligned scans and degrades from there.
Getting a clean sheet anyway
- Ask for the source. A statement, a report or a price list almost always came from a spreadsheet or a database; the PDF is a printout of it. One email saves an hour of cleaning.
- Convert, then fix in Excel rather than in the PDF. Sorting orphan rows, filling shifted columns with a formula, and removing repeated page headers are all fast in a spreadsheet and impossible in the PDF.
- For simple tables, use the plain-text export instead. Our PDF-to-text tool gives you the lines as text; pasting them into Excel and using Data → Text to Columns with a fixed-width split often gives a better grid than a geometric converter, because you choose the boundaries.
- Check the numbers, not just the layout. Thousands separators, negative numbers shown in parentheses or with a trailing minus, and dates in the source locale's format are all imported as text unless the converter recognises them. Reformat the columns and verify a subtotal against the PDF before trusting the file.
- For scans, OCR first, then convert, and expect to proofread. If the table is large and the scan is poor, retyping the totals and sampling a few rows may genuinely be faster.
Why the other direction is easy
Excel to PDF is the reverse problem and it is trivial by comparison: the spreadsheet knows exactly where every cell is, and the converter only has to draw them. The difficulty there is purely about pagination — which columns fit on a page, whether to repeat the header — not about meaning. The asymmetry is the whole lesson: a PDF is a final rendering, and every conversion out of it is reconstruction. When you have a choice, convert from the richer format to the poorer one, and keep the richer one.