What actually happens when you compress a PDF (with real numbers)
"Compress this PDF" sounds like one operation. In practice it is at least two entirely different things wearing the same label, and confusing them is why so many people end up with a file that got bigger, not smaller, after they tried to shrink it. We ran both kinds against two real documents and measured exactly what happened, because a specific number is worth more than a general claim.
The two things "compression" can mean
The first is structural optimisation: a PDF accumulates waste as it is edited and re-saved — objects nothing points to any more, duplicate font descriptors, a cross-reference table padded with the history of every incremental save. Cleaning that up is completely lossless. Not one pixel of an image or one character of text is touched; the file is simply rewritten without the debris.
The second is rasterisation: every page is rendered to a picture and the picture is saved back into a JPEG-based PDF, discarding whatever vector text and full-resolution images were there before. This shrinks files dramatically, but it is not free — it trades the original content for a photograph of it.
Most tools, ours included, offer both, and the second one usually goes by a friendlier name than "rasterise and discard the text layer." That naming is where the trouble starts.
The test
We built two documents specifically to expose the difference. The first was an 83,343-byte, 40-page report generated cleanly, with nothing but ordinary paragraph text — the kind of PDF that comes out of a word processor's own export. The second was an 33,725,479-byte, 8-page scan: eight full-resolution photographs at roughly 300 DPI, standing in for what a phone camera or a document scanner actually produces.
We ran each through the three compression levels our own tool offers, timed them, and measured the output size in bytes rather than describing it qualitatively.
Level 1: lossless structural optimisation
This is the level that behaves the way most people assume "compress" behaves: it removes waste and touches nothing else.
- Text report (83,343 bytes) → 73,573 bytes. A 12% reduction, in 129 milliseconds, with every character still selectable.
- Scanned document (33,725,479 bytes) → 33,722,980 bytes. Effectively 0%, in 97 milliseconds.
Both results are correct behaviour, not a malfunction. The text report had a little redundancy to remove and this mode found it. The scan is almost entirely image data with no structural waste to speak of, so there was nothing this mode is designed to touch — it does not resample or requantise images, by design, which is exactly why it cannot degrade them either.
Level 2: moderate rasterisation
This used to be the setting selected automatically when you opened the tool — it no longer is, for the reason this section is about to demonstrate. It renders each page at full scale and re-encodes it as a JPEG at quality 0.4.
- Text report (83,343 bytes) → 4,159,480 bytes. Roughly fifty times larger, in about a second.
- Scanned document (33,725,479 bytes) → 291,399 bytes. A 99.1% reduction, in about six seconds.
Read those two lines side by side. The identical setting, applied to two different kinds of document, produced a fifty-fold increase in one case and a hundred-fold decrease in the other. That is not noise or an edge case — it is the direct, predictable consequence of turning vector text into a full-resolution photograph of itself. A page of text stored as outlines and font references is astonishingly compact; the same page stored as a JPEG has to spend bits describing every pixel of white space between the letters.
Level 3: aggressive rasterisation
Half the render scale, quality 0.2 instead of 0.4.
- Text report (83,343 bytes) → 598,507 bytes. About seven times larger, in roughly a second.
- Scanned document (33,725,479 bytes) → 85,731 bytes. A 99.7% reduction, in about eighteen seconds.
Smaller than the moderate setting in both directions, as you would expect from a harsher render scale and quality value. Still catastrophic for the text document; now genuinely excellent for the scan, at the cost of visibly softer detail.
Why the same button produces opposite results
The mechanism explains the outcome completely once you see it laid out. Structural optimisation looks at the container — the object graph, the cross-reference table — and never touches what is drawn on the page. Its ceiling is set by how much waste happens to be in that particular file, which is why a clean export barely moves and a heavily-edited file can shrink noticeably.
Rasterisation ignores the container and photographs the contents instead. For an already-photographic page — a scan — that photograph is a more efficient JPEG than whatever the scanner produced, so the file shrinks. For a page whose content was already compact vector instructions, replacing it with a photograph is strictly worse: you are storing the same visual information in a fundamentally less efficient format, on purpose, and paying for it in bytes.
Put plainly: rasterisation is a tool for photographs, and a page of ordinary text is not a photograph, no matter how it arrived in the PDF.
The practical rule
Before compressing anything, ask one question: is this page mostly text, or mostly image? If you are not sure, open it and look — a scanned document looks like a photograph of paper because it is one; a report exported from Word or Google Docs looks crisp at any zoom level because it is vector text.
- Mostly text (reports, contracts, letters, invoices generated digitally): use lossless structural optimisation. It is the only mode that cannot make the file larger, and the only one that keeps the text searchable and selectable afterwards.
- Mostly image (scans, photographed documents, PDFs built from camera photos): rasterisation is the right tool, and the more aggressive setting is usually worth trying first, since the quality loss at 99%+ reduction is often invisible for reading purposes.
- Mixed documents: compress, then actually open the result and look at it before sending it anywhere. A number on a screen does not tell you whether page 14 is now unreadable.
If your tool defaults to an aggressive setting regardless of content — as many still do, ours did until this measurement — that default is a convenience for scan-heavy use cases and a trap for everyone else. Ours now defaults to lossless structural optimisation and only switches to rasterisation when you pick it yourself, and it refuses to hand you a "compressed" file that is actually bigger than what you uploaded. Check which mode is selected before you click through on a document you have not tested before, whichever tool you use.
What this does not fix
None of this is a substitute for reading text out of a scan — that is a job for OCR, not compression, and the two operate on completely different problems. It also will not recover quality already lost: once a page has been rasterised at low quality and saved, a second pass of any compression mode works on the JPEG that already exists, not the original vector content, which is long gone.
One more thing worth measuring rather than assuming: none of this took more than eighteen seconds, and all three modes ran on ordinary consumer hardware inside a browser tab, with the file never leaving the device. Whatever compression tool you use, that is a reasonable benchmark to hold it to — if it takes noticeably longer or needs an upload first, you are paying for infrastructure the operation does not actually require.