We covered a Social Security number with a black box. The text was still there.
Drawing a black rectangle over something you want hidden feels like it should work, and visually, it does: open the PDF, look at the page, the number is gone. We wanted to know what actually happens underneath that rectangle, so instead of describing it in the abstract, we built a test document, redacted it exactly the way a person would using an ordinary PDF editor, and then tried to read the hidden text back out. Here is precisely what we did and what came back.
The test document
We created a one-page PDF resembling a loan application summary, with a name, a Social Security number, an account number, and an approved amount, each on its own line of ordinary text. This is a stand-in for the kind of page that actually gets redacted in practice: a form with a few sensitive fields sitting among fields that are fine to share.
The redaction
Using the exact technique a typical PDF annotation tool uses — drawing a solid black rectangle on top of the page, positioned over the Social Security number line — we covered it completely. Viewed normally, the result looks exactly like a properly redacted document: a black bar sits where the number used to be, and nothing is visible through it.
The extraction
We then ran the covered document through ordinary PDF text extraction — the same operation any "copy text from PDF," "PDF to text," or "PDF to Word" tool performs, and the same thing a search engine's indexer or a basic screen reader does automatically. The extracted text read: "Loan Application — Applicant Summary Applicant name: Jordan Ellis Whitfield Social Security Number: 044-88-2671 Account number: 8827-4410-9931 Approved amount: $42,500.00."
The Social Security number came back in full. So did everything else. The black rectangle changed nothing about what the document contains — only what a human looking directly at the rendered page can see.
Why this happens, mechanically
A PDF page is not a flat image; it is a set of layered drawing instructions. Text is one instruction — draw these characters, in this font, at this position. A black rectangle is a separate, independent instruction — fill this region with this colour, at this position. When you draw a rectangle over text, you add a new instruction that happens to render visually on top of the old one. You do not remove, alter, or interact with the text instruction at all. It is exactly as present in the file as it was before you drew anything, sitting underneath a shape that merely happens to obscure it from a human eye looking at a rendered page.
Anything that reads the document's actual content, rather than looking at a rendering of it — text extraction, search indexing, copy-paste, screen readers, and any tool built to specifically strip "visual only" cover-ups — sees straight through the rectangle, because there is nothing there to see through. The text was never covered from that perspective; it was covered from one specific perspective only, a human looking at the page.
This is not a hypothetical risk
Redaction failures of exactly this kind — sensitive information covered visually but still present in the underlying file, later recovered by copying the text or examining the file's structure — have caused real, documented incidents in legal filings, government document releases, and corporate disclosures over the years. This is not an obscure edge case in PDF internals; it is one of the most common and well-known mistakes in handling sensitive documents, precisely because the visual result looks so convincingly finished.
What genuine redaction requires
Properly removing content means the drawing instructions for that content are deleted from the page, not covered by a new instruction layered on top. Software specifically designed for redaction does this: it identifies the actual text or image objects in the region you select and removes them from the page's content stream entirely, so there is nothing left underneath to extract, no matter how the file is subsequently examined.
- If your tool's own documentation describes drawing a shape, a highlight, or an annotation over content, that is very likely a cover-up, not a redaction, whatever the feature is named in the interface.
- If a tool explicitly describes removing or deleting the underlying content — not just obscuring it — and ideally lets you verify the result by running text extraction on the output, that is doing the real operation.
- When in doubt, run the test we ran: extract the text from your "redacted" document yourself before you send it anywhere. It takes under a minute and tells you, with certainty, whether the sensitive content is actually gone.
A note on the numbers used in this test
The name, Social Security number, and account number in our test document were invented specifically for this test and do not correspond to a real person or account. We are describing our own methodology openly, including sharing the exact extracted output, precisely so the claim is checkable rather than taken on faith: build the same kind of test file yourself, cover a line with a shape in any PDF editor, and run it through any text extraction tool. The result will be the same, because it follows from how the PDF format represents layered content, not from anything specific to the tool you use.
What this means for using our own tools
Our edit tool lets you draw shapes and text over a page, which is genuinely useful for annotating drafts, filling in forms, and marking up documents — and, as demonstrated above, it does not remove the content underneath, because that is not what drawing on a page does at any level. If you need to permanently eliminate a whole page, the remove pages tool actually deletes it from the resulting document rather than covering it. Removing a specific piece of text from within a page you are keeping — a name in the middle of an otherwise-shareable paragraph — is a harder problem that a shape tool cannot solve safely, and is worth treating with the caution this test should have earned it.
One more consequence worth flagging: the same layering applies to images, not just text. Blurring or pixelating a photograph is often done by placing a processed copy over the original at the display layer, which can leave the unprocessed original image object still embedded in the file underneath, recoverable by anyone who extracts embedded images rather than viewing the rendered page. The underlying lesson is the same regardless of content type: covering is not removing, and the only way to be sure something is gone is to check what the file contains, not what the page shows.