The metadata hiding in your PDF: author names, software, dates — and how to see and remove it

2026-09-15
7 min read

A recruiter receives a CV whose PDF properties say "Author: Maria", though the candidate is called Ana. A law firm sends a settlement proposal whose title field still reads "Draft v3 — do not send". A whistleblower's document lists the corporate laptop it was created on. None of these were on the visible pages; all of them were in the file. PDF metadata is small, mostly harmless and occasionally the single most sensitive thing in a document. It is worth knowing exactly where it lives.

The two places metadata lives

The older mechanism is the document information dictionary, usually just called the Info dictionary. It is a short list of key–value pairs defined by the PDF specification: Title, Author, Subject, Keywords, Creator (the application that made the original document, e.g. "Microsoft Word"), Producer (the software that wrote the PDF, e.g. "macOS Quartz PDFContext" or "pdfcpu"), CreationDate and ModDate. Every PDF viewer shows these in a "Properties" or "Document info" panel.

The newer mechanism is XMP, an XML block embedded in the file. Adobe introduced it so that the same metadata schema could travel across PDF, JPEG, TIFF and other formats. XMP can hold the same fields as the Info dictionary plus arbitrarily more: a full edit history (when the document was saved and by which application, entry by entry), a document identifier that stays constant across versions, PDF/A conformance claims, copyright notices, and whatever custom fields the producing software chose to write. It is where the interesting leaks tend to be, precisely because most viewers only show a summary of it.

Both blocks can disagree with each other. A file whose Info dictionary was cleaned by a tool that did not know about XMP will happily keep the original author in the XMP block, and a viewer that prefers XMP will display it. Any serious cleanup has to address both.

What typically ends up there

  • The name the operating system or office suite was registered to. Word, LibreOffice and Google Docs all write the account name into Author by default. Corporate installations often write the company name as well.
  • The original filename or title. Word writes the first heading or the filename into Title; if the document was renamed later, the old name stays.
  • Software and version numbers in Creator and Producer, which reveal the platform (Windows, macOS, a particular printer driver) and sometimes the licence type.
  • Creation and modification timestamps, including time zone offsets, which can contradict a claimed date or reveal working hours.
  • In XMP, a history of saves with application names and instance IDs, which can show that a "final" document was edited after it was supposedly finalised.
  • Keywords and Subject, which people fill in for internal indexing ("confidential", "litigation hold", a client code) and forget about.

What is not in metadata: the text of the pages, comments, tracked changes or hidden layers. Those are content, and removing them is a different problem — see the guide on why a black rectangle is not redaction. Metadata is the header of the envelope, not the letter inside.

How to inspect it

In Acrobat Reader: File → Properties → Description shows the Info dictionary; the "Additional Metadata" button (in Acrobat Pro) opens the XMP. In Preview on macOS: Tools → Show Inspector → the first tab. In Firefox and Chrome, which both have built-in PDF viewers: open the file, then look for "Document properties" in the viewer's menu. On the command line, exiftool reads both blocks in full and is the tool to use if you need to see everything, including custom XMP fields that viewers hide.

A minimal, no-install check: open the PDF in a plain text editor and search for "/Author" and "<x:xmpmeta". PDFs are partly text, and unless the file uses object streams (which compress those sections), you will see the raw fields. It is crude, but it makes the point that this information is just sitting in the file for anyone to read.

Three ways to remove it

1. Edit the fields before exporting

The cleanest option is to not write the metadata in the first place. Word's "Inspect Document" (File → Info → Check for Issues) strips author and personal information from the source document; LibreOffice has "Remove personal information on saving" under Options → Security. Do this on the source, then export, and the PDF starts clean. It does not help for PDFs you received rather than made.

2. Clear the fields in the PDF

Acrobat Pro has "Remove Hidden Information" (Tools → Redact → Sanitize), which clears both blocks and, importantly, also removes attachments, comments and hidden layers. Command-line tools do the same more surgically: exiftool -all= clears both Info and XMP in one pass, and qpdf or pdfcpu can rewrite the file without them. Whatever tool you use, re-inspect the result, because "clearing" the Info dictionary in some viewers only sets the fields to empty strings while leaving the XMP block untouched.

3. Re-create the file (the nuclear option)

If the PDF does not need to remain searchable or editable, converting every page to an image and assembling a new PDF from those images removes all metadata, all hidden content, all comments and all layers in one step — because the new file contains nothing but pictures. Our pdf-to-image tool exports each page as PNG or JPEG, and image-to-pdf builds a fresh document from them; both run locally, so the sensitive original never leaves your machine. The cost is real: the text becomes pixels (no search, no copy, no accessibility) and the file usually grows. For a document whose whole point is that nothing else leaks, that trade is often the right one.

Note that tools which rewrite a PDF normally set their own Producer field ("pdfcpu", "Quartz", and so on). That is not a leak — it names the software, not you — but if the goal is a file with no fingerprints at all, check the result rather than assuming.

When to bother

For most documents, never: the author field saying your name on a report you are signing anyway is not a problem. Bother when the file is going somewhere the metadata contradicts the story (a "final" document with an edit history after the deadline), when it is anonymous by design (a bid evaluated blind, a complaint, a submission to a journal with double-blind review), or when it was created on infrastructure that should not be associated with it. In those cases, clean it, verify with a second tool, and keep in mind that the pages themselves may still contain the same information in plain sight.

Need a file with nothing but the pages in it? Export each page as an image and rebuild the PDF — entirely on your device.

Convert PDF to images