Home › Guides › PDF metadata

How to check and remove PDF metadata

A PDF usually carries your name, your software, and when it was made. Often none of that is meant for the recipient.

PDFs carry a metadata block that most people never see. It typically holds the author's name, the producing application, creation and modification timestamps, and often the original filename and title.

This has embarrassed a lot of organisations. A document sent as anonymous carrying an author's name, or a tender response revealing it was created after the deadline, are both routine occurrences.

Step-by-step

  1. Add your PDF.
  2. Read what it contains.
  3. Strip the fields you do not want to share.
  4. Download the cleaned copy.

What is usually in there

What removing metadata does not do

It cleans the document's properties. It does not touch the content, and it is not a way of hiding information written into the page.

If you covered something with a black rectangle, removing metadata does nothing about it — the text is still underneath. That needs proper redaction, which deletes the content rather than drawing over it.

Before sending anything sensitive

Check the metadata, check for tracked changes if it came from a word processor, and check that anything obscured was actually removed rather than covered. Those three account for the overwhelming majority of accidental disclosures.

A PDF can carry its own earlier drafts

PDFs support incremental saving: rather than rewriting the file, an editor can append the changes to the end and add a new cross-reference table pointing at them. It is fast and it is safe against corruption, and it means the previous version of the content is often still physically present in the file, further up, simply no longer referenced.

The practical consequence is that a page which was edited — a figure replaced, a paragraph removed, a signature block changed — may still contain what was there before, invisible to every viewer and recoverable by anything willing to read the earlier revision. Stripping metadata does nothing about this, because it is not metadata; it is old content.

The reliable way to leave it behind is to write a fresh file rather than to save over the existing one. An export, a print-to-PDF, or any process that rebuilds the document from scratch produces a file with no history in it. If a document has been through several rounds of editing and is about to go somewhere sensitive, that rebuild is worth doing on principle.

Check the file, not the page

The general lesson behind both metadata and revision remnants is that a PDF is a container and the page you can see is a rendering of part of it. Author names, the software that produced it, creation and modification times, and sometimes a full filesystem path from the machine it was made on all sit outside anything that appears on screen — and so do attachments, form field values that were filled and cleared, and layers that are switched off.

None of that is exotic or hard to read; it is simply not shown. Inspecting a file before sending it takes seconds and is the only way to know what is actually in it.

Frequently asked questions

Is the PDF uploaded?

No. It is read and rewritten in your browser.

Does this remove text I covered with a rectangle?

No. That text is still in the file underneath the rectangle. Use the redaction tool, which removes the content itself.

Will removing metadata change how the document looks?

No. Metadata is separate from page content, so the document renders identically.

Open the PDF tools →