
Marta Kowalska
My starting observation is that the metadata is the part of a document nobody chooses to fill in and everybody publishes. The standard set includes a producer, a creation and modification date, and sometimes an author and a title. The author field is typed by whoever saved the file and is regularly a person's full name. Titles from authoring tools often append the application and the platform, which tells a reader more than intended. The document information dictionary is not the only place. XMP metadata is a separate XML packet in the file that many applications prefer and that survives some edits that clear the first dictionary. A tool that clears one and not the other leaves half the trail, so I cover both. Attachments, annotations and form values can all carry names, and embedded files can carry their own metadata. A redacted document that still carries a comment thread from review is a disclosure in the ordinary sense of the word. Stripping is straightforward and has one caveat worth stating: it changes the file, so you should compare before and after to confirm the document itself is intact. Removing metadata is not redaction, and a document with a clean metadata block and an intact text layer has not had anything removed. I also cover what metadata is useful for in practice, which is more than people assume. It drives search indexing and accessibility, so a blanket strip is not always the right move. I close by separating metadata from content, because the two are routinely confused. A file can have no metadata and still contain everything you wanted removed, and it can carry a full set of them with nothing sensitive in the document at all.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →