BASALT · JOURNAL

How to anonymize research documents in PDF form

2026-09-21 · anonymize pdf research documents

Anonymizing research documents is not the same as deleting names. A participant may be identifiable from an unusual role, location, event date, diagnosis, quotation, or combination of ordinary facts. The research protocol, consent, ethics approval, and applicable law determine the acceptable risk.

Separate direct and indirect identifiers

Direct identifiers include names, contact details, participant IDs linked to an identity table, signatures, faces, and account or record numbers. Indirect identifiers include precise dates, small locations, rare occupations, unique events, and combinations that narrow the population.

Create a transformation plan for each field: remove, generalize, replace consistently, or retain under the approved method. If stable pseudonyms are needed for analysis, keep the re-identification key outside the PDFs with tighter access.

Review free text as free text

Interview transcripts and case notes contain identifiers in narrative form that regular expressions cannot reliably classify. Search known names and places, then have a reviewer read for context. Consider whether distinctive verbatim quotations could be found elsewhere on the web and linked back to the speaker.

Scans need both pixel removal and OCR-layer cleanup. Remove metadata, comments, tracked-review artifacts converted into annotations, attachments, and source filenames.

Test the released corpus

Search across the entire output set, not just one PDF, because separate documents can combine to identify a participant. Review pseudonym consistency and look for rare combinations. Independently extract text and enumerate hidden structures, then document the method and residual-risk decision.

For health research, use authoritative requirements such as HHS de-identification guidance, not a generic blog summary.

Frequently asked questions

Is removing names enough to anonymize research data?

No. Indirect identifiers and combinations of facts can identify participants even when direct names are gone.

What is the difference between anonymization and pseudonymization?

Pseudonymization replaces identifiers while retaining a separate means of re-identification. Anonymization aims to prevent identification under the applicable risk standard.

Should the pseudonym key be embedded in the PDF?

No. Keep it in a separate, access-controlled system; embedding or attaching it defeats the separation.

Doing it in Basalt

Basalt searches a document, removes text and image content, clears OCR and hidden data, and produces reproducible output with a verifier certificate. It supports the file transformation but does not decide whether a dataset meets a legal or ethical anonymization standard. Learn more.

Redaction that proves itself

Basalt destroys the content you mark, then re-opens the file it wrote and proves the content is gone before it saves anything. Your documents never leave your Mac.

DOWNLOAD BASALT 2.3.0 BUY $29 FREE FOR 24 HOURS · MACOS 13+