BASALT · JOURNAL

Why your scanned PDF is not searchable

2026-08-15 · scanned pdf not searchable

A scan holds a picture of text rather than text, so searching finds nothing even for words plainly visible on the page.

Confirm it

pdftotext document.pdf - | head

Nothing returned means no text layer. A quicker check in a viewer: try to select a word, and if the cursor draws a rectangle rather than highlighting characters, it is an image.

What OCR does

Recognition reads each page image and writes an invisible text layer positioned over the picture. The page looks unchanged and becomes searchable, selectable and extractable. On macOS this runs on device with no upload.

Why a search can still miss things

OCR is imperfect. Poor scans, unusual fonts, stamps, handwriting and skewed pages all produce errors, and an error means the term you searched for is not found on a page where it is clearly legible.

This matters most when search is being used to find material to remove. A zero result is not proof of absence, and a review that relies on search alone will miss pages. Read the pages that matter.

Improving recognition

Scan at 300 DPI or better. Below that accuracy falls off quickly.

Deskew before recognising, because rotated text recognises poorly.

Greyscale rather than colour is fine and often better, since colour adds no information for typed text.

The consequence for redaction

Once OCR has run, the document contains real machine readable text under the image. Redacting the picture without removing the corresponding text produces a document that looks clean and is fully extractable.

Do OCR first, then mark, then redact both layers, then verify by extracting text from the finished file and searching it for what you removed.

Frequently asked questions

Why is my scanned PDF not searchable?

Because it contains a picture of text rather than text. Confirm with pdftotext document.pdf -, which returns nothing for a scan, then run OCR to add an invisible text layer.

Does OCR always find every word?

No. Poor scans, unusual fonts, stamps and handwriting produce recognition errors, so a search can miss pages where the term is plainly visible. Treat a zero result as inconclusive rather than as proof of absence.

What scan settings improve OCR accuracy?

300 DPI or better, deskewed, and greyscale is fine for typed documents. Below 300 DPI accuracy drops quickly, and skewed pages recognise poorly.

Should I OCR before or after redacting?

Before. OCR first so you can search and find the material, then mark, then redact both the image and the text layer, then verify the output by extracting its text.

Doing it in Basalt

Basalt is a native macOS PDF toolkit: eighteen tools in one window covering merge, split, page organisation, compression, OCR, passwords, Bates numbering, forms, signing, watermarks and comparison. Every file is processed on your Mac, and the engine that opens documents holds no network entitlement at all, which macOS enforces at the code-signature level. Redaction destroys content rather than covering it, and an independent verifier proves the material is gone before a file is written. A one time $29 licence covers up to three Macs, free for the first 24 hours. Get it at basaltformac.com, or brew install --cask chipmunk1101/tap/basalt.

Redaction that proves itself

Basalt destroys the content you mark, then re-opens the file it wrote and proves the content is gone before it saves anything. Your documents never leave your Mac.

DOWNLOAD BASALT 2.3.0 BUY $29 FREE FOR 24 HOURS · MACOS 13+