BASALT · JOURNAL

Running OCR on a large scanned production

2026-08-09 · ocr large pdf

A scanned production arrives with no text. Every page is a photograph, nothing is searchable, and the first instinct is to OCR the lot. That is usually right, and it has one consequence for redaction work that is worth understanding before you start.

What OCR adds

Optical character recognition reads each page image and writes an invisible text layer positioned over the picture. The page looks exactly as before and becomes searchable, selectable and extractable.

The layer is genuine text in the file, not an annotation. That is what makes it useful and what makes it a hazard.

The redaction trap

Redacting a scanned page means destroying pixels. Redacting an OCRed scanned page means destroying pixels and removing the corresponding text from the invisible layer.

Miss the second step and the result is a document where the picture of the name is gone and the machine readable name is still present, extractable by anyone with a command line tool. Visually it looks perfectly redacted. It is not.

This is a documented failure mode, not a theoretical one, and it is a reason to verify output rather than trust it. Extracting text from the finished file and searching for the removed terms takes seconds and catches exactly this.

How long OCR takes

Longer than most operations, because recognition is genuine image analysis on every page. Plan for a large scanned production to take a meaningful stretch of machine time, and run it as an overnight or background job rather than something you wait on.

The cost is per page and largely independent of how much text is on the page, so a production of mostly blank forms costs about the same as one of dense correspondence.

Accuracy and what it means for review

OCR is imperfect. Poor scans, unusual fonts, handwriting and stamps all produce errors, and an error means a term you searched for is not found on a page where it plainly appears.

The practical consequence: a search based review over OCRed documents will miss pages. That is an argument for treating search as an aid rather than a complete method, and for reading pages that matter rather than trusting a zero result.

Order of operations

OCR first, then review and mark, then redact. Running OCR after redaction re-reads the redacted images and produces a text layer for a document you have already finished, which at best wastes time and at worst reintroduces recognised text near the areas you removed.

Frequently asked questions

Should I OCR a scanned PDF before redacting it?

Yes, because you cannot search a document with no text layer, and searching is how most material is found. Do it before marking, not after, so the redaction removes both the pixels and the matching text.

Can OCR text survive redaction?

Yes, and it is a real failure. If the redaction destroys the page pixels but leaves the invisible OCR text in place, the document looks correctly redacted while the words remain machine readable and extractable.

How long does OCR take on a large PDF?

Long enough to schedule rather than wait for, because recognition analyses every page image. Cost is per page and roughly independent of how much text each page contains, so run large productions as a background job.

Is OCR accurate enough to rely on for review?

Not on its own. Poor scans, unusual fonts, stamps and handwriting all produce recognition errors, and an error means a search misses a page where the term is clearly visible. Treat search as an aid rather than proof of absence.

Doing it in Basalt

Basalt is a native macOS PDF toolkit with eighteen tools in one window. It opens large documents without loading them into memory, renders pages on demand, and copies files into its engine in fixed-size chunks, so peak memory follows the chunk size rather than the file size. Redaction destroys content rather than covering it, and an independent verifier re-opens every written file to prove the material is gone before the file is saved. A one time $29 licence covers up to three Macs, it is free for the first 24 hours, and the engine holds no network entitlement at all, which macOS enforces at the code-signature level. Download it at basaltformac.com.

Redaction that proves itself

Basalt destroys the content you mark, then re-opens the file it wrote and proves the content is gone before it saves anything. Your documents never leave your Mac.

DOWNLOAD BASALT 2.3.0 BUY $29 FREE FOR 24 HOURS · MACOS 13+