What is OCR (optical character recognition)?

OCR, or optical character recognition, is the technology that converts the image of text, in a scanned document or a photo, into machine-readable text that can be searched, copied, and processed.

What it actually means

When you scan a paper document, what you get is a picture. The page looks like text to a human, but to a computer it is just an image, with no idea what words are on it. OCR reads that image and works out the actual characters and words, producing a text layer that sits alongside the picture.

That text layer is what makes everything else possible. Once a scan has been through OCR, its words can be searched, a system can read it to figure out what kind of document it is, and specific terms like dates and clauses can be pulled out of it. Without OCR, a scanned archive is a stack of photographs of information, technically stored and practically unusable.

Why it matters

OCR is the bridge between paper and everything a modern document system can do. It is the difference between an archive you can only look at one document at a time and an archive you can query as a whole. For organizations digitizing decades of paper, OCR is what turns the effort of scanning into something that actually pays off, because it is what makes the scanned documents findable rather than merely stored.

What OCR does and does not do

OCR reads text reliably when the text is typed or printed and the scan is reasonably clear, including very old documents. It struggles more with handwriting and poor-quality scans, and it does not understand what a document means, it only recognizes the characters. Understanding what a document is and where it belongs is a separate step called classification. OCR reads the words; classification makes sense of them.

PaperlessZen™ runs OCR on every page during intake, so full text is searchable. See intake and search, and read why scanning without classification just moves the pile. Related terms: full-text search, document classification.