What is OCR?
OCR, optical character recognition, turns a picture of text into text you can edit, search and copy. Here is how it works, what affects how well it works, and where it stops.
How OCR works
- Clean up the picture. The page is turned into black and white, straightened and cleaned of specks, so the letters stand out from the paper.
- Find the text. The page is split into blocks, lines and words, and pictures and tables are set aside.
- Recognise each line. Modern OCR reads a whole line at once with a neural network trained on millions of lines of text, and uses the language's dictionary to settle doubtful letters.
- Hand back text and positions. Every word comes back with where it sits on the page, which is what lets a converter rebuild the layout around it.
What makes OCR accurate
Helps
- 300 dpi or more, and sharp focus
- Dark text on a plain, light background
- A straight, flat page, evenly lit
- The right language chosen
Hurts
- Low resolution, blur and shadows
- Handwriting, which print-trained OCR is not built for
- Decorative fonts, and text over pictures
- The wrong language: its letters get read as something else
Tesseract, the OCR engine here
This converter reads scanned pages with Tesseract, an open source OCR engine first built at Hewlett-Packard, released as open source in 2005 and developed with Google's support after that. Since version 4 it recognises lines with a neural network, and it has language data for more than 100 languages; English, Odia and Hindi are built in here.
OCR runs inside the converter, so a scanned page is not sent anywhere else to be read.
Convert your PDF to Word now
Free and without a sign-up. Choose a file, pick its language if it is a scan, and the Word document downloads when it is done.
Convert a PDF to Word