What is OCR?

OCR, optical character recognition, turns a picture of text into text you can edit, search and copy. Here is how it works, what affects how well it works, and where it stops.

How OCR works

  1. Clean up the picture. The page is turned into black and white, straightened and cleaned of specks, so the letters stand out from the paper.
  2. Find the text. The page is split into blocks, lines and words, and pictures and tables are set aside.
  3. Recognise each line. Modern OCR reads a whole line at once with a neural network trained on millions of lines of text, and uses the language's dictionary to settle doubtful letters.
  4. Hand back text and positions. Every word comes back with where it sits on the page, which is what lets a converter rebuild the layout around it.

What makes OCR accurate

Helps

  • 300 dpi or more, and sharp focus
  • Dark text on a plain, light background
  • A straight, flat page, evenly lit
  • The right language chosen

Hurts

  • Low resolution, blur and shadows
  • Handwriting, which print-trained OCR is not built for
  • Decorative fonts, and text over pictures
  • The wrong language: its letters get read as something else

Tesseract, the OCR engine here

This converter reads scanned pages with Tesseract, an open source OCR engine first built at Hewlett-Packard, released as open source in 2005 and developed with Google's support after that. Since version 4 it recognises lines with a neural network, and it has language data for more than 100 languages; English, Odia and Hindi are built in here.

OCR runs inside the converter, so a scanned page is not sent anywhere else to be read.

Convert your PDF to Word now

Free and without a sign-up. Choose a file, pick its language if it is a scan, and the Word document downloads when it is done.

Convert a PDF to Word