The challenges of converting PDF to Word

A PDF records where each letter goes on a page, not what the text means. Turning it back into a Word document means working out paragraphs, tables and reading order that the file never stored.

A printout, not a manuscript

A Word document is a manuscript: paragraphs, headings, lists and tables that flow onto as many pages as they need. A PDF is closer to a printout: it says "draw these letters in this font at this spot", page by page, and nothing more. That is why a PDF looks the same everywhere, and why getting an editable document back out of one is hard.

Below are the twelve problems any PDF to Word converter has to solve, and how this one deals with each. Where it does not handle something yet, it says so.

1Text is stored as letters, not paragraphs

A PDF places each letter, or each run of letters, at an exact position. There is no "paragraph" or "line" in the file, so a converter has to rebuild them from the geometry.

Handled Letters are grouped into lines by their baseline and size, and lines into paragraphs by their spacing and indent. Wrapped prose is joined again, while list items and labels stay separate.

2Reading order

The order in which a PDF draws its text need not be the order you read it. A page in two columns, or a sidebar next to the main text, has to be untangled into one sequence.

Partly Single-column pages come out in reading order. Pages in two or three columns are still read across, line by line, so their columns can mix.

3Fonts you do not have

A PDF usually carries the fonts it uses, often only the letters it needs. Word cannot use those, and the reader's computer may not have the original typeface, so lines can come out longer or shorter than they were.

Handled A missing typeface is replaced by Times New Roman, Arial or Courier New, and the character spacing is adjusted so each line still fits the width it had. The letters look different; the lines and pages fall in the same places.

4Text that does not say what it shows

To draw a letter, a PDF only needs the shape. The mapping back to real characters is optional, and some PDFs leave it out or get it wrong, so the text looks right on screen but copies out as nonsense.

Partly Such pages are not detected on their own: their text comes across as the PDF states it. Turn on Read every page as a scan before you convert, and the words are read from what is shown instead.

5Scanned pages

A scan is a picture of a page. There is no text in it at all, only pixels, so the words have to be recognised before anything else can happen.

Handled Pages with almost no text are read with Tesseract OCR, and the words go through the same layout step as text pages. How well it works depends on the scan; bold and italic cannot be told apart. Scanned PDF to Word

6Text turned into shapes

Some print drivers, "Microsoft: Print To PDF" among them, can turn letters into drawn outlines. The page has no text and no picture, so a simple check thinks it is empty.

Handled Such pages are drawn as a picture and read with OCR, like a scan.

7Tables that are only lines

A PDF has no tables, just text placed in a grid and lines drawn around it. A converter has to find the grid and put each piece of text in the right cell.

Partly Grids of ruling lines become real Word tables, with cells merged across columns and the colours of the lines kept. Tables drawn without lines, and cells merged down a column, are not handled yet. PDF tables to Word

8Line breaks or flowing text

Keep every line break and the Word pages match the PDF, but editing a sentence leaves ragged lines. Join the lines and the text reflows as you type, but the pages no longer match. A converter has to choose.

Handled Every line breaks where the PDF broke it, so the pages match. Run the open source program on your own computer and it can join each paragraph instead.

9Headers and footers

Titles, page numbers and running heads are drawn on every page like any other text. Word keeps them in a separate header and footer, so they have to be recognised as repeating.

Not yet They stay in the body of each page, where they were drawn.

10Indian and other complex scripts

Scripts such as Odia, Hindi, Bengali and Tamil join letters and marks into clusters, and need a font that has them. Without one, Word shows empty boxes.

Partly The Word file asks for Nirmala UI, the Windows font that covers these scripts, so the text renders. Odia and Hindi scans are read with their own OCR; Bengali, Tamil and other scripts are not read in scans yet. Text stored in an older font such as Akruti or Kruti Dev copies out as English letters: Read every page as a scan reads it from the page instead. Odia PDF to Word Hindi PDF to Word

11Colours and drawings

Coloured text, charts and diagrams are drawing instructions in a PDF, with no direct equivalent in a Word paragraph.

Not yet Text comes out in black, and vector drawings other than table lines are left out. Pictures embedded in the PDF are kept, at their size and in their place.

12Passwords

A PDF that needs a password to open cannot be read without it.

Not yet Such a PDF is refused as soon as you choose it, with a message saying so. Open it in a PDF reader that knows the password, save a copy without it, and convert that.

Convert your PDF to Word now

Free and without a sign-up. Choose a file, pick its language if it is a scan, and the Word document downloads when it is done.

Convert a PDF to Word