When you receive a PDF that was scanned from a physical document — a signed contract, an old report, a printed form — the content exists as an image, not selectable text. To convert it to an editable Word document, you need OCR (Optical Character Recognition).
Text-Based PDF vs Scanned PDF
Text-based PDF: Created from a digital document (Word, Google Docs, PowerPoint). The text is stored as actual characters. You can click and select words in any PDF viewer.
Scanned PDF: Created by photographing or scanning a physical document. The page content is an image. You cannot select the text.
To test which type you have: open the PDF in Chrome, Firefox, or any viewer, and try to highlight a word by clicking and dragging. If you can — it's text-based. If you cannot — it's scanned.
What Is OCR?
OCR (Optical Character Recognition) is software that analyses a raster image, identifies letter shapes, and converts them to machine-readable text. It is the technology behind scanning apps like Adobe Scan, Microsoft Lens, and online converters.
OCR works by:
- Pre-processing the image (contrast adjustment, deskewing)
- Detecting text regions
- Identifying individual characters
- Assembling characters into words and sentences
How to Convert a Scanned PDF to Word
If the PDF is a scan, a regular PDF-to-Word converter will not work — it will produce empty or near-empty output because there is no text layer to extract.
The correct process:
- Go to PDF to Word
- Upload your scanned PDF
- The converter automatically detects the scan and applies OCR
- Download the
.docxfile - Review and correct any OCR errors
What to Expect From the Output
High-quality scans
A high-resolution scan (300 DPI or higher) of clearly printed, undamaged text will convert with very high accuracy. The output is often usable with minimal correction.
Lower-quality scans
Faded text, skewed pages, low-resolution scans, or documents with irregular fonts will have higher OCR error rates. Expect to correct words and formatting manually.
Tables and columns
OCR converts tables and multi-column layouts into text, but the structure is rarely preserved perfectly. You will likely need to rebuild tables manually in Word.
Handwriting
Handwritten text is outside the scope of standard OCR tools. Do not expect accurate conversion of handwriting.
Improving OCR Results Before Converting
If you control the scanning process, these settings improve OCR accuracy:
- Scan at 300 DPI minimum — 600 DPI for small fonts
- Use black and white — not colour, for high-contrast text documents
- Ensure pages are flat and aligned — curled edges reduce accuracy
- Good lighting — avoid shadows across the page
- High-contrast originals — faded or low-contrast originals always produce more errors
After Converting: Review and Correct
Every OCR conversion should be proofread. Common errors to look for:
0(zero) confused withO(letter)l(lowercase L) confused withI(uppercase i) or1(one)rnread asm- Word spacing errors
- Line breaks inserted mid-sentence
- Numbers reversed or garbled
Use Word's spell check as a starting point, but it will not catch correctly spelled words that are the wrong word.
Related guides: How to Convert PDF to Word · Common PDF Conversion Problems