You received a scanned PDF. You need the text. But when you try to copy, you get nothing. Search finds nothing. The content is there—you just can't access it.
Scanned PDFs are images, not text. The document looks like pages, but underneath, it's just pixels. To access the text, you need OCR (Optical Character Recognition) — technology that converts images of text into actual, selectable, searchable text.
This guide covers text extraction from scanned PDFs: how OCR works, step-by-step instructions, accuracy factors, and tips for getting the best results.
Actual example: A genealogist was researching family history and found a scanned death certificate from 1920 in a university archive. The document was an image-only PDF — no text layer. Running OCR extracted the handwritten names, dates, and locations, letting her search and cross-reference records across dozens of documents instantly.
Why Scanned PDFs Don't Have Extractable Text
Scanned Documents Are Photographs of Pages
When you scan a document, the scanner takes a photograph of each page. The output is a series of images — one image per page — packaged in a PDF wrapper. What is created is an image file (JPEG, TIFF, or similar) per page with no text layer, just pixel data, and the PDF format contains images only. It looks like a document but acts like a photo.
No Text Layer Exists in a Scan
A normal PDF stores text as text — characters that can be selected, copied, and searched. A scanned PDF stores everything as images. Normal PDF structure includes text objects with character codes, font information for rendering, layout coordinates for positioning, and a searchable text layer. Scanned PDF structure includes only image objects per page, no text objects, no character information, and search requires image analysis (OCR).
Text Is Embedded as Pixels, Not Characters
When you scan a document at 300 DPI, each letter becomes a grid of colored squares (pixels). The computer sees colored dots — not letters. The pattern looks like an "A" to human eyes; the computer just sees colored dots.
Copy/Paste Returns Nothing
When you try to copy text from a scanned PDF, the copy operation looks for text objects — which don't exist. Nothing gets copied. Copy fails since there are no text objects to read, only image data is available, the PDF reader can't extract characters, and the result is an empty clipboard.
How OCR Text Extraction Works
What Is OCR?
Optical Character Recognition (OCR) is technology that converts images of text into actual text characters. Modern OCR uses machine learning to recognize patterns that represent letters, numbers, and symbols. The OCR process involves image preprocessing to clean up the image and adjust contrast, layout analysis to identify columns, paragraphs, and tables, character detection to find individual characters, pattern recognition to match shapes to characters, confidence scoring to rate accuracy of recognition, and output generation to produce text with formatting.
How Adobe Free Tool Extracts Text
Step 1: Upload Your Scanned PDF
Upload options are drag and drop from the computer, or click to browse the file system. Supported formats are scanned PDF, image PDF (embedded images), and multi-page PDFs.
Step 2: Automatic OCR Processing
What happens: PDF pages are analyzed, images are extracted from each page, the OCR engine processes each image, text patterns are recognized, characters are assembled into words and sentences, and layout structure is identified. Processing time: small PDFs (under 10 pages) take 10–30 seconds, medium PDFs (10–100 pages) take 1–5 minutes, and large PDFs (100+ pages) take 5–15 minutes.
Step 3: Review Extracted Text
Review options: view text in browser, check accuracy page by page, edit mistakes directly, and download or copy when satisfied. Accuracy expectations: clean printed text is 95–99% accurate, standard documents are 90–95% accurate, degraded documents are 70–85% accurate, and handwritten text is 50–80% accurate (varies widely).
Step 4: Copy Text or Download as TXT
Output formats are copy to clipboard (select text and copy), download as TXT (plain text file), download as DOCX (Word document, some tools), and download searchable PDF (original with text layer added).
Step 5: Use Extracted Text Anywhere
Applications are paste into Word for editing, import into Excel for data entry, add to a notes app, use in other documents, let search engine indexing, and enter into databases.
What Text Gets Extracted
All Readable Text from the Scan
OCR extracts all legible text visible in the document. Extracted content includes body text, headings, captions, footnotes, page numbers, marginal notes if legible, and tables (basic structure). Examples are contracts and agreements, reports and studies, articles and papers, letters and memos, invoices and receipts, and manuals and guides.
Preserved Paragraph Structure
OCR maintains document structure in most cases. Structure preserved includes paragraph breaks, line breaks where significant, columnar layout (basic), and headers and footers. Limitations are that complex multi-column layouts may need manual cleanup, nested structures may collapse, and some formatting nuances are lost.
Table Content (Needs Cleanup)
Tables present a challenge for OCR. What works well: simple tables with clear lines, data in rows and columns, and consistent spacing. What needs cleanup: tables without borders, complex merged cells, tables spanning columns, and hand-drawn tables. After extraction, tables often need manual column alignment, cell content verification, and border reconstruction.
Formatting Detection
OCR identifies basic text formatting. Detected formatting includes bold text (strong patterns), italic text (connected or slanted letters), underlined text, superscript and subscript, and font size variations. Accuracy varies: bold is 85–95% accurate, italic is 70–85% accurate, underline is 60–75% accurate, and complex formatting may not transfer.
Tips for Better Text Extraction from Scans
Use High-Resolution Scans (300 DPI or Higher)
Resolution directly impacts OCR accuracy. Resolution guide: 150 DPI is draft purposes only at 70–80% accuracy, 200 DPI is acceptable for basic OCR at 80–90% accuracy, 300 DPI is recommended for standard documents at 95–99% accuracy, and 400+ DPI is maximum accuracy for small text at 98–99% accuracy. Why resolution matters: higher DPI means more pixels, more detail, better character recognition, small text requires higher resolution, and low resolution causes character merging. Scanning settings: set the scanner to 300 DPI minimum, use 400 DPI for small text or poor quality originals, and use the maximum quality setting rather than "auto."
Good Contrast (Black Text on White Paper)
Contrast affects how clearly text stands out from the background. Good contrast is black text on white paper, dark blue text on white, and clear, clean documents. Poor contrast is gray text on light gray background, faded printing, yellowed paper, and colored paper with colored text. To improve contrast, use the scanner's auto-contrast feature, increase contrast in scanning settings, use threshold adjustment if available, and for old documents, use document enhancement.
Straighten Crooked Pages Before OCR
Skewed pages confuse OCR engines. Deskew matters since OCR expects horizontal text, slanted text reduces accuracy, words may be misread as different letters, and layout analysis becomes unreliable. Deskew before scanning: straighten originals before scanning, use automatic deskew if the scanner has it, and align paper properly in the feeder. Deskew after scanning: many OCR tools include deskew, image editors can rotate, and automatic deskew is standard in most OCR. Acceptable skew: 0–2° has no significant impact, 2–5° has minor accuracy reduction, 5°+ has significant accuracy loss, and 10°+ causes major recognition errors.
Select the Correct Language for Better Accuracy
OCR accuracy improves significantly with correct language selection. Language matters since OCR uses language models, wrong language means wrong character set, hyphenation and word breaks depend on language, and dictionary-based correction relies on language. Setting language: most OCR tools auto-detect language, manual selection improves accuracy, for mixed-language documents process separately, and common languages include English, Spanish, French, German, Chinese, and Japanese. Multi-language documents: identify the dominant language, process the document multiple times if needed, some tools support multiple languages, and accuracy may decrease for minority languages.
Clean, Undamaged Originals
Physical document condition affects scan quality. Best conditions for OCR are clean, unwrinkled paper, no stains or marks, clear printing or handwriting, flat (not curved) pages, and no shadows from binding. Problem conditions are torn or folded documents, water damage, highlighted text (may interfere), notes in margins (adds confusion), and folded or creased pages.
When Text Extraction Needs Help
Very Old or Degraded Documents
Age and condition affect OCR accuracy. Challenges with old documents are faded ink or toner, yellowed or stained paper, foxing (brown age spots), brittle or fragile pages, and non-standard printing. Improving results: use highest resolution (400+ DPI), use document enhancement features, process in sections if possible, accept lower accuracy (70–85%), and manual review and correction is key.
Handwritten Text (OCR Struggles)
Handwriting recognition is significantly harder than print. Challenges are huge variation in writing styles, individual letter variations, connected versus printed handwriting, cultural differences in writing, and legibility issues. Handwriting OCR accuracy: clear, printed handwriting is 60–80%, standard handwriting is 40–60%, poor handwriting is 20–40%, and cursive is often below 30%. Alternatives for handwriting are manual transcription, specialized handwriting OCR tools, human review required, and accepting limitations.
Non-Standard Fonts
Unusual fonts confuse OCR. OCR-friendly fonts are Times New Roman, Arial, Helvetica, Calibri, and standard serif and sans-serif. Problematic fonts are script and cursive fonts, decorative fonts, custom or proprietary fonts, very thin or light fonts, and complex display fonts. What to do: standard fonts work well at 95%+ accuracy, unusual fonts reduce accuracy, expect errors with decorative fonts, and manual correction is needed.
Multiple Languages in One Document
Mixed-language documents are challenging. Challenges are language detection issues, different character sets, switching between alphabets, and non-Latin scripts (Chinese, Arabic, etc.). Best practices are identify the primary language, process documents in sections, use multi-language OCR if available, for rare languages accuracy decreases, and manual review is key.
Using Extracted Text
Copy to Clipboard
The process is select desired text, right-click or tap Copy, and paste into the destination. Tips: select by paragraph for efficiency, use Ctrl+A (Cmd+A) for all text, and check for formatting artifacts in copy.
Export to Different Formats
Common exports are TXT (plain text, universal compatibility), DOCX (Word document, editable), XLSX (Excel, for tabular data), CSV (spreadsheet format for data), and PDF (searchable PDF with text layer).
Search Engine Indexing
After text extraction, search engines can index the content. For web content: convert to HTML, upload with accessible text, and search engines index text not images. For local files: add a text layer to make searchable, use desktop search tools, and file content becomes searchable.
Data Entry and Processing
Extract structured data from documents. Applications are invoice data entry, form field extraction, survey results compilation, database population, and contact information extraction. Best practices: verify extracted data accuracy, use validation rules, spot-check random samples, and build in a review process.
Free vs. Paid OCR Tools
Free Options
Browser-based tools (like our tool) require no software installation, no signup, work on any device, have good accuracy for standard documents, and may have limitations on file size or usage. Free software includes Google Keep (basic OCR), Windows 10+ built-in OCR, Mac built-in OCR (Preview), Tesseract (open source), and ABBYY Screenshot Reader (limited free). Accuracy is 85–95% for standard documents.
Paid Options
Professional OCR software includes Adobe Acrobat Pro, ABBYY FineReader, OmniPage, and Readiris. Accuracy is 95–99% for standard documents with features like batch processing, format preservation, and advanced features.
When to Pay vs. Free
Use free when: occasional, simple documents; standard printed text; no strict accuracy requirements; and budget is limited. Consider paid when: high-volume processing; complex layouts; strict accuracy requirements; need advanced features; and professional document workflows.
Related Tools and Resources
- OCR PDF to Searchable — Create searchable PDF
- Extract Images from PDF — Get images from PDF
- Compress PDF — Reduce file size
- Merge PDF — Combine documents
- Make PDF Searchable — Insight OCR technology
Read More

What Is OCR — Complete Guide to Text Recognition
Read article
Chat with PDF: Ask Questions & Get Answers Instantly (2026 Guide)
Read article



