OCR and accessibility
How to make a scanned PDF searchable with OCR
Understand how OCR turns scanned PDF pages into searchable text, how to improve recognition accuracy, and what to check before using the result.
Why some PDFs cannot be searched
A PDF can look like a normal document while containing only page-sized images. This commonly happens when paper records are scanned or photographed. Because the file contains pixels rather than encoded characters, a PDF reader cannot reliably search, select or copy the visible words.
Optical character recognition, usually shortened to OCR, analyses those pixels and estimates the letters, words and positions they represent. A searchable PDF normally keeps the original page appearance and adds an almost invisible text layer aligned with it.
Prepare the scan before recognition
OCR works best with straight, high-contrast pages and clearly separated characters. Very low resolution, shadows near a book spine, skewed pages, compression artefacts and handwriting all reduce accuracy. If possible, scan ordinary printed documents around 300 dpi and use greyscale or colour when faint marks would disappear in black and white.
- Rotate sideways pages before OCR.
- Crop large borders that do not contain useful content.
- Prefer an original scan over a screenshot of a scan.
- Avoid repeatedly saving the pages as low-quality JPEG images.
- Select only the pages that actually need recognition.
Choose speed or accuracy deliberately
Higher rendering resolution gives the OCR engine more detail, but it also increases memory use and processing time. Fast mode is useful for clean, long documents. Balanced mode suits most printed pages. Accurate mode is better reserved for small type or difficult scans where the additional processing produces a worthwhile improvement.
Browser OCR is computationally demanding because recognition happens on your device. File size alone does not predict duration: a 10 MB document with hundreds of scanned pages can take longer than a 40 MB document with a few high-quality pages.
Review what OCR produced
OCR output should be treated as a draft transcription. Names, dates, decimal points, reference numbers and pairs such as 0/O or 1/I are especially important to check. Search for a few known phrases, copy representative paragraphs, and compare important figures with the visible scan.
Paperuna skips pages that already contain a usable text layer and can process selected page ranges. The downloadable text file helps with review, while the searchable PDF preserves the scanned page as its visible source.
What OCR does not do
OCR does not restore the structure of the original word-processing document. Columns, tables and reading order may still be interpreted imperfectly, and adding a text layer does not automatically make a document fully accessible to screen readers. Proper accessibility may require headings, tags, alternative text and a carefully checked reading order.
Common questions
Does OCR change how the PDF looks?
A searchable PDF can retain the original scanned page as the visible layer while adding recognised text behind it.
Why does OCR take a long time?
Every image-only page must be rendered and analysed. Page count, image detail, recognition quality and device performance all affect processing time.
Try it privately
Put this guide into practice
Paperuna processes supported documents locally in your browser. Your file is not uploaded for processing.
OCR PDFThis guide provides general information, not legal, regulatory or professional advice. Review important outputs and keep an untouched copy of the original document.