OCR and accessibility
How to check whether a PDF has selectable text
Tell the difference between a text PDF and a scanned image, understand why search fails, and decide when OCR is needed.
Try selection, search and copy
Open the PDF and drag across a sentence. If individual characters or words become selected, the page probably contains a text layer. Search for a distinctive phrase and paste copied text into a plain-text editor to check whether the characters and reading order are sensible.
A page can contain both text and images. Headers may be selectable while a scanned letter in the middle is not, so test representative areas on several pages.
Why visible words may only be pixels
Scanners and cameras usually create images. A PDF can wrap those images into pages without recognising any characters. The viewer can display the words because people can read the pixels, but software has no encoded text to search or copy.
Some scans have already been through OCR but contain a poorly aligned or inaccurate hidden text layer. Successful selection therefore does not guarantee correct text.
Use OCR when appropriate
OCR analyses each page image and estimates its words and positions. The resulting searchable PDF can retain the original scan visually while adding nearly invisible text. Straight, high-resolution, high-contrast pages produce better results.
Rotate and crop pages before recognition. Process only the pages that need it, particularly for long documents, because browser OCR uses significant memory and processor time.
Check the recognised result
Search for known phrases and copy several passages after OCR. Compare names, dates, totals, reference numbers and similar characters such as zero and capital O with the visible source.
Searchable is not the same as accessible. Screen-reader accessibility may also require document tags, headings, alternative text and a verified reading order.
Hands-on practice
Test text selection
Use fictional data first, then repeat the workflow with your own document once you understand the result.
Common questions
Why can I select some PDF pages but not others?
The PDF may combine digitally created pages with scanned image pages, or only some pages may already have OCR text.
Does OCR guarantee accurate text?
No. Recognition quality depends on the scan and typography, so important text must be reviewed.
Try it privately
Put this guide into practice
Paperuna processes supported documents locally in your browser. Your file is not uploaded for processing.
Make a PDF searchable with OCRThis guide provides general information, not legal, regulatory or professional advice. Review important outputs and keep an untouched copy of the original document.