Quick answer
Check whether the file is image-only, then use OCR and searchable PDF in the scan workbench. Download searchable.pdf and verify searches, copied sentences and page completeness. Finding one word establishes that some text is searchable; it does not establish that the entire document was recognized correctly.
Separate missing text from reader problems
Try selecting a normal paragraph and searching for a clearly visible word. Repeat in another reader so that a selection mode is not mistaken for a file defect. If selectable text exists but copies as nonsense, ask the sender for a properly exported original before rebuilding every page.
Protection, encryption and usage rights can also affect copying. Work with an unencrypted copy you are authorized to process. OCR is not a way to remove access restrictions. Keep the original and use a distinct output name, especially for documents with signatures, attachments or important links.
Start with representative pages
Try a page containing ordinary text, numbers and small print. Check completeness, orientation, glare and occlusion. Rotation corrects an angle; it does not automatically repair every perspective distortion in a photographed sheet.
Tesseract’s output-quality guidance describes factors such as skew, noise and binarization. The site’s manual rotation, threshold and brightness-normalization options can help you prepare an image, but cannot reconstruct characters that the source does not contain.
For separate photographs, first follow the images-to-PDF completeness and page-order checks. Successful conversion or merging does not establish that OCR has occurred.
Generate a new text layer in the scan workbench
- 1Open the scan enhancement and OCR workbench and select a PDF or images. English and Chinese recognition resources need to load on first use. Complete a small sample before a full document.
- 2Adjust manual rotation if necessary. A zero black-and-white threshold retains grayscale processing. Compare small letters, punctuation and pale content when changing threshold or brightness normalization; a whiter background is not inherently a better result.
- 3Select OCR and searchable PDF. Enhance scans alone processes the images and does not complete this text-layer task.
- 4Wait for completion and download searchable.pdf from the output files. Save ocr.txt separately if you need plain text. Its filename and page separators help you trace content. Keep the source and enhanced images for comparison.
This entry recognizes English and simplified Chinese locally, page by page. Pixel scaling and device memory affect processing. It is not a lossless editor for an existing PDF.
Check text, pages and delivery requirements
- Search for words on several pages and check their actual locations. Include Chinese, English and numbers when relevant, rather than testing only a cover title.
- Copy a sentence containing punctuation and numbers into a plain-text editor. Compare digits, decimal points, dates, minus signs and line order with the image.
- Inspect page count, orientation, edges and readability. Columns, headers, footnotes and tables need separate reading-order checks. Selectable text does not establish correct structure.
- Check filename, size and recipient requirements. OCR does not guarantee compression, accessibility conformance or a particular archival standard.
Keep PDF and TXT corrections separate
The workbench rebuilds an OCR PDF from page images. Do not assume original links, signatures, attachments or document structure survive. Keep usable native-text PDFs whenever possible; this workflow primarily serves scans.
Editing TXT or the table-proofreading area does not update searchable.pdf’s text layer. For a critical text-layer error, use a tool that supports correction, or improve the source and recognize again. Recheck the final PDF.
If a single image only needs copyable text, use the image OCR and proofreading workflow. Handwriting, complex layouts and traditional Chinese may be unreliable. Obtain a clearer original instead of repeatedly recognizing a blurred source.
Handle loading failures and large documents
Recognition runs in your browser, but program and language resources still need to load. Network or cache issues can prevent startup. Check a small file first. For large PDFs, save manageable page-range copies, process them separately, then merge and check page completeness.
There is no fixed page count that works on every device. Keep the page open and avoid starting several jobs repeatedly. Save downloaded results, and follow your organization’s rules for handling sensitive material on the device.
Common questions
Why are some words still unsearchable?
Recognition can misread small, tilted or occluded text. Compare the copied text with the image, improve the source and try again. One successful search does not verify the entire document.
Will editing ocr.txt correct the PDF?
No. TXT is a separate output, and editing it does not automatically change the PDF text layer. Verify the file you will actually deliver.
Should a PDF with selectable text be OCRed again?
Usually keep the usable native-text version. Rebuilding scans can alter appearance, size and document structure. Process a copy only when you have a specific image-recognition need.
Further reading & sources
Put the guide into practice
Open the relevant tool, preview the result and check your exported file.


