OCR reads the shapes of letters and turns them back into text — which means how clearly those shapes were captured in the first place puts a hard ceiling on how accurate the result can be. No OCR engine recovers detail that a bad scan never captured. The settings below are the ones that actually move accuracy.
Resolution: 300 DPI
DPI (dots per inch) is the single biggest factor. 300 DPI is the standard recommendation for OCR on normal printed text — enough resolution for the engine to distinguish similar characters cleanly (“rn” from “m”, “cl” from “d”) without producing a file that’s needlessly huge.
- Below 200 DPI, accuracy drops noticeably, especially on small fonts, footnotes, and anything condensed.
- Above 300–400 DPI gives diminishing returns for normal text and mainly costs you file size, though very small or fine print (a receipt, a legal document’s fine print) can benefit from 400–600 DPI.
If you’re scanning with a phone camera rather than a flatbed scanner, DPI as a setting doesn’t apply the same way — get close enough that text fills a reasonable portion of the frame, and make sure focus is sharp, since blur has the same effect as low DPI.
Color mode: grayscale, not black-and-white or full color
Grayscale is the sweet spot for OCR on typical documents. Pure black-and-white (1-bit) thresholding can lose faint text and turn light pen marks or watermarks into noise, while full color captures information OCR doesn’t need and produces a much larger file for no accuracy benefit on plain text.
The exception: if the document has colored highlighting, stamps, or diagrams that matter visually, scan in color and accept the larger file — the accuracy trade-off for pure text is small, and you keep the visual information.
Straightness matters more than people expect
A page that’s noticeably skewed — scanned a few degrees off-square — measurably hurts recognition, because OCR engines expect text to run in roughly straight horizontal lines. A flatbed scanner with the paper aligned to the edge guide avoids this almost entirely; a phone camera needs a squarer, more deliberate shot than it feels like you need. Scan to PDF corrects this automatically when scanning with your phone, cropping and straightening each page as you capture it.
Lighting and contrast
Even, flat lighting beats bright lighting with shadows. A single overhead light or a window behind the camera creates a gradient of brightness across the page that can push OCR to misread text in the darker corners even when a human reading it wouldn’t notice a problem. If scanning with a phone, an overcast day near a window or diffuse indoor lighting beats direct sun or a single desk lamp.
File format: keep it lossless until OCR runs
Save or export the scan as PNG, TIFF, or an uncompressed PDF rather than a heavily compressed JPEG where possible. JPEG’s compression introduces small blocky artifacts around edges — including letter edges — and while modest JPEG compression is usually fine, an aggressively compressed one measurably degrades OCR accuracy on fine text.
Putting it together
Once you’ve got a clean scan — 300 DPI, grayscale, straight, evenly lit — run it through OCR PDF to add a searchable text layer, or start with Scan to PDF if you’re capturing with a phone, since it handles cropping and straightening as part of the same pass.
Related reading
- How to Make a Scanned PDF Searchable — what OCR does once your scan is ready
- Why You Can’t Select or Search Text in a PDF — the underlying reason scans aren’t searchable to begin with
Ready? Try the free OCR PDF tool now — no sign-up, no watermark, and your file is never stored.