Why OCR fails on skewed, dark or low-resolution scans (and what to fix first)

2026-09-15
8 min read

The usual OCR story goes: photograph a page with your phone, run it through a recogniser, get back a paragraph where "invoice" has become "irwoice" and every "l" is a "1". The temptation is to blame the engine. In practice the engine — Tesseract, the open-source recogniser behind our OCR tool and behind a large share of commercial products — reads clean, straight, 300-dpi text with an error rate well under one per cent. What it cannot do is compensate for an input that was never legible to begin with. Fixing the input is where almost all of the improvement is.

1. Resolution: the letters have to be tall enough

Recognition works on the shapes of individual characters, and shapes need pixels. The practical rule from Tesseract's own documentation is that lowercase letters should be at least 20 pixels tall; below about 10 pixels, recognition collapses. For ordinary 10–12 point body text, that translates to a scan resolution around 300 dpi: a 10-point "x" is roughly 1.8 mm tall, which at 300 dpi is about 21 pixels. At 150 dpi the same letter is 10 pixels and the engine starts guessing. Going beyond 300 dpi buys very little for text and multiplies processing time.

Phone cameras muddy this because they do not have a "dpi"; what matters is how many of the sensor's pixels land on the page. A 12-megapixel photo of an A4 sheet filling the frame gives roughly 4000 × 3000 pixels over 297 × 210 mm, about 340 pixels per inch — plenty. The same photo taken from a metre away, with the page occupying a quarter of the frame, drops to about 170 dpi and the small print becomes unreadable to the engine even though it still looks fine to you. Fill the frame, or crop tightly before recognising.

2. Skew: a few degrees is enough to break lines

Before it recognises characters, the engine has to find lines of text, and it does so by looking for horizontal bands of ink. A page rotated by 2–3 degrees still looks straight to a human, but across an A4 width a 3-degree tilt shifts the baseline by about 11 mm from one edge to the other — roughly three lines of text. Lines start to overlap the detector's bands, words from adjacent lines get merged, and the output turns into fragments in the wrong order. This is the single most common cause of "the OCR returned garbage" on phone scans.

Tesseract does attempt to estimate and correct small skews internally, but its tolerance is limited and it does nothing for perspective distortion (the page being wider at the bottom than at the top because the camera was tilted). Our scan tool addresses both: it detects the page edges and applies a perspective transform so the page is rectangular before anything else happens. If you are working from an existing photo instead, cropping to the page and straightening it in any image editor achieves the same result. Do this before adjusting anything else; skew correction on a low-contrast image is harder, not easier.

3. Contrast and lighting: shadows read as ink

Internally, recognition is binary: every pixel ends up classified as ink or paper. The engine chooses a threshold (Otsu's method, in Tesseract's case) that separates the two based on the distribution of brightness across the image. That works well when the page is evenly lit. It falls apart when a shadow from your hand or the phone darkens one corner: the threshold that is right for the bright half is wrong for the dark half, so either the shaded text disappears into "paper" or the shadow itself becomes a block of "ink" that swallows the words under it.

The fixes are cheap. Photograph in diffuse light rather than under a single lamp; avoid having your own shadow on the page; and convert the image to greyscale before recognising, which removes colour casts that confuse the threshold. Our scan tool applies a black-and-white threshold that adapts to the page's average brightness for exactly this reason, and the grayscale tool strips colour from an existing PDF. Coloured backgrounds (a form printed on yellow paper, a highlighted paragraph) are a special case: the highlighter is often closer in brightness to the ink than to the paper, and the affected words vanish. Greyscale with a contrast boost usually recovers them; when it does not, there is no software fix, only a better photograph.

4. Margins, noise and the things that are not text

The engine tries to read everything on the page. A hole-punch, a coffee ring, a staple shadow, the edge of the desk behind the sheet: each of these gets segmented and attempted as characters, and each produces a few stray symbols in the output. Worse, a dark border from the table around a page edge can be interpreted as a very large glyph or as a table rule, which shifts the layout analysis for the whole page. Cropping tightly to the page — which the scan tool does automatically — eliminates most of it. A white margin of a few millimetres around the text also helps, because the line detector expects text to be surrounded by paper.

Handwriting deserves a separate mention: Tesseract is a printed-text engine. It will produce something from neat block capitals and nothing usable from cursive. That is a limitation of the model, not of the scan, and no amount of preprocessing changes it.

The order to fix things in

  1. Retake the photo if the page does not fill the frame or the lighting is uneven. Everything downstream is easier with a good source.
  2. Crop to the page and straighten (or let the scan tool do it). Skew and perspective are the highest-leverage fix.
  3. Convert to greyscale and boost contrast. Check that small text is still solid black, not grey and broken.
  4. Confirm the text height: if lowercase letters are under 20 pixels in the processed image, upscale it ×2 before recognising. Upscaling adds no information but it does bring glyphs into the size range the engine was trained on, which measurably helps.
  5. Pick the right language model. Tesseract's language packs include the dictionary it uses to resolve ambiguous glyphs; running Spanish text through the English model turns every "ñ" and accented vowel into a guess.

What "good" looks like, and how to check

On a clean 300-dpi scan of a typed page, expect the output to need a handful of corrections per page, mostly in unusual proper nouns and in the "0/O" and "1/l/I" families where the shapes genuinely coincide in many fonts. On a phone photo that has been cropped, straightened and greyscaled, expect somewhat more, concentrated in the smallest text. If the output has errors in every other word, one of the four inputs above is still wrong; go back to the photo rather than trying to correct the text by hand.

One thing that helps assess a result quickly: our OCR tool returns the recognised text as plain text you can copy or download, next to the page it came from. Pick a line of small print you can read on the page and find it in the output. If it is there with at most a character or two wrong, the input was good enough; if whole words are missing or replaced, the recognition never saw them clearly and the page needs a better source, not a better spell-checker. The original PDF is never altered by OCR — you lose nothing by trying.

Photograph the page with the scan tool, which crops, straightens and binarises it before you recognise the text.

Scan a document