The three biggest scan problems and how to fix them
Skew is the most common issue. Even a 2-degree tilt causes the OCR engine to misread character boundaries, turning 'rn' into 'm' and 'cl' into 'd'. Most scanner software has a deskew option. Use it. Dark borders from photocopy edges create noise that OCR tries to interpret as text. Crop them out or use a border-removal preset if your scanner has one. Resolution below 300 DPI is the silent killer — 200 DPI might look readable to you, but the OCR sees mush. Always scan at 300 DPI minimum for text documents.
When to convert to grayscale first
Scanned documents with yellowed paper or colored backgrounds confuse OCR engines that are tuned for black text on white. Convert to grayscale and bump the contrast before conversion. Most image viewers have a one-click auto-contrast function. You are not losing useful information — you are removing noise that the OCR would otherwise try to read as characters.