The extracted text is wrong
OCR (optical character recognition) reads text from images the way a person squints at a blurry photo — usually right, sometimes not. When extraction comes back with errors or missing fields, here's what's going on and how to get a better result.
Why errors happen
Accuracy depends almost entirely on image quality. The most common culprits: blurry or shaky photos, receipts photographed at an angle, shadows or glare across the text, crumpled thermal paper, tiny fonts, and low-contrast printing (gray text on white). Handwriting is harder than print, and decorative or unusual fonts confuse the reader too.
How to get a cleaner extraction
- Retake the photo. Lay the document flat, fill the frame with it, and shoot straight down in bright, even light. No flash — it creates hotspots. This single step fixes most bad results.
- Flatten the paper. Smooth out curls and creases, especially on receipts. A second sheet of paper on top for a minute helps.
- Crop tightly. If the tool lets you, crop out the table, hands, and background before processing. Less noise means fewer misreads.
- Scan instead of photographing when accuracy matters (tax documents, contracts). A $0 phone scanner app beats the best camera technique.
What to do with the result you got
Always skim the output before using it — check totals, dates, and names especially. CSV and Excel outputs open in any spreadsheet app, where fixing a misread cell takes seconds. For W-2s and W-9s, verify every number against the original; masked SSNs/TINs are shown partially by design.
Good to know
- Stylized logos and decorative headers are the most commonly misread elements — and usually the least important ones.
- Tables with no visible gridlines (common in invoices) are harder to parse than ruled tables. If columns merge, that's why.
- Handwriting recognition works best with neat, separated letters. Connected cursive is its weak spot.
- Reprocessing the same image gives the same result — the fix is a better image, not a second run.
- If a clean, straight scan still extracts badly, send us the file through the report form. Real examples are how the models improve.
Last updated: October 4, 2026
Still stuck? If the text is still wrong after rescanning at 300 DPI or higher, report a problem and attach the original file if you can.