Guide
How to rename scanned documents automatically
You tried a renaming tool on a folder of scans and it did nothing, or named everything Untitled. This is why, and what to do instead.
Short answer
A scan is a photograph of a page, not text. Nothing can read it until you either run OCR to add a text layer, or use a tool that reads the picture directly.
Why your renaming tool did nothing
When a document is created digitally, the words are stored as text and a computer can read them. When a document is scanned or photographed, the page is stored as an image. The words are visible to you and invisible to software.
Most renaming tools read the text layer. On a scan there is nothing there, so they find nothing and either skip the file or name it with a blank.
You can check in two seconds. Open the file and try to select a line with your mouse. If you cannot highlight it, it is a scan.
Option 1: run OCR first, then rename
OCR, optical character recognition, looks at the picture and works out what the letters are, then writes a text layer into the file. After that, ordinary renaming tools work.
- 1OCRmyPDF is free and open source, runs on your own machine, and adds a text layer without changing how the page looks.
- 2Adobe Acrobat Pro has OCR built in under Scan and OCR, if you already have it.
- 3Your scanner's own software may have a setting called Searchable PDF. Turning it on is the cheapest fix of all, because it stops the problem happening again.
The best long term answer, because a searchable PDF is more useful forever. Turn it on at the scanner and the problem disappears for all future documents.
Option 2: a tool that reads the picture
Newer tools skip the OCR step and read the image directly, the way a person looking at it would. That handles crooked scans and phone photos better than traditional OCR, which expects a flat, straight page.
What makes scans go wrong, and what to do about it
- Crooked pages
- Traditional OCR degrades quickly on a tilted page. Straightening before OCR helps a lot, and most scanner software has a deskew option.
- Low resolution
- Below about 200 dpi, small print stops being readable at all. 300 dpi is the usual advice and it costs you nothing but file size.
- Photos rather than scans
- A phone photo has perspective distortion, shadows and uneven focus. These are much harder than a flat scan and are where most OCR gives up.
- Two pages in one image
- Scanning a double page spread as one image confuses almost everything. Scan pages singly.
Common questions
- How do I know if a PDF is scanned?
- Open it and try to select text with your mouse. If you can highlight a line of words, it has a text layer. If your cursor drags a box over a picture instead, it is a scan.
- Is there free OCR software?
- Yes. OCRmyPDF is free and open source and runs on your own computer. Tesseract, which it uses underneath, is also free. Many scanners can produce a searchable PDF directly, which avoids the problem entirely.
- Why does OCR get numbers wrong?
- Usually resolution or angle. Scan at 300 dpi and straighten the page before running OCR. Numbers are harder than words because there is no dictionary to check against, so a misread digit looks perfectly plausible.
Related
Solven builds custom software and automation for Australian businesses. If you have a job like this that no tool quite fits, that is the work we do.
Get a free teardown