Guide

How to rename scanned documents automatically

You tried a renaming tool on a folder of scans and it did nothing, or named everything Untitled. This is why, and what to do instead.

Short answer

A scan is a photograph of a page, not text. Nothing can read it until you either run OCR to add a text layer, or use a tool that reads the picture directly.

Why your renaming tool did nothing

When a document is created digitally, the words are stored as text and a computer can read them. When a document is scanned or photographed, the page is stored as an image. The words are visible to you and invisible to software.

Most renaming tools read the text layer. On a scan there is nothing there, so they find nothing and either skip the file or name it with a blank.

You can check in two seconds. Open the file and try to select a line with your mouse. If you cannot highlight it, it is a scan.

Option 1: run OCR first, then rename

OCR, optical character recognition, looks at the picture and works out what the letters are, then writes a text layer into the file. After that, ordinary renaming tools work.

  1. 1OCRmyPDF is free and open source, runs on your own machine, and adds a text layer without changing how the page looks.
  2. 2Adobe Acrobat Pro has OCR built in under Scan and OCR, if you already have it.
  3. 3Your scanner's own software may have a setting called Searchable PDF. Turning it on is the cheapest fix of all, because it stops the problem happening again.

The best long term answer, because a searchable PDF is more useful forever. Turn it on at the scanner and the problem disappears for all future documents.

Option 2: a tool that reads the picture

Newer tools skip the OCR step and read the image directly, the way a person looking at it would. That handles crooked scans and phone photos better than traditional OCR, which expects a flat, straight page.

What makes scans go wrong, and what to do about it

Crooked pages
Traditional OCR degrades quickly on a tilted page. Straightening before OCR helps a lot, and most scanner software has a deskew option.
Low resolution
Below about 200 dpi, small print stops being readable at all. 300 dpi is the usual advice and it costs you nothing but file size.
Photos rather than scans
A phone photo has perspective distortion, shadows and uneven focus. These are much harder than a flat scan and are where most OCR gives up.
Two pages in one image
Scanning a double page spread as one image confuses almost everything. Scan pages singly.

Common questions

How do I know if a PDF is scanned?
Open it and try to select text with your mouse. If you can highlight a line of words, it has a text layer. If your cursor drags a box over a picture instead, it is a scan.
Is there free OCR software?
Yes. OCRmyPDF is free and open source and runs on your own computer. Tesseract, which it uses underneath, is also free. Many scanners can produce a searchable PDF directly, which avoids the problem entirely.
Why does OCR get numbers wrong?
Usually resolution or angle. Scan at 300 dpi and straighten the page before running OCR. Numbers are harder than words because there is no dictionary to check against, so a misread digit looks perfectly plausible.

Related

Solven builds custom software and automation for Australian businesses. If you have a job like this that no tool quite fits, that is the work we do.

Get a free teardown