OCR a scan
The words read off a scanned page and laid invisibly back over it, so the document can be searched, selected and copied. It happens in your browser, and the scan is never uploaded.
Drop your PDF here
They stay on your computer. Nothing is sent to a server.
How to make a scanned PDF searchable
A scan is a photograph of a page. To your computer it holds no words at all - which is why Ctrl+F finds nothing in it, why you cannot copy a line out of it, and why a screen reader has nothing to read. OCR is the job of looking at the pixels and working out which letters they are.
- Drop the scanned PDF on the box above.
- Leave the tick in place if parts of the document are already typed, so those pages are not read again.
- Press Read the words and save. The engine arrives first, once, and then the pages are read one after another.
- Wait. This is the slow tool on the site: a few seconds a page on a fast machine, longer on a phone.
- The document is saved when the last page is done. Open it and press Ctrl+F to see the difference.
What comes out, and what it looks like
The same document, looking exactly as it did. No page is retyped and nothing is redrawn - the scan is still the scan. What is added is a layer of invisible text lying precisely over the picture of the words, so that selecting a line selects the words that are drawn there, searching finds them, and copying gives you real text.
That is what everybody means by a searchable PDF, and it is what the engine here produces itself. The letters are placed by the part of the engine that knows where in the picture it found them, rather than by us guessing afterwards, which is the difference between a text layer that lines up and one that is half a line out.
There is also a button to save the words on their own as a plain text file, or copy all of them, for when the text is what you were after rather than the document.
Why this is not uploaded, when every other free one is
OCR is the tool the free sites most want you to upload for. It is heavy work, so doing it on a server is easier for them - and the documents people need read back are payslips, passports, medical letters, contracts and old family papers. That is a remarkable collection of things to hand to a company whose business model you have not read.
The engine here is Tesseract, which is the open-source engine most of those services are built on top of anyway, compiled to run in a browser. It is served from this site's own address rather than somebody's content network, so no request leaves the page while you are reading a document, and once it has been fetched once the tool works with the internet switched off.
The cost of that honesty is the download - about seven megabytes of engine and language data, once - and the waiting. Those are real costs and they are the reason this tool is slower than the paid ones. The document staying on your machine is what you get for them.
What it reads well, and what it does not
Printed English on a straight, evenly lit page is read very well. A crooked page, a dim photograph, a low-resolution scan, a page of dense small print or a fax are all read worse, and you are shown how sure the engine was so you know which one you got.
Handwriting is not read well by this kind of engine, whatever any website claims. Nor are heavily stylised fonts, text printed over pictures, or a language the engine has not been given - it has English here.
The best thing you can do for a poor result is not a setting on this page. It is scanning the page again, straighter and brighter. The Scan to PDF tool on this site finds the page, crops it and flattens it, which is exactly the shape this engine reads best.
Good to know
- Is my scan uploaded?
- No. The engine runs inside your browser and the document never leaves your computer. Even the engine itself is served from this site rather than from somebody else, so no third party learns that you used it.
- Why is it so slow?
- Because it is doing the work on your machine instead of a server farm. Expect a few seconds a page on a laptop, longer on a phone, plus a one-off download of about seven megabytes the first time. A long book is a walk away from the screen.
- Does it change how the document looks?
- No. The picture of the page is untouched. The text is added underneath it, invisible, lined up with the words you can see.
- Can it read handwriting?
- Not reliably. This kind of engine is built for printed text. Neat block capitals sometimes come through; ordinary handwriting does not, here or in most paid tools.
- Which languages does it read?
- English. Other languages need their own data file, which is a few megabytes each. If you need one, say so and it can be added.
- What if the confidence is low?
- It means the page was hard to read - crooked, dim, blurred or very small print. Rescanning it straighter and brighter helps more than anything on this page. The Scan to PDF tool crops and flattens a photographed page, which is the shape this engine reads best.