← Back to all posts

OCR: pull text out of scanned PDFs and photos

nova-pdf-converterrelease

Nova PDF Converter now includes OCR (image to text). Turn a scanned PDF — or a photo of a document — into text you can select and search.

Text being extracted from a scanned document

What it does

  • Scanned PDFs and images alike: PDF pages are re-rendered as images before recognition, and you can also drop a PNG or JPEG straight in
  • 11 languages: Japanese, English, Chinese (Simplified/Traditional), Korean, Spanish, French, German, Portuguese, Italian and Russian
  • Copy or save the result: put the recognized text on your clipboard in one click, or download it as a text file
  • Multi-page PDFs: the first 50 pages are processed in one go, with per-page progress

The existing "Extract text" tool pulls out the text layer stored inside a PDF. A scanned PDF is really just images, so that tool finds nothing — when that happens, it now points you to OCR.

How to use it

  1. Open the OCR tool
  2. Drop in a PDF or an image
  3. Pick the language and press "Read text"
  4. Copy the result, or save it as a text file

No sign-up and nothing to install. The recognition data is downloaded once, on first use, and cached after that.

About your data

Recognition runs entirely inside your browser. Neither your PDF, nor your image, nor the recognized text is ever sent to a server. Speed depends on your device, but in exchange nothing leaves it.

Every feature is free.

When it comes in handy

  • Quoting a passage from a paper document you only have as a scan
  • Turning a photo of a handout or whiteboard into searchable text
  • Making an old image-only archive searchable again

👉 Try the OCR tool

Novare Orbis offers other free, browser-based tools as well — see the full list.