OCRmyPDF

Input—per 1M tokens
Output—per 1M tokens
Context—tokens
WeightsClosed

About

OCRmyPDF is ranked #2 of 40 in OCR software on Inferse. It runs on API, Linux, macOS, Self-hosted, Windows. There is a free plan.

Compared on OCR software

Free plan
Yesocrmypdf.readthedocs.io
Handwriting OCR
Noocrmypdf.readthedocs.io

Facts

Free plan
Yesocrmypdf.readthedocs.io · 23 Sept 2026
Searchable PDF
Yesocrmypdf.readthedocs.io · 23 Sept 2026
Handwriting OCR
Noocrmypdf.readthedocs.io · 23 Sept 2026
Primary platform
desktopocrmypdf.readthedocs.io · 23 Sept 2026
Supported inputs
pdfocrmypdf.readthedocs.io · 23 Sept 2026
Purpose
OCRmyPDF adds a searchable text layer to scanned PDF files while preserving the original PDF as much as possible.ocrmypdf.readthedocs.io · 2 Oct 2026
OCR engine
It uses Tesseract to recognize text in PDF page images.ocrmypdf.readthedocs.io · 2 Oct 2026
PDF/A
By default, OCRmyPDF generates PDF/A-2b archival PDFs, and users can select regular PDF output instead.ocrmypdf.readthedocs.io · 2 Oct 2026
Image processing
It offers image processing options such as deskew to improve visual quality and OCR accuracy.ocrmypdf.readthedocs.io · 2 Oct 2026
Existing text
Its processing modes can error on existing text, skip such pages, redo OCR, or force OCR across all pages.ocrmypdf.readthedocs.io · 2 Oct 2026
API and plugins
OCRmyPDF can be used as a Python library and supports plugins that customize processing steps.ocrmypdf.readthedocs.io · 2 Oct 2026
Installations
The documentation provides installation methods for Linux, macOS, Windows, FreeBSD, and Docker.ocrmypdf.readthedocs.io · 2 Oct 2026
Integrations
The documentation identifies Paperless-ngx and Nextcloud OCR as third-party integrations that use OCRmyPDF.ocrmypdf.readthedocs.io · 2 Oct 2026
Security
The project advises using OCRmyPDF only with PDFs users trust and says its Docker web service example has no security measures and is not intended for public internet deployment.ocrmypdf.readthedocs.io · 2 Oct 2026
OCR accuracy
The documentation notes that OCR accuracy may trail commercial solutions, handwriting is not recognized, and poor scans can produce poor results.ocrmypdf.readthedocs.io · 2 Oct 2026
Language support
Results may be poor when a document contains languages not specified in the language argument.ocrmypdf.readthedocs.io · 2 Oct 2026
Page time limit
By default, OCRmyPDF allows Tesseract three minutes per page and can skip images above a configured megapixel threshold.ocrmypdf.readthedocs.io · 2 Oct 2026
Commercial use and dependency
The documentation says users should comply with the project and dependency licenses and notes that Ghostscript, which OCRmyPDF requires in some workflows, is AGPLv3 licensed.ocrmypdf.readthedocs.io · 2 Oct 2026
Maintainer
The project metadata names James R. Barlow as an author.github.com · 2 Oct 2026

Best OCRmyPDF alternatives

See all 12