Tesseract OCR app icon

Office

Tesseract OCR

Open-source OCR engine for Windows with a command-line interface, 100-plus language models and scriptable batch text extraction.

WindowsmacOSLinuxApache-2.0

The OCR engine that other OCR programs quietly rely on

Tesseract is an open-source optical character recognition engine maintained by the Tesseract OCR project and released under the Apache License 2.0. It is not a polished desktop application; it is an engine with a command-line interface, and it is the recognition core inside a long list of other programs, including gImageReader, several document managers and many scanning pipelines. The Windows build published through the project releases installs the engine, a set of command-line tools and the training data for a first language, and it can recognise printed text in more than a hundred languages once the matching data files are installed.

What it does

Given an image or a multi-page TIFF, the engine returns plain text, hOCR, ALTO XML, TSV with word coordinates, or a PDF with a text layer laid over the original page image. Command-line options control the page segmentation mode, the recognition engine, the character whitelist, the dictionary behaviour and the output formats, and a configuration file can capture a repeated set of those choices. Training tools ship alongside the engine for users who want to fine-tune recognition on a particular typeface or document family.

The workflow

Prepare images at around 300 dots per inch with straight edges and good contrast, then run the engine over a folder with a small batch script that loops through the files and writes one output file per input. Start with the automatic page segmentation mode and change it when documents are single columns, sparse text or single lines. Review a sample of the output to tune resolution and preprocessing before running a large job, and keep the intermediate images so a failed batch can be re-run without rescanning.

Practical settings and limits

The engine has no graphical interface, no preview and no interactive correction, so it is normally paired with a front end or a script. Recognition accuracy collapses on skewed pages, low resolution, decorative typefaces and handwriting, and there is no layout analysis beyond paragraph and block detection, which means tables and multi-column pages need post-processing. Language data files are separate downloads and the total size grows quickly when many languages are installed. The Apache licence is permissive, so the engine can be embedded in commercial products.

Who should choose something else

Users who want a window, a preview and a save button should install a front end such as gImageReader rather than driving the engine directly. Organisations that need table reconstruction, form recognition or handwriting support should evaluate a commercial engine. Tesseract is the right foundation when recognition has to run locally, at scale, inside your own scripts, with no per-page cost and no data leaving the machine.

Best for
Local, scripted OCR at scale inside your own pipelines with no per-page cost.
Good to know
Apache-2.0, so it may be embedded commercially. It is a command-line engine with no graphical interface.

How to get started

  1. Open the Tesseract releases page and download the 64-bit Windows setup executable.
  2. Run the installer, accept the Apache 2.0 licence and keep the default component selection for the engine and language data.
  3. Add the installation folder to the system PATH so the tesseract command is available in any terminal window you open afterwards.
  4. Run tesseract --list-langs to confirm the installed language data and add more languages with the same installer when you need them.
  5. Recognise a test page with tesseract input.png output pdf and inspect the result.
  6. Wrap the command in a small batch file when a folder has to be processed nightly or on a schedule.

Questions & answers

Is Tesseract free for commercial products?

Yes. It is released under the Apache License 2.0, which permits commercial use and embedding without a royalty.

Can Tesseract make a searchable PDF from a scan?

Yes. Passing pdf as the output format writes the original page images with an invisible recognised text layer on top, so the file stays visually identical.

Does it need an internet connection?

No. Recognition is entirely local; the only optional download is the language data for additional scripts and alphabets.

More in office

View category