gImageReader app icon

Office

gImageReader

Free graphical front end for the Tesseract OCR engine, with PDF and image import, page previews, editable text and scan acquisition.

WindowsLinuxGPL-3.0-or-later

Desktop OCR with Tesseract underneath and a usable face on top

gImageReader is a free, open-source desktop application that puts a graphical interface on the Tesseract optical character recognition engine. Tesseract is excellent at recognising text and awkward to drive from a command line; gImageReader solves that by handling images, multi-page PDFs, scanned pages and clipboard captures, showing a live preview of the recognition boxes, and producing editable text, searchable PDF output or plain text files. It runs on Windows and Linux, and it exists in both a Qt and a GTK flavour depending on the build supplied by the project.

What it does

The window is split between a page list and a preview canvas. Pages can be imported from image files, dragged in from another application, scanned through a TWAIN or SANE device, or pulled from a multi-page PDF, which is rendered page by page. Recognition runs over the whole document or only the current selection, with a language chosen from the Tesseract data files that are installed. Results appear in an editable text pane beside the image, so the recognised characters can be corrected while the source is visible. Output can be saved as text, as a PDF with an invisible text layer, or as hOCR and similar formats that keep layout information.

The workflow

Scan or import the pages, check that they are level and that the correct language is selected, then run recognition over the document. Work through the text pane correcting the errors that matter - proper nouns, numbers, punctuation - because recognition quality depends heavily on scan resolution. Export a searchable PDF when the document needs to stay visually identical, or plain text when it is going into another program. For mixed-language documents, run the pages in batches with the appropriate language model.

Practical settings and limits

Recognition quality is dominated by input quality: 300 dots per inch, straight pages, good contrast and no skew produce dramatically better results than a phone photo under poor light. Additional language packs must be installed separately in the Tesseract data folder, and complex layouts with columns, tables or marginalia need manual reordering of the extracted text. The interface is functional rather than polished, the first run may require pointing the program at the Tesseract executable, and very large PDFs are slow because every page is rasterised before recognition.

Who should choose something else

Anyone who needs high-volume server-side OCR, layout analysis with table reconstruction or handwriting recognition should use a commercial engine. Users who only need to copy a paragraph from a screenshot will find a built-in operating-system tool faster. gImageReader is the right answer for people who want free, offline, local OCR over scanned documents with a real interface instead of command-line flags.

Best for
Turning scanned documents and PDFs into editable or searchable text with a desktop interface.
Good to know
GPLv3. The Tesseract engine and extra language data files must be installed separately.

How to get started

  1. Open the gImageReader releases page and download the qt6 x86_64 installer for the latest release.
  2. Run the installer and accept the GPLv3 licence.
  3. Install the Tesseract OCR engine for Windows so the engine executable and the language data files are present on the machine.
  4. Start gImageReader, open Settings, point it at the Tesseract executable and tick the recognition languages you need for your documents.
  5. Import a page or a PDF, run recognition and save the result as plain text, as hOCR or as a searchable PDF.
  6. Add a second language to the Tesseract data folder and re-run a sample page to compare the result.

Questions & answers

Does gImageReader include the OCR engine, or must I install it?

The front end calls Tesseract, so the engine and the language data files for your languages must be installed separately.

Can gImageReader produce a searchable PDF?

Yes. It can export the recognised text as an invisible layer over the original page images, which makes scanned documents searchable and copyable in any reader.

Does it work offline, and is my data uploaded?

Yes. Recognition runs entirely on the local machine; nothing is uploaded, no account is needed and the program works on a disconnected network.

More in office

View category