TesseractOCRParser Configuration

Configuration options for TesseractOCRParser, which runs the tesseract command-line program in a separate process. Component name: tesseract-ocr-parser.

Selecting the OCR engine (4.1.0+)

Name the engine in the top-level text-recognizers list; every default parser stays loaded, and image-parser and pdf-parser call it on embedded images and rendered pages:

{
  "engines": { "tesseract": { "tesseract-ocr-parser": { "language": "eng" } } },
  "text-recognizers": [ { "engine": "tesseract" } ]
}

language takes tesseract’s own names, joined with +: three-letter codes such as eng or chi_tra_vert, and script models such as Latin or Japanese_vert. A script model is written script/Latin when it lives in tessdata’s script/ subdirectory and Latin when it sits at the tessdata top level, which is where the Debian and Ubuntu tesseract-ocr-script-* packages put it; tesseract --list-langs prints the spelling that resolves. A script model reads every language of that script without a per-document choice, at roughly 1.6x the time of eng per page.

With no text-recognizers at all, Tesseract is still the OCR engine whenever the tesseract binary is found: the loader finds it among the loaded parsers because it implements TextRecognizer and advertises the image types it reads, and logs the choice at startup. When another classpath engine such as Tess4J claims the same types, one is used and a WARN at startup names the collision and the winner; an engine named under parsers (a VLM parser, for instance) takes precedence over Tesseract, with the deprecation WARN that shape draws since 4.1.0. To turn OCR off everywhere, set "text-recognizers": []. See Configuration for how the list behaves, and Wiring in Your Own OCR Engine to plug in an engine Tika does not ship.

Basic Configuration

{
  // Tesseract enriches every image Tika's parsers meet (embedded images, rendered PDF
  // pages) with OCR text. Configured once under "engines", named in "text-recognizers".
  // Every default parser stays loaded: "parsers" is not needed.
  "engines": {
    "tesseract": {
      "tesseract-ocr-parser": {
        "language": "eng",
        "timeoutMillis": 120000
      }
    }
  },
  "text-recognizers": [
    { "engine": "tesseract" }
  ]
}

Full Configuration

Every option with its default value and an inline comment describing it. Named under text-recognizers, the engine leaves every default parser loaded; the same options apply to an entry under parsers, the 4.0 shape, which is deprecated since 4.1.0 and unsupported in 4.2.0. ImageMagick is optional; it is used only when enableImagePreprocessing or applyRotation is true.

{
  // Named under "text-recognizers", Tesseract is invoked on every image Tika's parsers
  // meet; the default parsers stay loaded. Windows paths in JSON need forward slashes or
  // escaped backslashes.
  "text-recognizers": [
    {
      "tesseract-ocr-parser": {
        // Calculate skew and rotate (via ImageMagick) before OCR. Needs ImageMagick.
        "applyRotation": false,
        // Colorspace of the preprocessed image (preprocessing only).
        "colorspace": "gray",
        // Resolution (dpi) of the preprocessed image.
        "density": 300,
        // Bits per color sample in the preprocessed image.
        "depth": 4,
        // Run ImageMagick preprocessing (density/depth/colorspace/filter/resize) before OCR.
        "enableImagePreprocessing": false,
        // ImageMagick resize filter.
        "filter": "triangle",
        // Directory holding the ImageMagick program (empty = on PATH). ImageMagick is
        // OPTIONAL -- used only when enableImagePreprocessing or applyRotation is true.
        "imageMagickPath": "",
        // Write OCR output from embedded images inline into the parent document.
        "inlineContent": false,
        // Tesseract language(s); join multiple with '+', e.g. "eng+fra".
        "language": "eng",
        // Skip OCR for files larger than this many bytes.
        "maxFileSizeToOcr": 2147483647,
        // Skip OCR for files smaller than this many bytes.
        "minFileSizeToOcr": 0,
        // Additional raw Tesseract config variables (key-value).
        "otherTesseractConfig": {
          "preserve_interword_spaces": "1",
          "textord_initialx_ile": "0.75",
          "textord_noise_hfract": "0.15625"
        },
        // Tesseract output format.
        // Options: TXT, HOCR
        "outputType": "TXT",
        // Inserted between OCR'd pages (empty overrides Tesseract 4's form-feed).
        "pageSeparator": "",
        // Page segmentation mode (0-13); 1 = auto with orientation/script detection.
        "pageSegMode": "1",
        // Load Tesseract language data at startup instead of on first use.
        "preloadLangs": false,
        // Preserve interword spacing in the output.
        "preserveInterwordSpacing": false,
        // Scale percent (100-900) applied during preprocessing.
        "resize": 200,
        // Runtime kill-switch to disable OCR.
        "skipOcr": false,
        // Directory with tessdata language files (empty = Tesseract default).
        "tessdataPath": "",
        // Directory with the tesseract binary (empty = on PATH).
        "tesseractPath": "",
        // Max seconds to wait for the OCR process.
        "timeoutMillis": 120000
      }
    }
  ]
}

Arbitrary Tesseract variables go in otherTesseractConfig as a name-to-value map, as in the example above. The map is operator configuration only (tika-config.json, including its presets): a per-request tesseract-ocr-parser block that carries it is rejected, because variables such as debug_file name files the tesseract binary opens. Converting a 3.x config? See Migrating to 4.x — the automatic converter rewrites the old otherTesseractSettings list for you.