Package org.apache.tika.parser.enricher


package org.apache.tika.parser.enricher
Engines a container parser invokes on bytes it has already parsed: text recognizers (OCR, transcription) and annotators. Experimental in 4.1: the contract is one image per call, written into the caller's body as it runs; 4.2 batches recognition per document. The intent is additive: an engine that takes one image at a time keeps working unchanged, and one whose service batches opts in by declaring a batch size and implementing the list call 4.2 adds. Until then these interfaces may change in a minor release without a deprecation cycle. The "engines" and "text-recognizers" configuration that names an engine is stable.
  • Class
    Description
    Media-type-keyed registry of content enrichers: ordinary Parsers that a container parser invokes on bytes it has already parsed (OCR text for an image or a rendered PDF page), rather than being dispatched to by the composite parser.
    A Parser a container parser invokes on bytes it has already parsed (OCR on an image, an embedding of a rendered page) rather than one the composite dispatches to; its supported types are the types it can enrich.
    Resolves the content enricher for a media type.
    A parser that invokes content enrichers (OCR on its images or rendered pages).
    A text-recognizers entry's _min-width/_min-height: an image whose recorded dimensions (tiff:ImageWidth, tiff:ImageLength) fall short in either is not handed to the engine.
    A content enricher whose output is the document's text: what the image, audio, or video says, written into the caller's body as content the caller may treat as its own.
    The per-request switch for the "text-recognizers" list: {"parse-context": {"text-recognizers": {"enabled": false}}} runs no recognizer or annotator from the list for this parse, in every container.