Package org.apache.tika.parser.enricher
package org.apache.tika.parser.enricher
Engines a container parser invokes on bytes it has already parsed: text recognizers (OCR,
transcription) and annotators. Experimental in 4.1: the contract is one image per call,
written into the caller's body as it runs; 4.2 batches recognition per document. The intent
is additive: an engine that takes one image at a time keeps working unchanged, and one whose
service batches opts in by declaring a batch size and implementing the list call 4.2 adds.
Until then these interfaces may change in a minor release without a deprecation cycle. The
"engines" and "text-recognizers" configuration that names an engine is stable.-
ClassDescriptionMedia-type-keyed registry of content enrichers: ordinary
Parsers that a container parser invokes on bytes it has already parsed (OCR text for an image or a rendered PDF page), rather than being dispatched to by the composite parser.AParsera container parser invokes on bytes it has already parsed (OCR on an image, an embedding of a rendered page) rather than one the composite dispatches to; its supported types are the types it can enrich.Resolves the content enricher for a media type.Scope of aContentEnrichers.suspend(org.apache.tika.parser.ParseContext)orContentEnrichers.suspendRecognizers(org.apache.tika.parser.ParseContext); closing restores the prior state.A parser that invokes content enrichers (OCR on its images or rendered pages).Atext-recognizersentry's_min-width/_min-height: an image whose recorded dimensions (tiff:ImageWidth,tiff:ImageLength) fall short in either is not handed to the engine.A content enricher whose output is the document's text: what the image, audio, or video says, written into the caller's body as content the caller may treat as its own.The per-request switch for the"text-recognizers"list:{"parse-context": {"text-recognizers": {"enabled": false}}}runs no recognizer or annotator from the list for this parse, in every container.