Package org.apache.tika.parser.enricher
Interface TextRecognizer
- All Superinterfaces:
AutoCloseable,Closeable,ContentEnricher,Engine
- All Known Implementing Classes:
AbstractVLMParser,ClaudeVLMParser,GeminiVLMParser,OpenAIVLMParser,Tess4JParser,TesseractOCRParser
A content enricher whose output is the document's text: what the image, audio, or video
says, written into the caller's body as content the caller may treat as its own. Other
enrichers annotate (captions, embeddings, tags); a caller never substitutes their output
for text it already has. Only text recognizers count when a caller such as the PDF
parser's AUTO OCR strategy asks whether an engine can stand in for extracted text.
With no "text-recognizers" list, a recognizer on the classpath is found by this
interface (see ContentEnrichers.resolve(org.apache.tika.parser.Parser)). Decorators hide it: ask through
ContentEnrichers.asTextRecognizer(org.apache.tika.parser.Parser), not instanceof.
- Since:
- Apache Tika 4.1
-
Method Summary
Modifier and TypeMethodDescriptiondefault booleanrecognizesText(ParseContext context) Whether this configured instance recognizes text for this parse.
-
Method Details
-
recognizesText
Whether this configured instance recognizes text for this parse. An OCR engine told to skip OCR recognizes nothing; a VLM answers from its configuration, since only the operator knows whether its prompt transcribes or captions.
-