Class InferenceDispatcher

java.lang.Object
org.apache.tika.parser.inference.InferenceDispatcher
All Implemented Interfaces:
TransientParseState, ParseHook

public final class InferenceDispatcher extends Object implements ParseHook, TransientParseState
Matches offered units against the inference bindings and runs the tasks. Built once at config load and run as a ParseHook: every document, top-level or embedded, is offered once by the auto-detect parser. Units are buffered per binding for the whole top-level parse and flushed at its end, so a document is one request per binding, not one per unit; the buffer is bounded by each binding's maxChunks and maxBytes, and its bytes live in files the dispatcher owns until the flush. At the flush every unit is aimed at the metadata that is still read: under the recursive wrapper, the copy it kept for the unit's document (an inline part or a page render lands on its parent); outside it, where the parse hands back one metadata object and the embedded documents' own are discarded, the top-level document. A request narrows the bindings through InferenceSelection. Failures mark the document and never fail the parse. TEXT is not offered during the parse: whoever holds the finished metadata list runs text(java.util.List<org.apache.tika.metadata.Metadata>, org.apache.tika.parser.ParseContext) over it.
Since:
Apache Tika 4.1