Package org.apache.tika.parser.inference


package org.apache.tika.parser.inference
Inference bindings: what is offered to which engine, buffered per document tree and run at its end. Experimental in 4.1: the "inference" configuration, the unit and task contracts, and the metadata they write (tk:chunks, tk:inference-released) may change in a minor release without a deprecation cycle as 4.2 adds batched recognition, media inputs, and document-level tasks. The "engines" map is stable.
  • Class
    Description
    A model or service an inference binding calls: one endpoint or one local binary with its settings, named in the "engines" map.
    The "engines" map: user-chosen name to engine, built once at config load and closed once with it.
    One entry of the "inference" list: engine, input, tasks, filters, budget.
     
    Matches offered units against the inference bindings and runs the tasks.
    A binding resolved to its engine and tasks.
    Per-request choice among the configured bindings, under "inference" in parse-context: enabled: false runs none; bindings names the ones that run, every enabled binding when absent.
    What a binding asks of its engine, named in the binding's "tasks": one call per unit for embed, all units in one call for a document-level task.
    One thing offered to the dispatcher: the bytes of an image, a page render, a media clip, with the document they came from and, when known, its parent.
    What a binding is fed: a document's text, its rendered pages, an image, or media bytes.
    The "media" parse-context block: how audio and video are cut into segments before a MEDIA binding sees them.
     
    What an engine is shown for a unit, and so what a vector stands for.
    Cuts a document's text into the pieces a InputKind.TEXT binding sends to its engine.