Package org.apache.tika.parser.inference
Class InferenceUnit
java.lang.Object
org.apache.tika.parser.inference.InferenceUnit
One thing offered to the dispatcher: the bytes of an image, a page render, a media clip,
with the document they came from and, when known, its parent. Captured at the offer so
a task can run after the document has closed; the bytes live in a file the dispatcher
owns until the flush, and the id paths let the flush find the metadata the recursive
wrapper kept once the parser's own objects are no longer read. A task writes its result
on the
destination, which the dispatcher decides: the parent
for an inline part or a page render, the document itself otherwise, and the top-level
document for everything when the parse produces one metadata object.-
Constructor Summary
ConstructorsConstructorDescriptionInferenceUnit(MediaType type, Metadata target, Metadata parent, String text) AInputKind.TEXTunit: the target's extracted text, held in memory.InferenceUnit(InputKind kind, MediaType type, Metadata target, Metadata parent, Path path, int page) A unit that is one page of its target:pageis 1-based. -
Method Summary
Modifier and TypeMethodDescriptionbyte[]getBytes()Readable during a task run only; the file is deleted when the flush ends.Where a task writes this unit's results.getKind()intgetPage()The 1-based page this unit renders, forInputKind.PAGES; -1 otherwise.The document the target is embedded in; null at top level.getPath()longgetSize()The document the bytes belong to.getText()The text of aInputKind.TEXTunit; null for the rest.getType()booleanisLifted()Whether the results go on another document than the target, which a locator then names.static booleanWhether results for this document belong on the one it is part of, not on it.
-
Constructor Details
-
InferenceUnit
public InferenceUnit(InputKind kind, MediaType type, Metadata target, Metadata parent, Path path) throws IOException - Throws:
IOException
-
InferenceUnit
public InferenceUnit(InputKind kind, MediaType type, Metadata target, Metadata parent, Path path, int page) throws IOException A unit that is one page of its target:pageis 1-based.- Throws:
IOException
-
InferenceUnit
AInputKind.TEXTunit: the target's extracted text, held in memory.
-
-
Method Details
-
lifts
Whether results for this document belong on the one it is part of, not on it. -
getDestination
Where a task writes this unit's results. -
isLifted
public boolean isLifted()Whether the results go on another document than the target, which a locator then names. -
getText
The text of aInputKind.TEXTunit; null for the rest. -
getPage
public int getPage()The 1-based page this unit renders, forInputKind.PAGES; -1 otherwise. -
getKind
-
getType
-
getTarget
The document the bytes belong to. -
getParent
The document the target is embedded in; null at top level. -
getTargetIdPath
-
getParentIdPath
-
getPath
-
getSize
public long getSize() -
getBytes
Readable during a task run only; the file is deleted when the flush ends.- Throws:
IOException
-