Uses of Package
org.apache.tika.parser.pages
Packages that use org.apache.tika.parser.pages
Package
Description
The
"pages" block: one configuration for every parser that renders a document to
pixels -- how it renders, where a page's text comes from, what is released to inference,
whether renders are emitted.-
Classes in org.apache.tika.parser.pages used by org.apache.tika.parser.microsoftClassDescriptionWhat a parser does with a document's pages when the document has to be rendered to get pixels: how it renders (
render), where a page's text comes from (text), the recognizer budget and verdict (ocr), what is released to inference bindings (inference), and whether renders are emitted as embedded documents (emit). -
Classes in org.apache.tika.parser.pages used by org.apache.tika.parser.pagesClassDescriptionWhat a parser does with a document's pages when the document has to be rendered to get pixels: how it renders (
render), where a page's text comes from (text), the recognizer budget and verdict (ocr), what is released to inference bindings (inference), and whether renders are emitted as embedded documents (emit).WhenTextPolicy.AUTOsends a page to OCR: too few characters, or too many glyphs without a Unicode mapping.Whether, and for which documents, renders are emitted as RENDERING embedded documents.The recognizer side oftext: how many pages may be OCR'd, how AUTO decides.Where a page's text comes from: the document's own text, OCR of the rendered page, both, the per-page verdict, or nowhere. -
Classes in org.apache.tika.parser.pages used by org.apache.tika.parser.pdfClassDescriptionWhat a parser does with a document's pages when the document has to be rendered to get pixels: how it renders (
render), where a page's text comes from (text), the recognizer budget and verdict (ocr), what is released to inference bindings (inference), and whether renders are emitted as embedded documents (emit).WhenTextPolicy.AUTOsends a page to OCR: too few characters, or too many glyphs without a Unicode mapping.Where a page's text comes from: the document's own text, OCR of the rendered page, both, the per-page verdict, or nowhere.