Uses of Package
org.apache.tika.parser.pages

Package
Description
 
The "pages" block: one configuration for every parser that renders a document to pixels -- how it renders, where a page's text comes from, what is released to inference, whether renders are emitted.
 
  • Class
    Description
    What a parser does with a document's pages when the document has to be rendered to get pixels: how it renders (render), where a page's text comes from (text), the recognizer budget and verdict (ocr), what is released to inference bindings (inference), and whether renders are emitted as embedded documents (emit).
  • Class
    Description
    What a parser does with a document's pages when the document has to be rendered to get pixels: how it renders (render), where a page's text comes from (text), the recognizer budget and verdict (ocr), what is released to inference bindings (inference), and whether renders are emitted as embedded documents (emit).
    When TextPolicy.AUTO sends a page to OCR: too few characters, or too many glyphs without a Unicode mapping.
    Whether, and for which documents, renders are emitted as RENDERING embedded documents.
    The recognizer side of text: how many pages may be OCR'd, how AUTO decides.
    Where a page's text comes from: the document's own text, OCR of the rendered page, both, the per-page verdict, or nowhere.
  • Class
    Description
    What a parser does with a document's pages when the document has to be rendered to get pixels: how it renders (render), where a page's text comes from (text), the recognizer budget and verdict (ocr), what is released to inference bindings (inference), and whether renders are emitted as embedded documents (emit).
    When TextPolicy.AUTO sends a page to OCR: too few characters, or too many glyphs without a Unicode mapping.
    Where a page's text comes from: the document's own text, OCR of the rendered page, both, the per-page verdict, or nowhere.