Package org.apache.tika.parser.pdf
package org.apache.tika.parser.pdf
-
ClassDescriptionHow the parser uses a PDF's structure tree (marked content / tagged PDF).Deprecated.The 4.0 spelling of
ImageFormat.The 4.0 spelling ofImageType.What PDFBox draws when it renders a page; PDF-only, so it stays on the parser.The 4.0 spelling ofTextPolicy.The 4.0 spelling ofPagesConfig.Auto.This counts the number of pages that OCR would have been run or was run depending on the settings.Text extraction that follows a tagged PDF's structure tree.PDF parser.Config for PDFParser.Mode for checking document access permissions.Deprecated.One way to hand PDFBox a document: in-memory content is read in place through a zero-copy view, anything else through PDFBox's buffered file reader.
"pages"(seePagesConfig).