PDFParser Configuration
Configuration options for PDFParser.
Basic Configuration
{
"parsers": [
{
"pdf-parser": {
"extractInlineImages": true,
"sortByPosition": true
}
},
{
// Keep Tika's other default parsers. Without this, this config is PDF-only.
"default-parser": {}
}
]
}
Full Configuration
Every option with its default value and an inline comment describing it. The default-parser entry
keeps the rest of Tika’s parsers loaded — see Configuration for why
that matters.
{
// A "parsers" list loads ONLY the parsers it names; the "default-parser" entry at
// the bottom keeps all the others. Windows paths in JSON need forward slashes or
// escaped backslashes.
"parsers": [
{
"pdf-parser": {
// Enforce the PDF's access permissions. DONT_CHECK ignores them.
// Options: DONT_CHECK, ALLOW_EXTRACTION_FOR_ACCESSIBILITY, IGNORE_ACCESSIBILITY_ALLOWANCE
"accessCheckMode": "DONT_CHECK",
// Character-width tolerance for inserting spaces (PDFBox).
"averageCharTolerance": 0.3,
// Collect per-stream IOExceptions in metadata and rethrow after parsing.
"catchIntermediateIOExceptions": true,
// Detect and correct rotated (angled) text runs within a page.
"detectAngles": false,
// Line-height multiple that starts a new paragraph (PDFBox).
"dropThreshold": 2.5,
// Estimate where spaces belong between words (most PDFs lack explicit spaces).
"enableAutoSpace": true,
// Extract AcroForm field content.
"extractAcroFormContent": true,
// Extract PDF actions; JavaScript macros become embedded documents.
"extractActions": false,
// Extract annotation text (comments, form-field captions).
"extractAnnotationText": true,
// Extract outline / bookmark text.
"extractBookmarksText": true,
// Record font names in metadata.
"extractFontNames": false,
// Record metadata about incremental updates (whether present, how many).
"extractIncrementalUpdateInfo": true,
// Record inline-image metadata only, without rendering (faster than extractInlineImages).
"extractInlineImageMetadataOnly": false,
// Render and extract inline images from content streams.
"extractInlineImages": false,
// Emit each unique inline image (by object id) only once.
"extractUniqueInlineImagesOnly": true,
// If the PDF has an XFA form, process only it.
"ifXFAExtractOnlyXFA": false,
// Ignore content-stream space glyphs; rely on the spacing algorithm (PDFBOX-3774).
"ignoreContentStreamSpaceGlyphs": false,
// EXPERT: replace the inline-image factory; give a class implementing
// ImageGraphicsEngineFactory, e.g.:
// "imageGraphicsEngineFactoryClass": "com.example.MyImageGraphicsEngineFactory"
// What is done with the document's pages: how they are rendered, where a page's
// text comes from, what is released to inference, whether renders are emitted. The
// same block, under "parse-context", is the default for every rendering parser; this
// one overlays it for PDFs, field by field. The 4.0 "ocr" and "imageStrategy"
// spellings still load and land here.
"pages": {
// Where a page's text comes from. EXTRACT: the content stream, never OCR. AUTO:
// the content stream, OCR on pages the "ocr.auto" thresholds flag.
// EXTRACT_AND_OCR: both, every page. OCR: the engine only. NONE: no text at all;
// pages are still rendered for annotators, inference and emission. The engine
// comes from "text-recognizers".
// Options: EXTRACT, AUTO, EXTRACT_AND_OCR, OCR, NONE
"text": "AUTO",
// How pages are rendered for OCR and PAGES inference; "emit.render" overlays it for
// the renders emitted as embedded documents.
"render": {
// Target resolution.
"dpi": 300,
// A box the image must fit: scaled down to fit, never enlarged; -1 = no bound.
"maxWidth": -1,
"maxHeight": -1,
// Smaller than this at the target dpi and the page is not rendered at all.
"minWidth": 2,
"minHeight": 2,
// Options: RGB, GRAY
"imageType": "GRAY",
// Options: PNG, TIFF, JPEG
"imageFormat": "PNG",
// ImageIO quality (0.0-1.0): fidelity for JPEG, an inverted effort knob for PNG.
"imageQuality": 0.5,
// Skip a page that would still be larger than this area (w x h); -1 = no limit.
"maxImagePixels": 100000000
},
"ocr": {
// Max pages to OCR per document, counted from the first; -1 = no limit.
"maxPages": -1,
// Per-page character thresholds that send a page to OCR under "text": "AUTO".
"auto": {
"totalCharsPerPage": 10,
"unmappedUnicodeCharsPerPage": 10
}
},
// What the parser releases to the inference bindings (see the inference page):
// TEXT is the extracted text; PAGES renders every page for a PAGES binding.
"inference": ["TEXT"],
// Emit page renders as RENDERING embedded documents, at the end of each page.
"emit": {
"enabled": false,
// Max pages to emit, counted from the first and bounded by maxPages; -1 = no limit.
"maxPages": -1,
// Emit only for documents embedded as one of these resource types; empty = all.
"resourceTypes": [],
// Overlay on "render" for the emitted images only.
"render": {
"dpi": 96,
"imageType": "RGB"
}
}
},
// What PDFBox draws when it renders a page: everything, text only, no text, or
// vector graphics only.
// Options: NO_TEXT, TEXT_ONLY, VECTOR_GRAPHICS_ONLY, ALL
"renderingStrategy": "ALL",
// Tagged PDFs: follow the structure tree (headings, lists, tables) or use the plain
// text stripper. AUTO uses the tree on pages that pass a quality gate; TAGS always;
// NONE never. A PDF without a structure tree always uses the stripper.
// Options: AUTO, TAGS, NONE
"markedContent": {
"strategy": "AUTO",
// AUTO only: least fraction of a page's text the tree must claim.
"minCoverage": 0.5,
// AUTO only: most tree leaves for a page that may point at content the page lacks.
"maxDanglingRatio": 0.2
},
// Max incremental updates to parse when parseIncrementalUpdates is true.
"maxIncrementalUpdates": 10,
// Max memory to load a PDF before buffering to a temp file (default 512MB).
"maxMainMemoryBytes": 536870912,
// Max pages to process; -1 = no limit.
"maxPages": -1,
// Parse prior incremental-update versions as embedded documents.
"parseIncrementalUpdates": false,
// EXPERT: set the Sun KCMS color-management system property. Default false.
"setKCMS": false,
// Sort text by x/y position; helps some PDFs, can interleave columns in others.
"sortByPosition": false,
// Space-width tolerance for inserting spaces (PDFBox).
"spacingTolerance": 0.5,
// Remove text drawn twice over the same region (faked bold); can be slow.
"suppressDuplicateOverlappingText": false,
// Throw on an encrypted payload instead of skipping it.
"throwOnEncryptedPayload": false
}
},
{
// Keep Tika's other default parsers. Without this, this config is PDF-only.
"default-parser": {}
}
]
}
Pages (pages)
Where a page’s text comes from, how pages are rendered, how many may be OCR’d, what is released to
inference and whether page images are emitted are the pages block,
shared with every parser that renders. The PDF parser’s own pages entry overlays the one under
parse-context, field by field, and can be set per request:
{
"pdf-parser": {
"pages": {
"text": "AUTO",
"emit": { "enabled": true, "maxPages": 1, "render": { "maxWidth": 256, "maxHeight": 256, "imageType": "RGB" } }
}
}
}
renders a colour first-page thumbnail that fits 256 x 256 pixels while OCR keeps its 300 dpi
grayscale pages. The 4.0 ocr and imageStrategy spellings still load and land in pages; the
table at the end of that page maps them.
Two things are the PDF’s own:
renderingStrategy-
What PDFBox draws when it renders a page for OCR:
ALL(default),TEXT_ONLY,NO_TEXTorVECTOR_GRAPHICS_ONLY. A configuredrenderersengine always draws everything. The 4.0ocr.renderingStrategystill loads. extractInlineImages-
The images embedded in a PDF are a separate path from page renders: with it on they are parsed as embedded documents, so
image-parserenriches each one with the same engine. A scanned page that is one full-page image is therefore OCR’d twice when both page OCR and inline images are on. The 4.0imageStrategy: RAW_IMAGESis this switch.
The PDF parser does not OCR anything itself. It renders pages and hands each rendered image to
the text recognizer resolved for the render type (image/png by default): the engine named in
the top-level text-recognizers list, or, with no list, the one found among the loaded parsers
and logged at startup. See Configuration for
choosing the engine and Tesseract OCR for
its options. Emitted page renders are enriched by the PDF parser itself, once per page: the text
recognizer where text says so, and every other enricher (an image embedder, a tagging VLM) on
every page, with its results such as tk:chunks landing on the PDF. The embedded copy of a
render is parsed for its metadata but not enriched again.
pdf:ocr-page-count records how many pages tripped the verdict, whether or not an engine ran.
Tagged PDFs (markedContent)
A tagged PDF carries a structure tree: paragraphs, headings, lists, tables, links, and which
content is an artifact (running headers, footers, page numbers). The markedContent block says
whether the parser follows that tree. Every page’s words and spacing still come from the same
text stripper and the same sortByPosition, enableAutoSpace and tolerance settings; the tree
only decides which XHTML elements the words land in and, across blocks, their order.
strategy-
AUTO, the default, follows the tree on pages that pass the gate below and writes the other pages exactly as the stripper would.TAGSfollows the tree on every page that has tagged text.NONEignores the tree and writes what the stripper writes, the output of Tika before 4.1.0. A PDF with no structure tree is written as withNONEwhatever the strategy, anddetectAnglesalways wins. minCoverage,maxDanglingRatio-
The
AUTOgate, per page. A page uses its tags only when the tree claims at leastminCoverageof the page’s text (artifact content counts, since a producer that marks the body as artifact has not described the page), at mostmaxDanglingRatioof the tree’s leaves for the page point at content the page never draws, the tree puts at least some of the page’s text in an element that holds text (a paragraph, heading, cell or item: a tree of bare spans or divisions has no paragraphs to offer), and the tree does not cut the page’s words into pieces (some form generators put every glyph in its own paragraph, which would write each word as a column of letters; a page where more than 30% of its words span three blocks or more goes to the stripper). Both thresholds start loose (0.5 and 0.2) until a corpus says otherwise;minCoverageabove 1 turns every page back to the stripper.
With tags, a structure element becomes the XHTML element of the same meaning (P to <p>, H1
to <h1>, L to <ul> or <ol> by its ListNumbering, Table/TR/TH/TD to a table,
Link to <a href> when its annotation has a URI) and any other type becomes a <div> or
<span> with the type as its class. A Figure or Formula takes its place in the page even
when its content is an image, and its Alt text is written as its text, so the output reads
as a screen reader would read it. Text the tree does not reference follows the tagged content
of its page in <div class="untagged">, and artifact content in <div class="artifact">, so
nothing is lost and a consumer can drop headers and footers. Whitespace at the edges of a
block element is dropped; the separators between words stay outside links and spans. Text a
producer draws twice for a bold effect, tagging one copy, reads once per copy: the tagged
copy in its element, the other in the untagged division, where the stripper alone reads the
interleaved glyphs as doubled letters.
The output is well-formed XHTML whose nesting is the tree’s, repaired where a tree leaves text
or elements somewhere HTML gives them no place, the way an HTML parser would repair the same
markup. A block (a paragraph, table, list, figure) that the tree nests inside a paragraph or
heading closes it first, and text of the paragraph after the block gets a paragraph of its own,
so a chain of paragraphs nested in each other comes out as a sequence of paragraphs and a
paragraph inside a heading is not heading text. Text straight inside a
division (a Div, Sect, Part, Figure, a custom type) is written as paragraphs, split
where the stripper would split the flat page, so a tree of containers reads like the stripper’s
output inside them. Text straight inside a table becomes its <caption> (as does a Caption
element there), inside a row a <td>, inside a list an <li>, and anything else straight inside
a list is wrapped in an <li> too. A block inside a link or span, or straight inside a table or
row, becomes a <span> with the type as its class, with a line break at each edge, so the link
stays one and a chart’s labels inside a table land in its caption rather than in an invented
row; such spans do not nest, the innermost wins. An item outside a list gets its <ul>, a row
outside a table its <table>. Tables and lists inside a link or span stay where the tree put
them.
The metadata says what happened:
pdf:marked-content-pages-tagged, pdf:marked-content-pages-fallback, and
pdf:marked-content-rejections with the reason per page that fell back.
The old boolean extractMarkedContent still loads: true is TAGS, false is NONE.