Setting Limits
Untrusted documents can be pathological: deeply nested, self-expanding, or
crafted to run forever. Every limit below is configured in the parse-context
section of the JSON config, loaded into the ParseContext, and enforced
throughout the parse.
Limits at a glance
parse-context key |
Class | Bounds |
|---|---|---|
|
|
Depth and count of embedded documents. Unlimited by default. |
|
|
Extracted characters, XML nesting, package nesting, zip-bomb ratio. |
|
|
Total wall-clock budget per task and the stall detector. |
|
|
Total bytes written by |
|
|
Metadata size, per-field size, key size, values per field, field allow/deny lists. |
|
|
How much of an exception (stack trace, message, class name) is reported, and its length. |
Configuring limits
Every limit group is a parse-context block. AllLimitsTest
(tika-serialization/src/test/resources/configs/all-limits-test.json) exercises all of
these but exception-reporting, which has its own test:
{
"parsers": ["default-parser"],
"parse-context": {
"embedded-limits": {
"maxDepth": 10,
"throwOnMaxDepth": false,
"maxCount": 1000,
"throwOnMaxCount": false
},
"output-limits": {
"writeLimit": 100000,
"throwOnWriteLimit": false,
"maxXmlDepth": 100,
"maxPackageEntryDepth": 10,
"zipBombThreshold": 1000000,
"zipBombRatio": 100
},
"timeout-limits": {
"totalTaskTimeoutMillis": 3600000,
"progressTimeoutMillis": 60000
},
"standard-metadata-limiter-factory": {
"maxTotalBytes": 1048576,
"maxFieldSize": 102400,
"maxKeySize": 1024,
"maxValuesPerField": 100
},
"exception-reporting": {
"level": "MESSAGE_REDACTED",
"maxLength": 10000
}
}
}
TikaLoader.loadParseContext() loads them all; each class has a static
get(ParseContext) that returns the configured instance or defaults:
ParseContext context = TikaLoader.load(configPath).loadParseContext();
EmbeddedLimits embedded = EmbeddedLimits.get(context);
OutputLimits output = OutputLimits.get(context);
TimeoutLimits timeouts = TimeoutLimits.get(context);
To set them programmatically, construct and put them on the context — the same pattern for every limit class:
context.set(EmbeddedLimits.class, new EmbeddedLimits(10, true, 500, false));
context.set(OutputLimits.class, new OutputLimits(50000, true, 50, 5, 500000, 50));
context.set(TimeoutLimits.class, new TimeoutLimits(7200000, 120000));
Tests: AllLimitsTest, EmbeddedLimitsTest, OutputLimitsTest, TimeoutLimitsTest and
ExceptionReportingConfigLoadTest under tika-serialization/src/test/java/org/apache/tika/config/;
ExceptionReportingParseTest (tika-core) drives the policy through a real parse.
Embedded document limits
EmbeddedLimits bounds how deep and how many embedded documents are parsed.
| Setting | Default | Description |
|---|---|---|
|
-1 (unlimited) |
Maximum nesting depth. Recursion stops at the limit; siblings at the current level still parse. |
|
false |
Throw |
|
-1 (unlimited) |
Maximum total embedded documents. Processing stops immediately when reached. |
|
false |
Throw |
With maxDepth=1, depth-1 siblings all parse and their children do not:
container.zip (depth 0)
├── doc1.docx (depth 1) PARSED
│ ├── image1.png (depth 2) NOT PARSED
│ └── embed.xlsx (depth 2) NOT PARSED
├── doc2.pdf (depth 1) PARSED
└── doc3.txt (depth 1) PARSED
Output limits
OutputLimits bounds extracted text and structural expansion.
| Setting | Default | Description |
|---|---|---|
|
-1 (unlimited) |
Maximum characters of text to extract. Extraction stops when reached. |
|
false |
Throw |
|
100 |
Maximum XML element nesting depth. Guards against XML bombs. |
|
10 |
Maximum depth of nested package entries (zip within zip). |
|
1,000,000 |
Extracted characters before the zip-bomb ratio check activates. |
|
100 |
Maximum ratio of extracted characters to input bytes read before flagging a zip bomb. |
Timeout limits
TimeoutLimits applies two independent bounds: one on total wall-clock time,
one on time since the parser last reported progress.
| Setting | Default | Description |
|---|---|---|
|
3,600,000 (1 hour) |
Wall-clock budget for the whole task, embedded documents included. Every per-parser timeout is clipped to what remains of this budget, however it is itself configured. |
|
120,000 (2 minutes) |
How long the task may go silent before it counts as hung. Enforced — the task
actually killed — only when the parse runs in a forked JVM (Pipes, or tika-app
|
|
false |
Throw |
|
Waits on a bounded external call checkpoint progress periodically, so a
legitimate 10-minute Budgets compose recursively — a PDF inside a zip inside an email draws from the
one |
Embedded byte extraction limits
ParseMode.UNPACK writes embedded bytes out; UnpackConfig.maxUnpackBytes
caps the total.
| Setting | Default | Description |
|---|---|---|
|
10 GiB |
Maximum total bytes extracted from all embedded documents per file. |
At the limit, extraction stops for the remaining embedded documents and already-extracted bytes are kept. The parse still reports success: the truncation shows up only as a WARN in the fork’s log, not in the metadata or the result status.
{
"pipes": { "parseMode": "UNPACK" },
"parse-context": {
"unpack-config": { "maxUnpackBytes": 104857600 }
}
}
See Extracting Embedded Bytes for the rest of
UnpackConfig, and UnpackModeTest in
tika-pipes/tika-pipes-integration-tests.
Metadata limits
Configuring a MetadataWriteLimiterFactory in the ParseContext makes
Metadata.newInstance(parseContext) return a Metadata with limits already
applied, so every subsequent write is filtered.
StandardMetadataLimiterFactory factory = new StandardMetadataLimiterFactory();
factory.setMaxTotalBytes(1024 * 1024);
factory.setMaxFieldSize(100 * 1024);
factory.setMaxValuesPerField(100);
ParseContext context = new ParseContext();
context.set(MetadataWriteLimiterFactory.class, factory);
Metadata metadata = Metadata.newInstance(context);
| Setting | Default | Description |
|---|---|---|
|
10 MB |
Total estimated size of all metadata in UTF-16 bytes. Further metadata is
dropped and |
|
100 KB |
Maximum size of a single field’s value(s) in UTF-16 bytes. Longer values are truncated. |
|
1024 |
Maximum metadata key length in UTF-16 bytes. Longer keys are truncated. |
|
10 |
Maximum values on a multi-valued field. Extra values are dropped. |
|
empty (all) |
If non-empty, only these fields are stored, plus the always-included fields below. |
|
empty (none) |
Never stored, unless always-included. |
|
false |
Whether to store empty or null values. |
Always-included fields
StandardMetadataLimiter.ALWAYS_SET_FIELDS and ALWAYS_ADD_FIELDS bypass
includeFields/excludeFields and the total-size budget, because parser
dispatch and error reporting depend on them. Everything except tk:content is
still truncated at max(maxFieldSize, 300).
Always set: Content-Type, Content-Length, Content-Encoding,
Content-Disposition, tk:content-type-override,
tk:content-type-parser-override, tk:content-type-hint, tk:content,
tk:resource-name, tk:exception:container-exception,
access-permission:extract-content,
access-permission:extract-for-accessibility.
Always added (multi-valued): tk:parsed-by, tk:exception:embedded-exception.
Exception reporting
ExceptionReporting controls how much detail Tika reports when a parse fails:
the tk:exception:container-exception, tk:exception:embedded-exception,
tk:exception:warn and tk:exception:embedded-stream-exception metadata
values. Stack traces and messages can carry file paths, hostnames and fragments
of the document, and the metadata limiter is no backstop: with no
MetadataWriteLimiterFactory configured these values are unbounded, and with
one, only container-exception and embedded-exception are always-included
(cut at max(maxFieldSize, 300) UTF-16 bytes); warn and
embedded-stream-exception are ordinary fields, subject to
includeFields/excludeFields and dropped past maxValuesPerField (see
Always-included fields). The limiter’s cut appends no marker, so keep
maxLength (characters) under half maxFieldSize (UTF-16 bytes) if you want
the …[truncated] marker to survive.
| Setting | Default | Description |
|---|---|---|
|
|
|
|
-1 (unlimited) |
Length at which the formatted exception is cut; a truncated value is
|
"exception-reporting": {
"level": "MESSAGE_REDACTED",
"maxLength": 10000
}
The same policy applies to the container and to every embedded document, to
parser-level warnings (tk:exception:warn,
tk:exception:embedded-stream-exception), to every pipes result message
(fetch/emit/crash inside the worker, and the results the parent fabricates when
the worker never reports one), and to every tika-server error body built from an
exception — 422 and 500. It is loaded from the config only and is
rejected in a per-request parse-context — tika-server /rmeta/config and
friends, /pipes and /async answer 400 with the reason, tika-grpc’s
parse_context_json answers INVALID_ARGUMENT — so a caller cannot turn
redaction back off.
Programmatically, parsers and callers should format exceptions through
ExceptionUtils.format(Throwable, ParseContext). A parser still calling the
deprecated two-argument EmbeddedDocumentUtil.recordException /
recordEmbeddedStreamException — including any third-party parser compiled
against 4.0.0 — reports FULL whatever the policy says. To set a policy in a
parser test:
ParseContext context = new ParseContext();
context.set(ExceptionReporting.class,
new ExceptionReporting(ExceptionReporting.Level.MESSAGE_REDACTED, 500));
Recommendations
-
Set limits whenever the content is untrusted.
-
Set
exception-reportingtoMESSAGE_REDACTEDwith a finitemaxLengthwhen the caller is not the operator. -
Use
includeFieldsto capture only the metadata you need. -
Check
tk:warn:truncated-metadatarather than guessing. -
Combine with process isolation — limits protect against memory blowups, process isolation protects against crashes.
-
Test with adversarial files;
MockParser(tika-core test jar) simulates hangs, OOMs, and huge output.
See also
-
Robustness — process isolation and fault tolerance
-
Configuration — general Tika configuration
-
Extracting Embedded Bytes — the rest of
UnpackConfig -
Timeouts — the full timeout model