Pipes Configuration
The pipes section of the JSON config controls the pipeline process itself:
how many forked JVMs to run, timeouts, memory management, and parse behavior.
{
"pipes": {
"numClients": 4,
"socketTimeoutMillis": 60000,
"maxFilesProcessedPerProcess": 10000,
"parseMode": "RMETA",
"forkedJvmArgs": ["-Xmx512m"]
}
}
Process Management
| Field | Default | Description |
|---|---|---|
|
CPU-derived |
Number of parallel forked JVMs. Each processes one document at a time. Defaults to |
|
|
JVM arguments for forked processes (e.g., |
|
|
Path to the Java executable for forked processes. |
|
|
Restart forked processes after this many files. Prevents slow-building memory leaks in parsing libraries. |
|
system default |
Deprecated since 4.1, removal planned for 5.0. Set |
Where temporary files go
Set -Djava.io.tmpdir on the parent JVM. It has to be a launch flag because of Tika’s
dependencies: POI, PDFBox and the rest create temp files through the JDK, which reads
java.io.tmpdir once at JVM start, so nothing Tika sets at runtime reaches them.
Tika, every library it calls, and its forked pipes servers all honor the flag: each fork gets a private subdirectory under it (its whole
java.io.tmpdir points there, so spooled input, unpacked embedded files and JVM crash logs
all land inside), and the parent deletes that subdirectory when the fork is torn down.
Tika checks the directory exists and is writable at config load and refuses to start
otherwise, rather than failing on the first document that spools.
If the parent dies, a surviving fork deletes its own subdirectory as it exits. Only when the
whole process family is killed at once (SIGKILL of the group, container stop) do
subdirectories survive, with whatever the forks were spooling inside; Tika never deletes
directories another process created. Point java.io.tmpdir at a disk-backed directory on a
volume where filling it does not take out the OS, and apply your own retention to
pipes-server-* there.
|
DO NOT USE tmpfs ( Spool size is bounded by the input, not by any Tika setting: one large archive expanding
into tmpfs can exhaust memory for every process on the host or get a container evicted,
and a fork directory orphaned by a killed parent pins that RAM until someone deletes it.
A slow run is recoverable; a lost host is not. If you must, give tmpfs a dedicated mount
with |
Embedded-object cache memory budget
Each fork holds a process-wide in-memory budget for stream caching (the rewind buffers used
when digesting embedded documents, and the in-memory views parsers read random-access
content from), so small embedded objects stay in RAM instead of spilling to a temp file at
the per-object 1MB threshold. The default is a quarter of the fork’s max heap, so raising
-Xmx raises it; a value set via the system property is clamped to that same quarter-heap
ceiling. The effective value and where it came from are logged at fork startup. Tune it with
a system property in forkedJvmArgs (plain bytes, no unit suffix; ⇐0 disables the
budget and restores the per-object threshold):
"forkedJvmArgs": ["-Xmx1g", "-Dtika.pipes.cacheMemoryBudgetBytes=134217728"]
Size -Xmx so that three quarters of it covers the parsing working set; the remaining
quarter is what the budget may hold. It is a ceiling on live cached bytes, not a
preallocation. In per-client mode every fork holds its own budget (numClients x a quarter
of each fork’s heap in aggregate). In shared-server mode all concurrent parses in the single
forked server share one pool, so a single large document can take most of it and push its
siblings to disk for a while. With no explicit heap flag the fork’s heap is sized from the
host (see forkedJvmArgs), and the budget scales with it — on a large host that is
several GB per fork by default.
Timeouts
See also Timeouts for the full timeout model.
| Field | Default | Description |
|---|---|---|
|
|
Maximum time (ms) to wait for data from a forked process. If no heartbeat or result is received within this window, the parse is considered hung. Also serves as the fork’s idle-shutdown timer: a fork that receives no work for this long exits and is restarted transparently on next use. |
|
|
Interval (ms) between heartbeats sent from the forked process. Must be significantly less than |
|
|
Socket read timeout (ms) applied to a freshly-connected fork until it completes its READY handshake, after which |
|
|
Ceiling for request-supplied timeout limits: a per-request |
|
|
Maximum time (ms) to wait for an available forked process when all are busy. |
Parse Behavior
| Field | Default | Description |
|---|---|---|
|
|
How embedded documents are handled: |
|
|
What to do when a parse fails: |
|
|
When |
How much of a failure is written into tk:exception:* metadata and into a PipesResult
message is set by parse-context.exception-reporting, not here — see
Exception reporting.
Async / Emit Batching
These settings control how parsed results are batched before sending to emitters.
| Field | Default | Description |
|---|---|---|
|
|
Number of emitter threads. |
|
|
Size of the fetch/emit tuple queue. |
|
|
Flush the emit batch if nothing has been emitted within this many milliseconds, even if the batch is not full. |
|
|
Flush the emit batch when the estimated size reaches this many bytes. |
|
|
When |
IPC and Inline Payload Limits
| Field | Default | Description |
|---|---|---|
|
|
Maximum size in bytes of a single IPC message between the client and the fork. This limit is bidirectional: it applies both to parse results returned from the fork (FINISHED) and to requests sent from the client (NEW_REQUEST). Raising it lets very large documents pass over IPC; set the forked JVM |
|
|
Largest document carried inline to the fork instead of being written to a file first. A host that already holds the content (tika-server’s |
Emit Strategy
emitStrategy controls whether parsed extracts are emitted directly from the forked PipesServer or passed back to the parent process first. The default is balanced for typical workloads — tune only if you have a memory or throughput problem.
{
"pipes": {
"emitStrategy": {
"type": "DYNAMIC",
"thresholdBytes": 100000
}
}
}
| Field | Default | Description |
|---|---|---|
|
|
One of |
|
|
Only used when |
Distributed Config Store
For multi-host pipelines (e.g., shared-server clusters) you can store fetcher/emitter configuration in a distributed backend instead of memory. Most users should leave the defaults.
| Field | Default | Description |
|---|---|---|
|
|
Backend for storing fetcher/emitter configurations. |
|
|
JSON object (as a string) with backend-specific parameters. Structure depends on |
Shared Server Mode (Experimental)
| Field | Default | Description |
|---|---|---|
|
|
When |
See Shared Server Mode for details.
Complete examples
Worked-out end-to-end configs from the test tree, so the syntax stays current. The tests that cover them assert that the JSON parses; they do not instantiate every component, so a config can load cleanly and still name something that is missing from a release classpath.
Filesystem-to-filesystem pipeline
{
"content-handler-factory": {
"basic-content-handler-factory": {
"type": "TEXT",
"writeLimit": -1,
"throwOnWriteLimitReached": true
}
},
"fetchers": {
"fsf": {
"file-system-fetcher": {
"basePath": "FETCHER_BASE_PATH",
"extractFileSystemMetadata": false
}
}
},
"emitters": {
"fse": {
"file-system-emitter": {
"basePath": "EMITTER_BASE_PATH",
"fileExtension": "json",
"onExists": "EXCEPTION"
}
}
},
"pipes-iterator": {
"file-system-pipes-iterator": {
"basePath": "FETCHER_BASE_PATH",
"countTotal": true,
"fetcherId": "fsf",
"emitterId": "fse"
}
},
"pipes": {
"parseMode": "RMETA",
"onParseException": "EMIT",
"numClients": 4,
"emitIntermediateResults": "EMIT_INTERMEDIATE_RESULTS",
"forkedJvmArgs": ["-Xmx512m"],
"emitStrategy": {
"type": "DYNAMIC",
"thresholdBytes": 1000000
}
},
"auto-detect-parser": {
"throwOnZeroBytes": false
},
"parse-context": {
"mock-digester-factory": {},
"timeout-limits": {
"progressTimeoutMillis": 5000
}
},
"plugin-roots": "PLUGINS_PATHS"
}
Tokens (FETCHER_BASE_PATH, EMITTER_BASE_PATH, PLUGINS_PATHS, EMIT_INTERMEDIATE_RESULTS) are substituted by the test harness — replace them with real values in production configs. The first three are paths; EMIT_INTERMEDIATE_RESULTS is the boolean emitIntermediateResults flag.
Emit-all variant
{
"fetchers": {
"fsf": {
"file-system-fetcher": {
"basePath": "FETCHER_BASE_PATH",
"extractFileSystemMetadata": false
}
}
},
"emitters": {
"fse": {
"file-system-emitter": {
"basePath": "EMITTER_BASE_PATH",
"fileExtension": "json",
"onExists": "EXCEPTION"
}
}
},
"pipes": {
"numClients": 1,
"forkedJvmArgs": [
"-Xmx256m"
],
"emitStrategy": {
"type": "EMIT_ALL"
}
},
"parse-context": {
"timeout-limits": {
"progressTimeoutMillis": 60000
}
},
"plugin-roots": "PLUGINS_PATHS"
}
Shared-server (YOLO) mode
{
"content-handler-factory": {
"basic-content-handler-factory": {
"type": "TEXT",
"writeLimit": -1,
"throwOnWriteLimitReached": true
}
},
"fetchers": {
"fsf": {
"file-system-fetcher": {
"basePath": "FETCHER_BASE_PATH",
"extractFileSystemMetadata": false
}
}
},
"emitters": {
"fse": {
"file-system-emitter": {
"basePath": "EMITTER_BASE_PATH",
"fileExtension": "json",
"onExists": "REPLACE"
}
}
},
"pipes-iterator": {
"file-system-pipes-iterator": {
"basePath": "FETCHER_BASE_PATH",
"countTotal": true,
"fetcherId": "fsf",
"emitterId": "fse"
}
},
"pipes": {
"parseMode": "RMETA",
"onParseException": "EMIT",
"numClients": 4,
"useSharedServer": true,
"emitIntermediateResults": "EMIT_INTERMEDIATE_RESULTS",
"forkedJvmArgs": ["-Xmx512m"],
"emitStrategy": {
"type": "DYNAMIC",
"thresholdBytes": 1000000
}
},
"auto-detect-parser": {
"throwOnZeroBytes": false
},
"parse-context": {
"mock-digester-factory": {},
"timeout-limits": {
"progressTimeoutMillis": 5000
}
},
"plugin-roots": "PLUGINS_PATHS"
}
See Shared Server Mode for the trade-offs.
Tika Pipes config template
{
"content-handler-factory": {
"basic-content-handler-factory": {
"type": "TEXT",
"writeLimit": -1,
"throwOnWriteLimitReached": true
}
},
"parsers": [
{
"default-parser": {}
},
{
"pdf-parser": {
"extractActions": true,
"extractInlineImages": true,
"extractIncrementalUpdateInfo": true,
"parseIncrementalUpdates": true
}
},
{
"ooxml-parser": {
"includeDeletedContent": true,
"includeMoveFromContent": true,
"extractMacros": true
}
},
{
"office-parser": {
"extractMacros": true
}
}
],
"fetchers": {
"fsf": {
"file-system-fetcher": {
"basePath": "FETCHER_BASE_PATH",
"extractFileSystemMetadata": false
}
}
},
"emitters": {
"fse": {
"file-system-emitter": {
"basePath": "EMITTER_BASE_PATH",
"fileExtension": "json",
"onExists": "EXCEPTION"
}
}
},
"pipes-iterator": {
"file-system-pipes-iterator": {
"basePath": "FETCHER_BASE_PATH",
"countTotal": true,
"fetcherId": "fsf",
"emitterId": "fse"
}
},
"pipes": {
"parseMode": "RMETA"
},
"plugin-roots": "PLUGIN_ROOTS"
}
For per-plugin pipeline examples (S3, OpenSearch, JDBC, Kafka, etc.), see the relevant page under Plugins.