Pipes Configuration
The pipes section of the JSON config controls the pipeline process itself:
how many forked JVMs to run, timeouts, memory management, and parse behavior.
{
"pipes": {
"numClients": 4,
"socketTimeoutMillis": 60000,
"maxFilesProcessedPerProcess": 10000,
"parseMode": "RMETA",
"forkedJvmArgs": ["-Xmx512m"]
}
}
Process Management
| Field | Default | Description |
|---|---|---|
|
CPU-derived |
Number of parallel forked JVMs. Each processes one document at a time. Defaults to |
|
|
JVM arguments for forked processes (e.g., |
|
|
Path to the Java executable for forked processes. |
|
|
Restart forked processes after this many files. Prevents slow-building memory leaks in parsing libraries. |
|
system default |
Directory for temporary files. Each fork gets a subdirectory here, and the fork’s whole |
The parent deletes a fork’s subdirectory when that fork is torn down or fails to start, so a
crashing fork does not accumulate them. A parent killed abruptly (SIGKILL, container stop)
cannot, and its subdirectories survive. On a RAM-backed filesystem those leaks consume memory
rather than disk, and /dev/shm is commonly sized at half of RAM — so if you point
tempDirectory at one, sweep it on service start.
Embedded-object cache memory budget
Each fork holds a process-wide in-memory budget for stream caching (chiefly the rewind
buffers used when digesting embedded documents), so small embedded objects stay in RAM
instead of spilling to a temp file at the per-object 1MB threshold. The default is 256MB per
fork, clamped to a quarter of the fork’s max heap; the effective value is logged at fork
startup. Tune it with a system property in forkedJvmArgs (plain bytes, no unit suffix;
⇐0 disables the budget and restores the per-object threshold):
"forkedJvmArgs": ["-Xmx1g", "-Dtika.pipes.cacheMemoryBudgetBytes=134217728"]
Size -Xmx with the budget in mind: the budget is additional heap the fork may use on top
of its parsing working set, and in per-client mode every fork holds its own budget
(numClients x budget in aggregate). In shared-server mode all concurrent parses in the
single forked server share one budget, so each in-flight document gets a smaller slice of
the same value.
Timeouts
See also Timeouts for the full timeout model.
| Field | Default | Description |
|---|---|---|
|
|
Maximum time (ms) to wait for data from a forked process. If no heartbeat or result is received within this window, the parse is considered hung. Also serves as the fork’s idle-shutdown timer: a fork that receives no work for this long exits and is restarted transparently on next use. |
|
|
Interval (ms) between heartbeats sent from the forked process. Must be significantly less than |
|
|
Socket read timeout (ms) applied to a freshly-connected fork until it completes its READY handshake, after which |
|
|
Ceiling for request-supplied timeout limits: a per-request |
|
|
Maximum time (ms) to wait for an available forked process when all are busy. |
Parse Behavior
| Field | Default | Description |
|---|---|---|
|
|
How embedded documents are handled: |
|
|
What to do when a parse fails: |
|
|
When |
Async / Emit Batching
These settings control how parsed results are batched before sending to emitters.
| Field | Default | Description |
|---|---|---|
|
|
Number of emitter threads. |
|
|
Size of the fetch/emit tuple queue. |
|
|
Flush the emit batch if nothing has been emitted within this many milliseconds, even if the batch is not full. |
|
|
Flush the emit batch when the estimated size reaches this many bytes. |
|
|
When |
IPC and Inline Payload Limits
| Field | Default | Description |
|---|---|---|
|
|
Maximum size in bytes of a single IPC message between the client and the fork. This limit is bidirectional: it applies both to parse results returned from the fork (FINISHED) and to requests sent from the client (NEW_REQUEST). Raising it lets very large documents pass over IPC; set the forked JVM |
|
|
Largest document carried inline to the fork instead of being written to a file first. A host that already holds the content (tika-server’s |
Emit Strategy
emitStrategy controls whether parsed extracts are emitted directly from the forked PipesServer or passed back to the parent process first. The default is balanced for typical workloads — tune only if you have a memory or throughput problem.
{
"pipes": {
"emitStrategy": {
"type": "DYNAMIC",
"thresholdBytes": 100000
}
}
}
| Field | Default | Description |
|---|---|---|
|
|
One of |
|
|
Only used when |
Distributed Config Store
For multi-host pipelines (e.g., shared-server clusters) you can store fetcher/emitter configuration in a distributed backend instead of memory. Most users should leave the defaults.
| Field | Default | Description |
|---|---|---|
|
|
Backend for storing fetcher/emitter configurations. |
|
|
JSON object (as a string) with backend-specific parameters. Structure depends on |
Shared Server Mode (Experimental)
| Field | Default | Description |
|---|---|---|
|
|
When |
See Shared Server Mode for details.
Complete examples
Worked-out end-to-end configs from the test tree, so the syntax stays current. The tests that cover them assert that the JSON parses; they do not instantiate every component, so a config can load cleanly and still name something that is missing from a release classpath.
Filesystem-to-filesystem pipeline
{
"content-handler-factory": {
"basic-content-handler-factory": {
"type": "TEXT",
"writeLimit": -1,
"throwOnWriteLimitReached": true
}
},
"fetchers": {
"fsf": {
"file-system-fetcher": {
"basePath": "FETCHER_BASE_PATH",
"extractFileSystemMetadata": false
}
}
},
"emitters": {
"fse": {
"file-system-emitter": {
"basePath": "EMITTER_BASE_PATH",
"fileExtension": "json",
"onExists": "EXCEPTION"
}
}
},
"pipes-iterator": {
"file-system-pipes-iterator": {
"basePath": "FETCHER_BASE_PATH",
"countTotal": true,
"fetcherId": "fsf",
"emitterId": "fse"
}
},
"pipes": {
"parseMode": "RMETA",
"onParseException": "EMIT",
"numClients": 4,
"emitIntermediateResults": "EMIT_INTERMEDIATE_RESULTS",
"forkedJvmArgs": ["-Xmx512m"],
"emitStrategy": {
"type": "DYNAMIC",
"thresholdBytes": 1000000
}
},
"auto-detect-parser": {
"throwOnZeroBytes": false
},
"parse-context": {
"mock-digester-factory": {},
"timeout-limits": {
"progressTimeoutMillis": 5000
}
},
"plugin-roots": "PLUGINS_PATHS"
}
Tokens (FETCHER_BASE_PATH, EMITTER_BASE_PATH, PLUGINS_PATHS, EMIT_INTERMEDIATE_RESULTS) are substituted by the test harness — replace them with real values in production configs. The first three are paths; EMIT_INTERMEDIATE_RESULTS is the boolean emitIntermediateResults flag.
Emit-all variant
{
"fetchers": {
"fsf": {
"file-system-fetcher": {
"basePath": "FETCHER_BASE_PATH",
"extractFileSystemMetadata": false
}
}
},
"emitters": {
"fse": {
"file-system-emitter": {
"basePath": "EMITTER_BASE_PATH",
"fileExtension": "json",
"onExists": "EXCEPTION"
}
}
},
"pipes": {
"numClients": 1,
"forkedJvmArgs": [
"-Xmx256m"
],
"emitStrategy": {
"type": "EMIT_ALL"
}
},
"parse-context": {
"timeout-limits": {
"progressTimeoutMillis": 60000
}
},
"plugin-roots": "PLUGINS_PATHS"
}
Shared-server (YOLO) mode
{
"content-handler-factory": {
"basic-content-handler-factory": {
"type": "TEXT",
"writeLimit": -1,
"throwOnWriteLimitReached": true
}
},
"fetchers": {
"fsf": {
"file-system-fetcher": {
"basePath": "FETCHER_BASE_PATH",
"extractFileSystemMetadata": false
}
}
},
"emitters": {
"fse": {
"file-system-emitter": {
"basePath": "EMITTER_BASE_PATH",
"fileExtension": "json",
"onExists": "REPLACE"
}
}
},
"pipes-iterator": {
"file-system-pipes-iterator": {
"basePath": "FETCHER_BASE_PATH",
"countTotal": true,
"fetcherId": "fsf",
"emitterId": "fse"
}
},
"pipes": {
"parseMode": "RMETA",
"onParseException": "EMIT",
"numClients": 4,
"useSharedServer": true,
"emitIntermediateResults": "EMIT_INTERMEDIATE_RESULTS",
"forkedJvmArgs": ["-Xmx512m"],
"emitStrategy": {
"type": "DYNAMIC",
"thresholdBytes": 1000000
}
},
"auto-detect-parser": {
"throwOnZeroBytes": false
},
"parse-context": {
"mock-digester-factory": {},
"timeout-limits": {
"progressTimeoutMillis": 5000
}
},
"plugin-roots": "PLUGINS_PATHS"
}
See Shared Server Mode for the trade-offs.
Tika Pipes config template
{
"content-handler-factory": {
"basic-content-handler-factory": {
"type": "TEXT",
"writeLimit": -1,
"throwOnWriteLimitReached": true
}
},
"parsers": [
{
"default-parser": {}
},
{
"pdf-parser": {
"extractActions": true,
"extractInlineImages": true,
"extractIncrementalUpdateInfo": true,
"parseIncrementalUpdates": true
}
},
{
"ooxml-parser": {
"includeDeletedContent": true,
"includeMoveFromContent": true,
"extractMacros": true
}
},
{
"office-parser": {
"extractMacros": true
}
}
],
"fetchers": {
"fsf": {
"file-system-fetcher": {
"basePath": "FETCHER_BASE_PATH",
"extractFileSystemMetadata": false
}
}
},
"emitters": {
"fse": {
"file-system-emitter": {
"basePath": "EMITTER_BASE_PATH",
"fileExtension": "json",
"onExists": "EXCEPTION"
}
}
},
"pipes-iterator": {
"file-system-pipes-iterator": {
"basePath": "FETCHER_BASE_PATH",
"countTotal": true,
"fetcherId": "fsf",
"emitterId": "fse"
}
},
"pipes": {
"parseMode": "RMETA"
},
"plugin-roots": "PLUGIN_ROOTS"
}
For per-plugin pipeline examples (S3, OpenSearch, JDBC, Kafka, etc.), see the relevant page under Plugins.