unpack-config: Extracting Embedded Document Bytes
When processing container files (ZIP, DOCX, PDF with attachments, etc.), you may want to
extract the raw bytes of embedded documents in addition to parsing them. The
unpack-config component (Java: UnpackConfig) controls how embedded bytes are
extracted and emitted.
Quick Start
To turn on byte extraction for every document the pipeline processes, set
parseMode to UNPACK in the pipes section of your tika-config.json.
That’s the minimum configuration — extraction defaults are fine for most cases.
{
"pipes": {
"parseMode": "UNPACK"
}
}
To tune extraction (size limits, naming, ZIP output, etc.), add an unpack-config
block under the top-level parse-context section. All the options listed below
live inside that block:
{
"pipes": {
"parseMode": "UNPACK"
},
"parse-context": {
"unpack-config": {
"maxUnpackBytes": 104857600,
"zipEmbeddedFiles": true
}
}
}
This extracts both metadata (like RMETA mode) and embedded document bytes.
|
You can also set |
Configuration Options
All options below are fields of the unpack-config block — nest them inside
parse-context.unpack-config as shown in the Quick Start.
| Property | Type | Default | Description |
|---|---|---|---|
|
String |
(from FetchEmitTuple) |
Emitter name for embedded bytes. Falls back to the FetchEmitTuple’s emitterId. |
|
long |
10 GiB |
Maximum total bytes to extract per file. Set to |
|
boolean |
|
Include the container document itself in the output. |
|
boolean |
|
Collect all embedded files into a single ZIP archive, emitted at the container’s emit key plus |
|
boolean |
unset: on for |
Metadata for every extracted file, in either format: one |
|
boolean |
— |
Deprecated alias for |
|
int |
|
Zero-pad embedded IDs in output names (e.g., |
|
NONE, EXISTING, DETECTED |
|
How to determine file extensions for extracted files. See Suffix Strategies. |
|
String |
|
Separator between |
|
DEFAULT, CUSTOM |
|
Strategy for generating emit keys. See Key Base Strategies. |
|
String |
|
Custom base path when |
|
REGULAR, FRICTIONLESS |
|
Output format for the ZIP archive. See Frictionless Data Package Output. |
|
ZIPPED, DIRECTORY |
|
|
|
boolean |
— |
Deprecated alias for |
Examples
ZIP Output with Metadata
Collect all embedded files into a ZIP with metadata:
{
"pipes": {
"parseMode": "UNPACK"
},
"parse-context": {
"unpack-config": {
"zipEmbeddedFiles": true,
"includeMetadata": true,
"includeOriginal": true
}
}
}
Custom Naming
embeddedIdPrefix only applies to keyBaseStrategy=CUSTOM, so set both:
{
"pipes": {
"parseMode": "UNPACK"
},
"parse-context": {
"unpack-config": {
"zeroPadName": 8,
"suffixStrategy": "DETECTED",
"keyBaseStrategy": "CUSTOM",
"emitKeyBase": "document",
"embeddedIdPrefix": "-embed-"
}
}
}
Produces names like document-embed-00000001.pdf. Under the DEFAULT strategy the same
zeroPadName/suffixStrategy settings would produce <containerKey>-embed/00000001.pdf.
Suffix Strategies
NONE-
No file extension added to extracted files.
EXISTING-
Use the file extension from the embedded document’s resource name.
DETECTED-
Use the file extension based on the detected MIME type.
Key Base Strategies
DEFAULT-
Output key is
{containerKey}-embed/{id}{suffix}. The-embed/separator is fixed;embeddedIdPrefixis not consulted. CUSTOM-
Output key is
{emitKeyBase}{embeddedIdPrefix}{id}{suffix}.
Safety Limits
maxUnpackBytes bounds zip bombs and other files that expand to enormous sizes. The 10 GiB
default suits most corpora; lower it for untrusted input.
Hitting the limit is silent apart from a log line. The embedded file being written is
truncated at the remaining budget, each later embedded file is skipped, and each case
logs a WARN — but nothing is stamped on the metadata and the result status is unchanged
(PARSE_SUCCESS / EMIT_SUCCESS). Watch the log, not the status, for truncation.
maxUnpackBytes: -1 (or any negative value) disables the limit — not recommended for
untrusted input. 0 is not "unlimited": it caps extraction at zero bytes.
Frictionless Data Package Output
The UNPACK mode can output files in Frictionless Data Package format,
a standard for packaging data files with their metadata. This format includes a datapackage.json
manifest with file checksums and MIME types, making it easy to verify and process extracted files.
Enabling Frictionless Output
Set outputFormat to FRICTIONLESS in your unpack-config:
{
"pipes": {
"parseMode": "UNPACK"
},
"parse-context": {
"unpack-config": {
"outputFormat": "FRICTIONLESS",
"includeMetadata": true
}
}
}
Output Structure
When using Frictionless output format, the ZIP archive contains:
output.zip
├── datapackage.json # Manifest with file list, SHA256 hashes, mimetypes
├── metadata.json # Full RMETA metadata (unless includeMetadata=false)
└── unpacked/
├── 00000001.pdf
├── 00000002.png
└── ...
The tk:content field inside metadata.json (and inside the per-file .metadata.json
sidecars includeMetadata adds to REGULAR zip output) is Markdown by default. On tika-server
change it with /unpack/all/{handlerType} or a content-handler-factory in the config
part; in a config, the factory alone. Each entry records the handler it was written with in
tk:content-handler-type:
{
"unpack-config": { "outputFormat": "FRICTIONLESS" },
"basic-content-handler-factory": { "type": "XML" }
}
To get metadata with no extracted text at all — one sidecar per file, but no tk:content — select the IGNORE handler. On the test corpus below this drops the sidecars from 20,257 to
11,132 bytes and is the way to avoid shipping every embedded document’s text alongside its
bytes:
{
"unpack-config": { "includeMetadata": true },
"basic-content-handler-factory": { "type": "IGNORE" }
}
On tika-server the handler is also a path segment on /unpack/all/{handlerType}, which
needs no config part and so no allowPerRequestConfig. Plain /unpack carries no
metadata for a handler to render, so it takes no segment.
curl -T container.docx http://localhost:9998/unpack/all/ignore # files + metadata, no text
Naming the handler in both the path and a config part is a 400, as on /tika and
/rmeta.
What each surface produces
The five knobs interact differently per surface. One default changes from 4.0.0: a
Frictionless package carries metadata.json unless includeMetadata is false (4.0.0 left
it out unless includeFullMetadata was true). REGULAR output is as it was.
tika-server /unpack |
tika-app -z / -Z |
pipes JSON config | |
|---|---|---|---|
format |
REGULAR; set |
REGULAR; |
|
packaging |
always one zip ( |
loose files; |
|
depth |
|
|
|
metadata |
Frictionless: by default; REGULAR: |
Frictionless: by default; REGULAR zip: |
|
source document’s bytes |
|
only via |
|
handler for |
|
|
|
/unpack/all means the source document’s bytes plus metadata in both formats. Under REGULAR
that is where the per-file sidecars come from; under Frictionless metadata.json is already
there, so /all adds only the source document. outputMode: DIRECTORY is refused on
/unpack: an HTTP response is one zip, so a request asking for it gets a 400, and a server
whose own unpack-config says so fails at startup if the unpack endpoint is enabled — use
/pipes or /async for per-file emission.
Output trees for doc.docx with three attachments:
REGULAR, tika-server (one zip) REGULAR, tika-app -Z (loose)
1.emf out/doc.docx.json
2.zip out/doc.docx-embed/1.emf
3.bin out/doc.docx-embed/2.zip
+ with /all: 0.docx and one out/doc.docx-embed/3.bin
<name>.metadata.json per file
FRICTIONLESS, tika-server (one zip) FRICTIONLESS, tika-app -Z --unpack-format=FRICTIONLESS
datapackage.json out/doc.docx.json
metadata.json out/doc.docx/datapackage.json
unpacked/1.emf ... out/doc.docx/metadata.json
+ with /all: unpacked/0.docx out/doc.docx/unpacked/1.emf ...
The top-level doc.docx.json on the command line is the normal pipes output for the parse;
the package’s metadata.json carries the same rows again so it is self-contained. Set
includeMetadata: false in a -c file to drop it.
The datapackage.json file contains:
-
List of all extracted files as "resources"
-
SHA256 hash for each file
-
MIME type for each file
-
File size in bytes
CLI Usage
Extract files in Frictionless format using the CLI. The -Z flag turns on recursive
unpack (the Pipes-mode counterpart of standard-mode -z), and -i/-o are the
Pipes input/output directories:
java -jar tika-app.jar -Z --unpack-format=FRICTIONLESS -i /path/to/input -o /path/to/output
-i expects a directory of containers to unpack, not a single file. For
one-off unpacking of a single document, see the standard-mode -z/--extract
flag — though as of 4.x that path also routes through the Pipes machinery and
expects an input directory.
|
Code Examples
For working code examples, see:
-
tika-pipes/tika-pipes-integration-tests/src/test/java/org/apache/tika/pipes/core/UnpackModeTest.java -
tika-server/tika-server-standard/src/test/java/org/apache/tika/server/standard/TikaPipesTest.java
These test files demonstrate all configuration options with assertions.