unpack-config: Extracting Embedded Document Bytes

When processing container files (ZIP, DOCX, PDF with attachments, etc.), you may want to extract the raw bytes of embedded documents in addition to parsing them. The unpack-config component (Java: UnpackConfig) controls how embedded bytes are extracted and emitted.

Quick Start

To turn on byte extraction for every document the pipeline processes, set parseMode to UNPACK in the pipes section of your tika-config.json. That’s the minimum configuration — extraction defaults are fine for most cases.

{
  "pipes": {
    "parseMode": "UNPACK"
  }
}

To tune extraction (size limits, naming, ZIP output, etc.), add an unpack-config block under the top-level parse-context section. All the options listed below live inside that block:

{
  "pipes": {
    "parseMode": "UNPACK"
  },
  "parse-context": {
    "unpack-config": {
      "maxUnpackBytes": 104857600,
      "zipEmbeddedFiles": true
    }
  }
}

This extracts both metadata (like RMETA mode) and embedded document bytes.

You can also set UnpackConfig programmatically per request from Java code by calling parseContext.set(UnpackConfig.class, …​) on the ParseContext attached to your FetchEmitTuple. The JSON parse-context section above is the declarative equivalent.

Configuration Options

All options below are fields of the unpack-config block — nest them inside parse-context.unpack-config as shown in the Quick Start.

Property Type Default Description

emitter

String

(from FetchEmitTuple)

Emitter name for embedded bytes. Falls back to the FetchEmitTuple’s emitterId.

maxUnpackBytes

long

10 GiB

Maximum total bytes to extract per file. Set to -1 for unlimited (not recommended). 0 is not unlimited — it means zero bytes, so the first embedded file’s extraction is immediately capped.

includeOriginal

boolean

false

Include the container document itself in the output.

zipEmbeddedFiles

boolean

false

Collect all embedded files into a single ZIP archive, emitted at the container’s emit key plus -embedded.zip.

includeMetadata

boolean

unset: on for FRICTIONLESS, off for REGULAR

Metadata for every extracted file, in either format: one metadata.json under Frictionless (a data package without the parse is half a package, so it is on unless you set false), a .metadata.json sidecar per file under REGULAR zip output (opt-in; tika-server’s /unpack/all sets it). REGULAR loose output (no zip) writes none.

includeMetadataInZip

boolean

 — 

Deprecated alias for includeMetadata (removal in 5.0). Still accepted so 4.0.0 configs load.

zeroPadName

int

0

Zero-pad embedded IDs in output names (e.g., 8 produces 00000001).

suffixStrategy

NONE, EXISTING, DETECTED

NONE

How to determine file extensions for extracted files. See Suffix Strategies.

embeddedIdPrefix

String

"-"

Separator between emitKeyBase and the embedded ID. Read only when keyBaseStrategy=CUSTOM; the DEFAULT strategy uses a fixed -embed/ separator instead.

keyBaseStrategy

DEFAULT, CUSTOM

DEFAULT

Strategy for generating emit keys. See Key Base Strategies.

emitKeyBase

String

""

Custom base path when keyBaseStrategy=CUSTOM.

outputFormat

REGULAR, FRICTIONLESS

REGULAR

Output format for the ZIP archive. See Frictionless Data Package Output.

outputMode

ZIPPED, DIRECTORY

ZIPPED

ZIPPED packages everything into one archive; DIRECTORY emits each extracted file to the emitter as its own item.

includeFullMetadata

boolean

 — 

Deprecated alias for includeMetadata (removal in 5.0). Still accepted so 4.0.0 configs load.

Examples

ZIP Output with Metadata

Collect all embedded files into a ZIP with metadata:

{
  "pipes": {
    "parseMode": "UNPACK"
  },
  "parse-context": {
    "unpack-config": {
      "zipEmbeddedFiles": true,
      "includeMetadata": true,
      "includeOriginal": true
    }
  }
}

Custom Naming

embeddedIdPrefix only applies to keyBaseStrategy=CUSTOM, so set both:

{
  "pipes": {
    "parseMode": "UNPACK"
  },
  "parse-context": {
    "unpack-config": {
      "zeroPadName": 8,
      "suffixStrategy": "DETECTED",
      "keyBaseStrategy": "CUSTOM",
      "emitKeyBase": "document",
      "embeddedIdPrefix": "-embed-"
    }
  }
}

Produces names like document-embed-00000001.pdf. Under the DEFAULT strategy the same zeroPadName/suffixStrategy settings would produce <containerKey>-embed/00000001.pdf.

Suffix Strategies

NONE

No file extension added to extracted files.

EXISTING

Use the file extension from the embedded document’s resource name.

DETECTED

Use the file extension based on the detected MIME type.

Key Base Strategies

DEFAULT

Output key is {containerKey}-embed/{id}{suffix}. The -embed/ separator is fixed; embeddedIdPrefix is not consulted.

CUSTOM

Output key is {emitKeyBase}{embeddedIdPrefix}{id}{suffix}.

Safety Limits

maxUnpackBytes bounds zip bombs and other files that expand to enormous sizes. The 10 GiB default suits most corpora; lower it for untrusted input.

Hitting the limit is silent apart from a log line. The embedded file being written is truncated at the remaining budget, each later embedded file is skipped, and each case logs a WARN — but nothing is stamped on the metadata and the result status is unchanged (PARSE_SUCCESS / EMIT_SUCCESS). Watch the log, not the status, for truncation.

maxUnpackBytes: -1 (or any negative value) disables the limit — not recommended for untrusted input. 0 is not "unlimited": it caps extraction at zero bytes.

Frictionless Data Package Output

The UNPACK mode can output files in Frictionless Data Package format, a standard for packaging data files with their metadata. This format includes a datapackage.json manifest with file checksums and MIME types, making it easy to verify and process extracted files.

Enabling Frictionless Output

Set outputFormat to FRICTIONLESS in your unpack-config:

{
  "pipes": {
    "parseMode": "UNPACK"
  },
  "parse-context": {
    "unpack-config": {
      "outputFormat": "FRICTIONLESS",
      "includeMetadata": true
    }
  }
}

Output Structure

When using Frictionless output format, the ZIP archive contains:

output.zip
├── datapackage.json      # Manifest with file list, SHA256 hashes, mimetypes
├── metadata.json         # Full RMETA metadata (unless includeMetadata=false)
└── unpacked/
    ├── 00000001.pdf
    ├── 00000002.png
    └── ...

The tk:content field inside metadata.json (and inside the per-file .metadata.json sidecars includeMetadata adds to REGULAR zip output) is Markdown by default. On tika-server change it with /unpack/all/{handlerType} or a content-handler-factory in the config part; in a config, the factory alone. Each entry records the handler it was written with in tk:content-handler-type:

{
  "unpack-config": { "outputFormat": "FRICTIONLESS" },
  "basic-content-handler-factory": { "type": "XML" }
}

To get metadata with no extracted text at all — one sidecar per file, but no tk:content — select the IGNORE handler. On the test corpus below this drops the sidecars from 20,257 to 11,132 bytes and is the way to avoid shipping every embedded document’s text alongside its bytes:

{
  "unpack-config": { "includeMetadata": true },
  "basic-content-handler-factory": { "type": "IGNORE" }
}

On tika-server the handler is also a path segment on /unpack/all/{handlerType}, which needs no config part and so no allowPerRequestConfig. Plain /unpack carries no metadata for a handler to render, so it takes no segment.

curl -T container.docx http://localhost:9998/unpack/all/ignore   # files + metadata, no text

Naming the handler in both the path and a config part is a 400, as on /tika and /rmeta.

What each surface produces

The five knobs interact differently per surface. One default changes from 4.0.0: a Frictionless package carries metadata.json unless includeMetadata is false (4.0.0 left it out unless includeFullMetadata was true). REGULAR output is as it was.

tika-server /unpack tika-app -z / -Z pipes JSON config

format

REGULAR; set outputFormat in the server config for Frictionless

REGULAR; --unpack-format=FRICTIONLESS

outputFormat

packaging

always one zip (ZIPPED is pinned; an HTTP response is one body)

loose files; DIRECTORY when Frictionless is asked for without --unpack-mode

outputMode governs Frictionless, zipEmbeddedFiles governs REGULAR

depth

embedded-limits in config

-z = direct attachments, -Z = recursive

embedded-limits

metadata

Frictionless: by default; REGULAR: /unpack/all

Frictionless: by default; REGULAR zip: --unpack-include-metadata

includeMetadata (unset: on for Frictionless, off for REGULAR)

source document’s bytes

/unpack/all

only via -c (unpack-config.includeOriginal)

includeOriginal

handler for tk:content

/unpack/all/{handlerType} or the config part

--handler

content-handler-factory

/unpack/all means the source document’s bytes plus metadata in both formats. Under REGULAR that is where the per-file sidecars come from; under Frictionless metadata.json is already there, so /all adds only the source document. outputMode: DIRECTORY is refused on /unpack: an HTTP response is one zip, so a request asking for it gets a 400, and a server whose own unpack-config says so fails at startup if the unpack endpoint is enabled — use /pipes or /async for per-file emission.

Output trees for doc.docx with three attachments:

REGULAR, tika-server (one zip)        REGULAR, tika-app -Z (loose)
  1.emf                                 out/doc.docx.json
  2.zip                                 out/doc.docx-embed/1.emf
  3.bin                                 out/doc.docx-embed/2.zip
  + with /all: 0.docx and one           out/doc.docx-embed/3.bin
    <name>.metadata.json per file

FRICTIONLESS, tika-server (one zip)   FRICTIONLESS, tika-app -Z --unpack-format=FRICTIONLESS
  datapackage.json                      out/doc.docx.json
  metadata.json                         out/doc.docx/datapackage.json
  unpacked/1.emf ...                    out/doc.docx/metadata.json
  + with /all: unpacked/0.docx          out/doc.docx/unpacked/1.emf ...

The top-level doc.docx.json on the command line is the normal pipes output for the parse; the package’s metadata.json carries the same rows again so it is self-contained. Set includeMetadata: false in a -c file to drop it.

The datapackage.json file contains:

  • List of all extracted files as "resources"

  • SHA256 hash for each file

  • MIME type for each file

  • File size in bytes

CLI Usage

Extract files in Frictionless format using the CLI. The -Z flag turns on recursive unpack (the Pipes-mode counterpart of standard-mode -z), and -i/-o are the Pipes input/output directories:

java -jar tika-app.jar -Z --unpack-format=FRICTIONLESS -i /path/to/input -o /path/to/output
-i expects a directory of containers to unpack, not a single file. For one-off unpacking of a single document, see the standard-mode -z/--extract flag — though as of 4.x that path also routes through the Pipes machinery and expects an input directory.

Code Examples

For working code examples, see:

  • tika-pipes/tika-pipes-integration-tests/src/test/java/org/apache/tika/pipes/core/UnpackModeTest.java

  • tika-server/tika-server-standard/src/test/java/org/apache/tika/server/standard/TikaPipesTest.java

These test files demonstrate all configuration options with assertions.