Apache Tika 4.1.0
Tika 4.1.0 is the first minor release of the 4.x line. The most notable changes over 4.0.0 are summarised below; the CHANGES-4.1.0.txt file carries the complete list. Please read "Breaking Changes and Deprecations" before upgrading, especially if you run tika-pipes plugins or the Kafka pipes iterator.
Upgrading from 3.x? Start with the migration guides. The 3.x line remains supported. See the download page for both.
Highlights
- OCR and enrichment engines are selected by name in a new "text-recognizers" list ("[]" turns enrichment off; no list resolves one engine per media type at startup); the image and PDF parsers invoke them rather than the composite dispatching to them. The image/ocr-* pseudo-types are retired; a third-party engine still advertising them is a legacy text recognizer with a WARN. VLM parsers gain "textRecognizer" (TIKA-4872, TIKA-4884).
- New "exception-reporting" parse-context config redacts and bounds exception text in metadata, tika-server error bodies and pipes/grpc messages; FileSystemEmitter writes atomically. Compat: tk:exception:* values may now begin with the TikaException wrapper line (TIKA-4848).
Performance
- Detection (magic ~35% faster on unmatched input, override keys honored before magic, cached type->parser map), markdown output ~30x faster, .doc cleanup, CSV sniffing, zip legacy-method rewinds, pipes ACK overlap, raw UTF-8 content passback for tika-server (content-bytes-config). Compat: with a Content-Type override set, DefaultDetector no longer lets a more specific magic result overrule it (TIKA-4868).
- PDF incremental-update scanning reads in blocks; ~4x less CPU (TIKA-4898).
- Embedded documents in Office files, PDFs, PST and the zip-family containers are re-opened from their container on rewind instead of cached (TIKA-4878).
- Improve spooling/decrease number of spills to disk. The pipes cache memory budget defaults to a quarter of the fork heap (-Dtika.pipes.cacheMemoryBudgetBytes overrides) (TIKA-4835).
- Digesting embedded documents no longer spools each to a temp file; a process-wide CacheMemoryBudget keeps them in memory. Zip-family parsing uses seekable channels, so hasFile() may be false afterward (TIKA-4828).
Breaking Changes and Deprecations
- Pipes IPC carries inline bytes as a raw binary field, not in the tuple; 4.0.0 tuples with an "inline-bytes" parse-context entry are rejected (TIKA-4829).
- Pipes plugins no longer bundle Jackson; plugin config is parsed by a shared strict mapper that rejects unknown and duplicate keys (TIKA-4840).
- OpenAIVLMParser no longer auto-registers via SPI; per-request config for the embedding filters works and locks baseUrl/apiKey/model; inline PDF page OCR accumulates tk:chunks from every page (TIKA-4871).
- Temp files follow -Djava.io.tmpdir on the parent JVM (Tika, its libraries and forks); pipes.tempDirectory is deprecated. A fork whose parent dies deletes its own temp dir; failure-path temp leaks fixed (TIKA-4877).
- Kafka pipes iterator no longer stops on the first empty poll (assignmentTimeoutMs, drainIdleMs); groupInitialRebalanceDelayMs is deprecated (TIKA-4833).
Inference (Experimental)
- tika-inference and tika-vlm are still experimental: classes, config keys and output may change in minor releases. Stable: the "engines" map, the "text-recognizers" list and pdf-parser "text" (TIKA-4884, TIKA-4895).
- Refactor OCR and inference to use engines for inference and page rendering (TIKA-4889).
- Enable audio/video chunking with ffmpeg and inference; the default segment grid is 25 s windows with 5 s overlap (TIKA-4900).
Server, Pipes and Operations
- tika-server: named configuration presets, and catalog presets that are inert until the config names them (TIKA-4856).
- Add Micrometer reporting and opt-in endpoint for tika-server (TIKA-4839).
- Enable environment variable interpolation in tika-config.json (TIKA-4909).
- tika-server and tika-async-cli accept // and /* */ comments in config during override merging (TIKA-4834).
- New file-system-jsonl-reporter pipes reporter records a batch run's crashes; tika-eval Profile/Compare read that ledger and the run-info json to classify NO_EXTRACT_FILE by cause (TIKA-4846, TIKA-4847).
- PipesForkParser no longer drops javaPath and socketTimeoutMillis set in code; the fork ran java from the PATH with a 60s socket timeout regardless (TIKA-4931).
- Pipes no longer copies the caller's Content-Type hint back over the detected type on output (TIKA-4932).
Extraction and Formats
- Improve extraction of tagged PDFs. New tika-eval-structure tool compares two sets of XHTML extracts block by block (TIKA-4891).
- OneNote: document-order extraction, superseded revisions omitted, embedded BLOBs extracted, bounded recursion; malformed files fall back to the legacy string dump via Henry Lindeman (TIKA-4814).
- Raw camera formats (NEF/NRW, PEF/PTX, ARW/SRF/SR2, SRW, DNG, RAF, RW2, MRW, ORF) are detected by content instead of as image/tiff via Dominik Schmidt (TIKA-4861).
- AVIF images are parsed by HeifParser (dimensions, EXIF, XMP) via Dominik Schmidt (TIKA-4870).
- New GeoGebraParser for *.ggb/*.ggs/*.ggt with content-based detection and THUMBNAIL embedded documents via Dominik Schmidt (TIKA-4831).
- The video of a Google/Android motion photo is emitted as an ATTACHMENT embedded document via Dominik Schmidt (TIKA-4869).
- RawTiffParser extracts embedded JPEG previews as THUMBNAIL embedded documents; image/x-raw-* are now subtypes of image/tiff ("extractPreviews": false disables) via Dominik Schmidt (TIKA-4824).
- Apple Mail emlx files are detected as message/x-emlx and parsed by RFC822Parser; a file name no longer turns plain text into a message type once the type's magic has rejected the bytes (TIKA-4890).
- Audio cover art (ID3/FLAC/Vorbis front cover, MP4 covr) is emitted as a THUMBNAIL embedded document rather than INLINE via Dominik Schmidt (TIKA-4850).
- EpubParser emits the OPF cover image as a THUMBNAIL embedded document via Dominik Schmidt (TIKA-4852).
- DWGReadParser emits THUMBNAILIMAGE as THUMBNAIL instead of INLINE via Dominik Schmidt (TIKA-4853).
- iWork '09 and '18 preview images are emitted as THUMBNAIL embedded documents via Dominik Schmidt (TIKA-4854).
- New poi-metafile-renderer rasterizes EMF/WMF (opt-in "renderImage"); OfficeParser emits the OLE2 SummaryInformation thumbnail as a THUMBNAIL embedded document ("extractThumbnail": false disables) via Dominik Schmidt (TIKA-4855).
Other Changes
- Ogg, Vorbis, Opus and FLAC comment fields are read independently of the JVM's default locale: on a Turkish JVM TITLE and ARTIST were silently lost. The OCR image-rotation value is formatted the same way (TIKA-4921).
- The Ignite config store quotes its table name, so lookups work under any default locale (TIKA-4922).
- tika-annotation-processor is no longer a transitive dependency of tika-pipes-core or the tika-langdetect modules (provided scope).
- Setting progressTimeoutMillis equal to totalTaskTimeoutMillis, the documented way to disable stall detection, no longer logs a warning on every parse (TIKA-4919).
- Retry a 429, 502, 503 or 504 answer from a hosted inference engine with a jittered backoff (TIKA-4912).
- Mp3Parser no longer writes "null" (a missing album) or the duration into the body (TIKA-4907).
- VLM parsers emit a raw HTML block in the model's markdown as XHTML (TIKA-4906).
- The forked-worker CPU cap (-XX:ActiveProcessorCount) is clamped at its floor of 2 (TIKA-4905).
- Tesseract script models are accepted by their bare name ("Latin", "eng+Japanese_vert"); per-request tesseract-ocr-parser config no longer accepts otherTesseractConfig (TIKA-4903, TIKA-4904).
- ICC tone-curve and lookup-table tags are no longer written to icc:* by default; "includeIccCurvesAndLuts": true on the image parsers restores them. Curve values are formatted independently of the JVM's default locale, and single-entry gamma curves decode correctly (TIKA-4902).
- Metadata never stores an unpaired UTF-16 surrogate (TIKA-4897).
- unpack-config gains includeMetadata; tika-server /rmeta gains config/{handlerType} (TIKA-4881).
- PDF bookmark outlines are walked without recursion; lists nest at most 50 deep (TIKA-4894).
- The image embedder writes a picture's vector onto the document the picture appears in, with an "embedded" locator; liftToParent: false keeps it on the picture (TIKA-4888).
- A PDF page rendered as an embedded document is enriched once, by the PDF parser; NO_OCR now means no page OCR at all (TIKA-4887).
- PDF AUTO OCR no longer emits a page's extracted text next to the OCR output that superseded it; AUTO without a text recognizer behaves as NO_OCR. (TIKA-4883).
- tika-app -m and --json no longer drop metadata keys added after endDocument (TIKA-4885).
- A missing OOXML relationship target (xlsx threaded comments, vsdx pages) no longer aborts the whole file (TIKA-4879).
- tika-eval Profile/Compare speedups: langdetect preprocessing, H2 page cache sizing, a per-interval rate in the status log (TIKA-4875).
- Embedded documents carry Content-Length where the stream knows it via Dominik Schmidt (TIKA-4873).
- The default plugins directory (tika-server, PipesForkParser, async CLI) and tika-grpc's plugin-roots fallback are resolved against the install layout instead of the working directory via Dominik Schmidt (TIKA-4864, TIKA-4865).
- Allow image compression settings in PDFBox-based renderer (TIKA-4862).
- The tika-server full and tika-grpc Docker images add ffmpeg and fonts-noto-cjk, and set OMP_THREAD_LIMIT=1 to avoid oversubscribing the CPU under forked parse workers (TIKA-4910, TIKA-4866, TIKA-4863).
- embedded-limits maxDepth counts embedding levels again; values above 1 used to stop one level early via Dominik Schmidt (TIKA-4857).
- Enum values in JSON configuration are matched case-insensitively via Dominik Schmidt (TIKA-4859).
- Per-request config for parsers that lock fields (Tess4J, VLM, OpenAI image-embedding) threw even when empty (TIKA-4843).
- OOXML: new msoffice:has-unreferenced-parts and msoffice:unreferenced-part-names (structural only; not XPS) (TIKA-4837).
- Shared pipes server: a worker death could trigger a spurious second restart; forks now carry a generation (TIKA-4844).
- Docs and javadocs reconciled with the code; fixes an extractFontNames NullPointerException on a PDF page with no /Resources (TIKA-4842).
- Pipes carries the client Content-Type into the forked worker as a detection hint for every forked endpoint via Dominik Schmidt (TIKA-4825).
- MP4 audio and video track codecs are exposed as FourCCs (audio:fourcc, video:fourcc) via Dominik Schmidt (TIKA-4838).
- FilenameUtils returns the default value for an unknown extension via Tim Grein (TIKA-4882).
- ProcessUtils does not start a subprocess when the granted timeout is <= 0 via Tim Grein (TIKA-4886).
- Remove the unnecessary jaxb-runtime dependency from the pdf and miscoffice modules via Thorsten Heit (TIKA-4849).
Contributors
The following people have contributed to Tika 4.1.0 -- via code, patches, pull requests, bug reports, review and discussion. Thank you!
- Dominik Schmidt
- Gary D. Gregory
- Henry Lindeman
- Jarek Potiuk
- Lewis John McGibbney
- Luca Foppiano
- Nicholas DiPiazza
- Thorsten Heit
- TianHengZhuang
- Tilman Hausherr
- Tim Allison
- Tim Grein
- Tim Scheckenbach
- Vasiliy Mikhailov
See https://s.apache.org/xp719 for more details on these contributions.


