Configuration

Tika 4.x is configured with a JSON file — parsers, detectors, content handlers, server behavior and the Tika Pipes pipeline. // and /* */ comments are allowed.

Converting a 3.x tika-config.xml? See the Migration Guide, and tika-app’s --convert-config-xml-to-json` flag.

Top-level JSON structure

A tika-config.json is a single JSON object whose keys are the sections below. Every section is optional; anything you omit uses its defaults.

{
  "parsers": [ /* parser declarations */ ],
  "detectors": [ /* detector declarations */ ],
  "encoding-detectors": [ /* encoding detector declarations */ ],
  "metadata-filters": [ /* metadata filter declarations */ ],
  "renderers": [ /* page renderer declarations */ ],
  "text-recognizers": [ /* OCR engines etc., selected by name; see below */ ],
  "translator": { /* translator declaration */ },
  "content-handler-factory": { /* handler type for emitted content */ },
  "auto-detect-parser": { /* AutoDetectParser options */ },
  "parse-context": {
    "timeout-limits": { /* progress + total task timeouts */ },
    "exception-reporting": { /* how much exception detail is reported */ },
    "unpack-config": { /* embedded-byte extraction */ },
    "pages": { /* rendering, OCR policy and page images for every parser that renders; see below */ }
    /* other SelfConfiguring components by component name */
  },
  "server": { /* tika-server options: allowPipes, allowPerRequestConfig, cors, tlsConfig, ... */ },
  "grpc": { /* tika-grpc options */ },
  "pipes": { /* Pipes process management: numClients, parseMode, ... */ },
  "fetchers": { /* named fetcher instances */ },
  "emitters": { /* named emitter instances */ },
  "pipes-iterator": { /* iterator (one per pipeline) */ },
  "pipes-reporters": { /* per-document status reporters */ },
  "plugin-roots": "/path/to/plugins",
  "metadata-list": { /* Jackson read limits for metadata JSON */ },
  "service-loader": { /* SPI load-failure policy */ },
  "xml-reader-utils": { /* XML parser pool size */ }
}

That list is exhaustive: an unrecognized top-level key is fatal at load time, and the error names the valid ones.

parsers, detectors, encoding-detectors, metadata-filters, content-handler-factory and parse-context are covered under Topics below. For server see Tika Server; for grpc see Tika gRPC; for pipes, fetchers, emitters, pipes-iterator, pipes-reporters and plugin-roots see Pipes Configuration.

Where a typo is caught, and where it is not

Unknown keys are rejected almost everywhere: at the top level, inside a component’s config body, and inside a FetchEmitTuple. Four places swallow them silently instead, so a misspelling there produces no error and no effect:

  • the metadata-list body — only maxStringLength, maxNestingDepth and maxNumberLength are read;

  • the body of a self-configuring component under parse-context — it is handed to the component untouched;

  • _mime-include / _mime-exclude when the value is neither a JSON array nor a string — a bare string is read as a one-element list, but any other shape yields an empty filter rather than a failure;

  • a non-textual entry inside default-parser’s `exclude list — the list expects plain strings, so an object or array entry such as [{"pdf-parser": {}}] is skipped and the component stays enabled. (An unregistered name in that list, and an unknown key in that same block, are both fatal.)

If a setting appears to do nothing, check its spelling against one of those four first.

Component names

Everything you configure is referenced by its component name — a kebab-case identifier such as pdf-parser, file-system-fetcher or unpack-config, derived from the class’s simple name unless the component declares its own. (developers/ calls the derived form a friendly name; it is the same string.) Class names are not accepted in their place: only components registered with @TikaComponent can be instantiated from JSON, which is what stops a config file from loading arbitrary classes. tika-app --list-parser-names prints the registered parser names, and an unknown name fails at load time listing the ones that are registered.

The parsers list and default-parser

A parsers list loads only the parsers it names — every other parser is dropped. To customize one parser while keeping all the others, add a default-parser entry:

{
  "parsers": [
    { "pdf-parser": { "sortByPosition": true } },
    { "default-parser": {} }
  ]
}

Configuring a parser automatically excludes its default copy, so there is no duplication. Omit default-parser only when you want a Tika limited to the parsers you listed.

detectors works the same way, with a default-detector entry. encoding-detectors has a default-encoding-detector, but it must not be mixed with explicit detector entries — see Encoding Detectors.

The text-recognizers list (4.1.0+)

Text recognizers are engines that produce a document’s text from its pixels: an OCR engine run on an embedded image or a rendered PDF page, a VLM prompted to transcribe. A container parser invokes them on bytes it has already parsed. Two recipes cover almost every deployment.

This list and the engines map it names are stable: an entry keeps loading and meaning "this engine recognizes text for these types", and if 4.2 changes the spelling the 4.1 form loads as an alias for at least one minor release. What may change is when the engine is called and where its text lands: 4.2 batches recognition per document, so text may be placed after the walk rather than as each image is met. The Java contract behind an engine (TextRecognizer) is experimental, with batching meant as an opt-in addition an engine that takes one image at a time never sees; see Wiring in Your Own OCR Engine.

Choose the engine. Configure it once under engines and name it here; every default parser stays loaded, and image-parser and pdf-parser call it:

{
  "engines": {
    "tesseract": { "tesseract-ocr-parser": { "language": "eng" } }
  },
  "text-recognizers": [
    { "engine": "tesseract" }
  ]
}

An entry naming an engine takes only engine, _mime-include, _mime-exclude, _min-width and _min-height; the engine’s own settings live under engines, where an inference binding can name the same engine. The 4.0 form, the engine inline, still works: { "tesseract-ocr-parser": { "language": "eng" } }, and takes the same underscore keys.

_min-width/_min-height keep an image whose recorded size (tiff:ImageWidth, tiff:ImageLength) falls short in either dimension away from the engine, in pixels; an image of unknown size is sent. Whatever the entry says, an image under 2 x 2 (a spacer in an email signature) never reaches any engine; the engines' own byte gates (minFileSizeToOcr) are unchanged.

Turn OCR off. An empty list is authoritative:

{ "text-recognizers": [] }

Turn it off for one request. The list stays configured; the request switches it off, in every container, and no engine is missed as an error:

{ "parse-context": { "text-recognizers": { "enabled": false } } }

Engine types: tesseract-ocr-parser, tess4j-parser, openai-vlm-parser, claude-vlm-parser, gemini-vlm-parser. At startup Tika logs one line per engine with the media types it enriches, so the effective engine is never a guess.

How the list behaves

A recognizer advertises its real media types (image/png, …​) and does not compete with the parser registered for those types: image-parser still parses the image and calls the recognizer. An entry’s _mime-include/_mime-exclude name real types too (image/tiff); the pre-4.1 image/ocr- spellings are refused at load. *Every entry matching a media type runs, in the order listed. Two recognizers on one type both run and their text is concatenated; a WARN at startup names them, since that is almost never intended. The list is authoritative: a media type no configured recognizer matches gets no OCR, never a classpath engine you did not name. A named engine that reports no media types at startup (missing native binary, unreachable inference server) fails config load rather than going silently inert; that check is made against the engine itself, so a _mime-include list cannot stand in for an unreachable engine. Failures are best-effort: one recognizer failing does not stop the others, and every failure is still reported through the parser’s normal exception handling (timeouts abort the chain immediately).

With no list

When the key is absent, the loader resolves one engine per media type from the recognizers among the loaded parsers and injects that, so Tesseract on the classpath OCRs images with no configuration at all. The rules: an engine configured under parsers beats one default-parser found on the classpath; among those the last registered wins (user-supplied classes register after Tika’s); an entry’s _mime-include/_mime-exclude applies, and on the default-parser entry it applies to every engine inside it. When several engines in the winning tier claim a type, a WARN at startup names them and the winner; an engine under parsers shadows a classpath engine (and draws the deprecation WARN below). Name the engine in text-recognizers to stop depending on any of this. An AutoDetectParser built in code without the loader applies the same rules at parse time.

An engine under parsers (deprecated)

Naming an engine under parsers is how 4.0 configured Tesseract’s options. It is deprecated since 4.1.0 and unsupported in 4.2.0, when parsers holds parsers only and an engine there fails config load. Until then it keeps working, with a WARN at startup that says what the entry does today, because it means two things at once: the entry configures the engine, and the engine is also a parser for the types no other parser there claims ("parsers": [{"tesseract-ocr-parser": {}}] alone sends every image straight to OCR; claude-vlm-parser next to a default-parser that excludes pdf-parser makes Claude the PDF parser). When every type the engine advertises is claimed by another parser it is never dispatched to and acts only as the text recognizer, or never runs at all. Configure the engine under engines, name it in text-recognizers, and keep parsers for parsers.

Replacing extracted text

An engine that produces the document’s text implements org.apache.tika.parser.enricher.TextRecognizer. Only such engines count when a caller asks whether an engine can stand in for text it already has: the PDF parser’s AUTO OCR strategy replaces a page’s extracted text only if the engine recognizes text and actually wrote some, and keeps the extracted text otherwise. The bundled OCR engines declare the capability; a VLM parser declares it through its textRecognizer option (default true). An entry that recognizes no text (the experimental image-embedding parser, a VLM with textRecognizer: false) still runs on the images and pages it is offered and never displaces text; startup warns about it, and 4.2 gives such annotators a list of their own. When the PDF parser also emits its page renders as embedded documents, it enriches each render itself and the embedded copy is not enriched again (see PDF Parser).

For engine authors

Recognizers are found by interface, not by media type: implement org.apache.tika.parser.enricher.ContentEnricher (or TextRecognizer, which extends it) and advertise the real types you enrich; Wiring in Your Own OCR Engine walks through a complete engine. An engine the classpath supplied is never dispatched to: tesseract-ocr-parser advertising image/png does not take PNGs away from image-parser, it is called by it. The image/ocr-* pseudo-types that routed OCR before 4.1 are retired: a third-party engine still advertising them is treated as a text recognizer for the real types, with a WARN at startup, until 5.0. Nothing else in Tika produces them; a tk:content-type-parser-override naming one matches no parser and falls to the empty fallback.

Secrets from the environment

A string value in the startup config may name an environment variable as ${env:NAME}, alone or inside a longer string, and the loader substitutes the variable’s value when the config is read:

{
  "engines": {
    "jina": { "openai-embedding-engine": { "baseUrl": "https://api.jina.ai",
              "apiKey": "${env:JINA_API_KEY}", "model": "jina-embeddings-v5-omni-small" } }
  },
  "emitters": {
    "es": { "es-emitter": { "esUrl": "http://${env:ES_HOST}:9200/tika", "apiKey": "${env:ES_API_KEY}" } }
  }
}

Only that exact form is touched: ${NAME}, $NAME, ${sys:…​} and any other text stay literal, so a prompt or a pattern that happens to contain ${ is safe. A reference to a variable that is not set fails the load with the variable’s name and the JSON path, never a value, so a missing secret stops the server at startup rather than reaching a service as an empty key. A literal value works as before.

The forked parse workers of tika-server and Pipes inherit the parent’s environment and read the same config file, so a reference resolves there too. Presets written inline in the config file resolve at startup with the rest of it; catalog presets and per-request configuration (/config endpoints, pipes tuples) never do: a request cannot read the server’s environment.

A reference keeps a secret out of tika-config.json, not out of every file: with a persistent config store (configStoreType file or Ignite) the resolved fetcher and emitter configurations are written to that store. The default in-memory store writes nothing.

Windows file paths

JSON treats the backslash as an escape character, so path options (tesseractPath, imageMagickPath, …​) need forward slashes (C:/Tools/…​) or escaped backslashes (C:\\Tools\...). A single backslash is a JSON parse error.

Adding extra jars (tika.extras.dir)

To add components — extra EncodingDetector or Parser implementations, or their dependencies — without repackaging the application, drop their jars in a directory and point the tika.extras.dir system property at it:

java -Dtika.extras.dir=/path/to/extras -jar tika-app.jar ...
java -Dtika.extras.dir=/path/to/extras -jar tika-server-standard.jar ...

Every *.jar in that directory joins the classpath Tika’s service loading scans, so SPI-registered components are picked up automatically. The jars are forwarded to the forked JVMs that Pipes and tika-server run, too, so they are present where parsing happens.

This is off by default: nothing is loaded unless tika.extras.dir is set, and there is no implicit default directory. Treat the directory as a trusted code location — anything in it runs with the full privileges of the Tika process, so it must not be writable by less-trusted principals, nor (on a server) reachable by request handling.

In the Docker images, mount a directory to /tika-extras instead. PF4J-managed plugins are a separate mechanism — see Pipes Plugins.

Topics

Parser configuration:

  • Pages — rendering, OCR policy and page images, for every parser that renders

  • PDFParser — PDF parsing options

  • Image Parsers — ICC metadata options for JPEG, TIFF, HEIF, BPG and WebP

  • TesseractOCRParser — OCR for image-based text extraction

  • Tess4J OCR Parser — in-process OCR via tess4j JNA bindings; advanced users only, most should prefer TesseractOCRParser

  • VLM Parsers — Claude, Gemini, OpenAI, Ollama, vLLM

  • External Parser — wrap external tools (ffmpeg, exiftool, …​)

Other configuration: