Chunk Strategies for Search Engines
Tika 4.x introduces a unified chunking and embedding pipeline (tika-inference) that puts a
tk:chunks array on each document’s metadata. This page describes how those chunks reach
Elasticsearch and OpenSearch, and why the shipped strategy is the one it is.
tika-inference is experimental. Its classes, configuration keys and the tk:chunks
layout may change in minor releases without a deprecation cycle, and no compatibility shims
are kept for earlier shapes.
|
The tk:chunks Field
All chunk data lands in one metadata field, tk:chunks, as a JSON array. There is no
separate field for image embeddings versus text embeddings — a chunk is a chunk, identified
by which locators it carries:
| Chunk kind | Produced by, and what it carries |
|---|---|
Text |
A |
Image |
An |
Media segment |
Two |
SpatialLocator (bounding box) is defined and round-tripped by ChunkSerializer; no
component produces it yet. EmbeddedLocator (id_path, name) marks a chunk that was
lifted from an embedded document into its parent; see Where a Chunk Lands: the Parent.
A chunk carries one vector and says where it came from three ways: producer is the binding
that wrote it, modality is what the engine saw (text, visual, audio), and
correlator is the id of the unit it was cut from, assigned by the chunker that defined the
unit, so chunks of one unit written by different bindings can be grouped (a segment’s sound
and picture; later, a page region’s OCR text and its crop):
{
"text": "Revenue grew 15% year-over-year...",
"vector": "base64-encoded-float32-be",
"producer": "text-vectors",
"modality": "text",
"locators": {
"text": [{"start_offset": 0, "end_offset": 120}],
"paginated": [{"page": 1}]
}
}
When several components produce chunks for the same document, they append to the same
array via ChunkSerializer.mergeInto(). The 4.0 filters set neither producer, modality
nor correlator. A modality value this Tika does not know (a newer writer’s) reads as
unset, so a foreign chunk never breaks the array for the chunks after it.
What Tika Emits Today
Each file — the container and each embedded file — is a separate Elasticsearch or
OpenSearch document, matching the existing SEPARATE_DOCUMENTS attachment strategy. Chunks
ride along inside their file’s document as a structured JSON array: not exploded into
separate documents, and not a stringified blob.
{"index":{"_id":"email.msg"}}
{"title":"Re: Q4 report","mime":"message/rfc822","tk:chunks":[
{"text":"Hi team, see attached...","vector":"...","locators":{"text":[{"start_offset":0,"end_offset":35}]}}
]}
{"index":{"_id":"email.msg-<uuid>"}}
{"title":"Q4-report.pdf","mime":"application/pdf","parent":"email.msg","tk:chunks":[
{"vector":"...","locators":{"paginated":[{"page":1}]}},
{"vector":"...","locators":{"paginated":[{"page":2}]}},
{"text":"Revenue grew...","vector":"...","locators":{"text":[{"start_offset":0,"end_offset":120}]}},
{"text":"Operating costs...","vector":"...","locators":{"text":[{"start_offset":121,"end_offset":300}]}}
]}
Mechanically:
-
The emitter’s
AttachmentStrategy(SEPARATE_DOCUMENTSorPARENT_CHILD) decides how embedded files are emitted; each embedded file becomes its own document either way. -
tk:chunkson each document holds all of that document’s chunks, text and image alike. -
The emitter clients (
ESClient,OpenSearchClient) recognizetk:chunksand write it as raw JSON rather than an escaped string, so the nested objects are indexable.
Why this shape:
-
One document per file matches how users think about documents, and an embedded file (a PDF inside an email) gets its own document with its own chunks.
-
Structured JSON lets Elasticsearch/OpenSearch index the vectors and locators natively.
-
No join field and no routing are needed for the chunk relationship.
-
It stays compatible with
nestedkNN when you want to search within a document’s chunks.
The trade-offs are the nested kNN caveats below, and large documents for files with many
pages.
Where a Chunk Lands: the Parent
An embedded document that is part of its parent, a picture in the body of a docx or an email, should not be a search hit of its own: its vector belongs to the document it appears in. The image embedder therefore writes a picture’s chunk onto the parent, not onto the picture, so a docx carries one vector per picture in its body and an email the vectors of its inline pictures. Nothing to configure:
{
"engines": {
"clip": { "openai-embedding-engine": { "baseUrl": "http://clip-server:8000",
"model": "jinaai/jina-clip-v2" } },
"text": { "openai-embedding-engine": { "baseUrl": "http://embed-server:8000",
"model": "text-embedding-3-small" } }
},
"inference": [
{ "id": "pictures", "engine": "clip", "input": "IMAGES", "tasks": ["embed"] },
{ "id": "text", "engine": "text", "input": "TEXT", "tasks": ["embed"] }
]
}
(The 4.0 openai-image-embedding-parser and openai-embedding-filter, deprecated since 4.1.0,
place chunks by the same rule.)
The rules:
-
INLINEandRENDERINGchildren hand their chunks to the parent. Attachments keep their own: a PDF inside an email is still its own document with its own vectors, as above. -
The parent is the immediate container, so a picture inside an attached docx reaches the docx, not the email.
-
Every chunk written to a parent carries an
embeddedlocator,{"id_path": "/1", "name": "image1.png"}, naming the child it came from, beside any locator it already had. -
"liftToParent": falseon the deprecatedopenai-image-embedding-parserkeeps every picture’s vector on the picture, for an index where each picture is a hit of its own; theIMAGESbinding has no such switch. -
In the recursive (
/rmeta,-J, pipesRMETA) parse the wrapper keeps every document’s metadata, so a chunk lands on its parent’s. In the single-object parse (/tikaas JSON,-j, pipesCONCATENATE) only the top-level document’s metadata is handed back, so every chunk lands there, with the sameembeddedlocator naming the part it came from. A task writes where the dispatcher aims it (InferenceUnit.getDestination()), throughChunkTarget(tika-inference).
PDF page renders follow the same rule by a different route: the PDF parser enriches its own
renders and their vectors land on the PDF directly, one per page with a paginated locator,
whether the pages were rendered for OCR or emitted as embedded documents (see
PDF Parser).
Elasticsearch/OpenSearch Mapping
Two ways to index chunks. The emitters' chunkStrategy: DOCUMENTS writes one document per
unit next to the file’s document: chunks sharing a correlator merge into one document, a
chunk without one is its own, the locators are flattened to fields and every vector is
decoded to a float array under v.<producer>:
{
"file_id": "s3://bucket/talk.mp4", "container_id": "s3://bucket/talk.mp4",
"correlator": "t:20000-45000", "start_ms": 20000, "end_ms": 45000,
"v": { "jina-video": [0.01, ...], "jina-audio": [0.02, ...] }
}
A hit is then a chunk, and Elasticsearch adds up scores per document, so two knn clauses
(or an rrf retriever) over v.jina-video and v.jina-audio rank a segment that matches
in both channels above one that matches in either: search one channel with one clause, both
with two. That cannot be done with segments nested inside the file’s document, where each
clause scores the file by its best child, possibly a different child per clause. The rrf
retriever is a licensed Elasticsearch feature and answers 403 on a basic license; a top-level
knn array works on any license, but it adds the scores, so a segment with two vector fields
outranks a text chunk with one even when the text is the better match. Search one channel at
a time when that matters. Elasticsearch 9 leaves dense_vector values out of _source by
default, so a fetched document shows no v; read them back with fields or check them with an
exists query. Mapping:
{
"mappings": {
"properties": {
"file_id": {"type": "keyword"},
"container_id": {"type": "keyword"},
"correlator": {"type": "keyword"},
"start_ms": {"type": "long"},
"end_ms": {"type": "long"},
"page": {"type": "integer"},
"text": {"type": "text"},
"v": {
"properties": {
"jina-video": {"type": "dense_vector", "dims": 1024, "index": true, "similarity": "cosine"},
"jina-audio": {"type": "dense_vector", "dims": 1024, "index": true, "similarity": "cosine"}
}
}
}
}
}
Group hits back to files with collapse on container_id, the emit key, which is stable
across runs; file_id is the emitter’s id for the document the chunk came from, and for an
attachment that id is minted per run. Text chunks carry start_offset, end_offset and
text; page chunks page (and bbox when the locator has one); lifted chunks
embedded_id_path and embedded_name. modality and producer are not fields on a merged
document: the vector’s key under v is the producer.
Chunk documents carry no join field, so DOCUMENTS needs attachmentStrategy:
SEPARATE_DOCUMENTS; the emitter refuses the pair at load. Re-indexing a file replaces its
chunk documents by id (<id>-chunk-0 …) but never deletes the tail of an earlier run that
had more of them, and under UPSERT a removed binding’s v.<producer> survives on the
documents that are rewritten: delete by container_id before re-indexing a changed grid or
binding set. A chunk whose vector is not base64 float32 keeps its other fields and loses
the vector, with a warning; a tk:chunks value that is not a JSON array stays on the file’s
document as INLINE would write it.
The default, chunkStrategy: INLINE, keeps tk:chunks inside the file’s document as a
JSON array. With nested kNN support, a mapping like this works:
{
"mappings": {
"properties": {
"title": {"type": "text"},
"mime": {"type": "keyword"},
"content": {"type": "text"},
"parent": {"type": "keyword"},
"tk:chunks": {
"type": "nested",
"properties": {
"text": {"type": "text"},
"producer": {"type": "keyword"},
"modality": {"type": "keyword"},
"correlator": {"type": "keyword"},
"vector": {
"type": "dense_vector",
"dims": 1024,
"index": true,
"similarity": "cosine"
},
"locators": {
"properties": {
"text": {
"type": "nested",
"properties": {
"start_offset": {"type": "integer"},
"end_offset": {"type": "integer"}
}
},
"paginated": {
"type": "nested",
"properties": {
"page": {"type": "integer"}
}
}
}
}
}
}
}
}
}
Inline, each vector is a base64-encoded big-endian float32 array and the emitter
writes it as-is. Elasticsearch 9.3+ accepts that encoding for dense_vector; earlier
versions need an ingest pipeline to decode it, or chunkStrategy: DOCUMENTS, which decodes.
|
Alternatives Considered
| Option | Approach | Why not |
|---|---|---|
A |
One document per file with |
|
B |
Each chunk becomes its own document with a |
Many more documents, and parent metadata must be either denormalized onto every chunk or fetched in a second lookup. Still wanted — see Future Work. |
C |
Chunks are children of the container document via the ES/OpenSearch
join field,
so parent metadata is never duplicated and |
Join queries are expensive, kNN combined with |
Future Work
-
Option B support — a
ChunkStrategy.SEPARATE_DOCUMENTSthat explodes chunks into individual ES/OpenSearch documents at emit time, for simpler kNN search withoutnestedqueries. -
Hybrid search — combining kNN vector search on chunks with BM25 text search on the parent document’s
contentfield. -
Chunk-level metadata — propagating selected parent metadata onto each chunk for filtering during vector search.