OCR and Inference Recipes
- Scanned pages only
- Hosted OCR: jina-ocr-v1
- Text vectors for every document
- Page vectors, no OCR
- OCR when needed and a vector on every page
- PDFs by page vectors, everything else by text vectors
- Text, pictures, audio and video in one index (Jina omni)
- OCR off for PDFs only
- Everything off for one request
- Not yet
Each recipe is a complete config. They combine four sections:
text-recognizers (what produces text from
pixels), engines and inference (what is called, on what),
and pages (where a page’s text comes from, what a PDF releases,
how pages are rendered). Every default parser stays loaded in all of them.
engines, text-recognizers and pages are stable; inference (and
pages.inference) is experimental in 4.1.0 and may change in a minor release without a
deprecation cycle as 4.2 adds batched recognition, media inputs and document-level tasks. The
recipes are correct for the release they ship with; check the release notes before moving.
|
Scanned pages only
OCR a PDF page only when its extracted text is poor, and every embedded image. Tesseract is
configured once and named as the recognizer; the PDF’s AUTO strategy is the default.
{
"engines": { "tesseract": { "tesseract-ocr-parser": { "language": "eng" } } },
"text-recognizers": [ { "engine": "tesseract" } ]
}
Hosted OCR: jina-ocr-v1
A hosted document-OCR model in place of Tesseract, through the OpenAI chat-completions shape.
It is a text recognizer like Tesseract: images and, under AUTO, the PDF pages whose extracted
text is poor. The key comes from the environment
(Secrets from the environment).
{
"engines": {
"jina-ocr": { "openai-vlm-parser": {
"baseUrl": "https://api.jina.ai",
"model": "jina-ocr-v1",
"apiKey": "${env:JINA_API_KEY}",
"prompt": "Transcribe the provided document image into a clean Markdown format, preserving the natural reading order.",
"maxTokens": 8192 } }
},
"text-recognizers": [ { "engine": "jina-ocr" } ]
}
The model returns markdown with tables as raw HTML, whatever the prompt says; Tika parses both
into XHTML, so a table is a <table> in every output format and a pipe table in markdown. Body
cells come back as <th> and are passed through as such. The load-time health check reads the
public /v1/models, so a wrong key passes startup and surfaces on the first request as the
document’s parse exception (HTTP 401 in the message).
Text vectors for every document
One embedding request per document tree, chunks of the container and its attachments mixed,
each chunk on its own document with a text locator.
{
"engines": {
"embedder": { "openai-embedding-engine": { "baseUrl": "http://embed:8000", "model": "bge-m3" } }
},
"inference": [
{ "id": "text-vectors", "engine": "embedder", "input": "TEXT", "tasks": ["embed"],
"chunker": { "markdown-chunker": { "maxChunkChars": 1500, "overlapChars": 200 } } }
]
}
Page vectors, no OCR
Every PDF page rendered once for the embedding engine, no text recognizer anywhere. The render
settings live under ocr and are sized here for an image model rather than Tesseract.
{
"engines": {
"clip": { "openai-embedding-engine": { "baseUrl": "http://clip:8000", "model": "siglip2" } }
},
"inference": [
{ "id": "page-vectors", "engine": "clip", "input": "PAGES", "tasks": ["embed"], "maxChunks": 50 }
],
"text-recognizers": [],
"parsers": [
{ "default-parser": {} },
{ "pdf-parser": { "pages": { "inference": ["PAGES"], "text": "EXTRACT",
"render": { "dpi": 100, "imageType": "RGB" } } } }
]
}
OCR when needed and a vector on every page
One render per page feeds both: Tesseract on the pages the AUTO verdict picks, the embedding
engine on all of them. The settings serve both consumers, so keep the dpi Tesseract needs.
{
"engines": {
"tesseract": { "tesseract-ocr-parser": { "language": "eng" } },
"clip": { "openai-embedding-engine": { "baseUrl": "http://clip:8000", "model": "siglip2" } }
},
"text-recognizers": [ { "engine": "tesseract" } ],
"inference": [
{ "id": "page-vectors", "engine": "clip", "input": "PAGES", "tasks": ["embed"] }
],
"parsers": [
{ "default-parser": {} },
{ "pdf-parser": { "pages": { "inference": ["PAGES"], "render": { "imageType": "RGB" } } } }
]
}
PDFs by page vectors, everything else by text vectors
Two bindings. The PDF releases its pages, records that it did, and the text stage skips it;
every other document is embedded by its text. List ["TEXT", "PAGES"] on the PDF to get both.
{
"engines": {
"embedder": { "openai-embedding-engine": { "baseUrl": "http://embed:8000", "model": "bge-m3" } },
"clip": { "openai-embedding-engine": { "baseUrl": "http://clip:8000", "model": "siglip2" } }
},
"inference": [
{ "id": "text-vectors", "engine": "embedder", "input": "TEXT", "tasks": ["embed"],
"chunker": { "markdown-chunker": { "maxChunkChars": 1500, "overlapChars": 200 } } },
{ "id": "page-vectors", "engine": "clip", "input": "PAGES", "tasks": ["embed"], "maxChunks": 50 }
],
"parsers": [
{ "default-parser": {} },
{ "pdf-parser": { "pages": { "inference": ["PAGES"], "text": "EXTRACT",
"render": { "dpi": 100, "imageType": "RGB" } } } }
]
}
Text, pictures, audio and video in one index (Jina omni)
One engine, jina-embeddings-v5-omni-small, embeds every input kind into one 1024-dimension
space: the extracted text of every document, every picture, and audio and video cut into
segments. Four bindings share it. mediaInput: "object" is required for the MEDIA bindings,
normalized makes cosine scores comparable across the four fields, and task selects the
indexing adapter (a query uses retrieval.query with the same model).
{
"engines": {
"jina": { "openai-embedding-engine": {
"baseUrl": "https://api.jina.ai",
"apiKey": "${env:JINA_API_KEY}",
"model": "jina-embeddings-v5-omni-small",
"mediaInput": "object",
"maxBatchSize": 8,
"requestParameters": { "task": "retrieval.passage", "dimensions": 1024, "normalized": true } } }
},
"inference": [
{ "id": "text", "engine": "jina", "input": "TEXT", "tasks": ["embed"],
"chunker": { "markdown-chunker": { "maxChunkChars": 1500, "overlapChars": 200 } } },
{ "id": "pictures", "engine": "jina", "input": "IMAGES", "tasks": ["embed"] },
{ "id": "jina-video", "engine": "jina", "input": "MEDIA", "modality": "visual" },
{ "id": "jina-audio", "engine": "jina", "input": "MEDIA", "modality": "audio" }
]
}
Media needs ffmpeg on the PATH; the -full Docker image ships it, the minimal one does not.
The default grid, 25 s windows with 5 s overlap, stays under the 30 s cap Jina puts on an audio
item; a media block that raises seconds to 30 or more has every audio segment refused. A
segment’s two vectors share a correlator, and the Elasticsearch and OpenSearch emitters'
chunkStrategy: DOCUMENTS merges them into one document per segment (see
Chunk strategies).
OCR off for PDFs only
The recognizer stays on for embedded and standalone images; the PDF parser never renders.
{
"engines": { "tesseract": { "tesseract-ocr-parser": {} } },
"text-recognizers": [ { "engine": "tesseract" } ],
"parsers": [
{ "default-parser": {} },
{ "pdf-parser": { "pages": { "text": "EXTRACT" } } }
]
}
Everything off for one request
Both lists stay configured; the request switches them off. Wire-safe, so it works as a
preset, on /rmeta/config, and on a pipes tuple.
{ "parse-context": { "text-recognizers": { "enabled": false },
"inference": { "enabled": false } } }
inference.bindings selects a subset instead of all-or-nothing; a PDF can be switched to pages
for one request with "pages": { "inference": ["PAGES"] } in the same block.
Not yet
A caption or tag per picture (a chat task on a chat engine), OCR configured but promoted per request, and separate render settings for OCR and for an image model are 4.2 work; the Inference page says what each binding can do today.