VLM (Vision-Language Model) Parsers

Tika includes a family of parsers that delegate OCR and document understanding to remote Vision-Language Model (VLM) endpoints. These parsers send images (or PDFs) to an external API and convert the model’s markdown response into structured XHTML.

tika-vlm is experimental. Its classes, configuration keys and metadata output may change in minor releases without a deprecation cycle, and no compatibility shims are kept for earlier shapes.

Three implementations are provided out of the box. None is auto-loaded: each must be named explicitly in your configuration. (Changed in 4.1.0: openai-vlm-parser previously auto-registered via SPI.) To use a VLM as the OCR engine for embedded images and rendered PDF pages, configure it under engines and name it in the text-recognizers list — see Configuration. Naming it under parsers is deprecated since 4.1.0 and unsupported in 4.2.0.

Parser Endpoint Config key

OpenAIVLMParser

Any OpenAI-compatible chat completions endpoint (vLLM, Ollama, local FastAPI, OpenAI)

openai-vlm-parser

ClaudeVLMParser

Anthropic Messages API

claude-vlm-parser

GeminiVLMParser

Google Gemini generateContent API

gemini-vlm-parser

All three enrich the standard image types (image/png, image/jpeg, …​). ClaudeVLMParser and GeminiVLMParser additionally declare application/pdf, so under parsers they can process PDFs natively with the model’s vision capabilities once the default pdf-parser is excluded (see below, deprecated); an enricher never displaces the parser for a type.

Module dependency

<dependency>
  <groupId>org.apache.tika</groupId>
  <artifactId>tika-vlm</artifactId>
  <version>4.1.0</version>
</dependency>
To run a local open-source VLM without cloud API keys, see Running a Local VLM Server.

OpenAI-compatible (vLLM, Ollama, etc.)

Basic Configuration

{
  "engines": {
    "vlm": {
      "openai-vlm-parser": {
        "baseUrl": "http://127.0.0.1:8000",
        "model": "jinaai/jina-vlm",
        "timeoutMillis": 300000
      }
    }
  },
  "text-recognizers": [
    {
      "engine": "vlm"
    }
  ]
}

Full Configuration

{
  "engines": {
    "vlm": {
      "openai-vlm-parser": {
        "baseUrl": "http://127.0.0.1:8000",
        "completionsPath": "/v1/chat/completions",
        "model": "jinaai/jina-vlm",
        "prompt": "Extract all visible text from this image. Return the text in markdown format, preserving the original structure (headings, lists, tables, paragraphs). Do not describe the image. Only return the extracted text.",
        "maxTokens": 4096,
        "timeoutMillis": 300000,
        "apiKey": "",
        "inlineContent": true,
        "skipOcr": false,
        "textRecognizer": true,
        "minFileSizeToOcr": 0,
        "maxFileSizeToOcr": 52428800,
        "maxImagePixels": 100000000,
        "allowRuntimePrompt": false
      }
    }
  },
  "text-recognizers": [
    {
      "engine": "vlm"
    }
  ]
}

The OpenAIVLMParser works with any server that exposes an /v1/chat/completions endpoint in the OpenAI format. This includes:

  • vLLM

  • Ollama

  • A local FastAPI / Flask wrapper around a Hugging Face model

  • OpenAI itself

Authentication uses a standard Authorization: Bearer <apiKey> header. Leave apiKey empty to skip authentication (typical for local servers).

Anthropic Claude

Basic Configuration

{
  "engines": {
    "claude": {
      "claude-vlm-parser": {
        "apiKey": "YOUR_ANTHROPIC_API_KEY",
        "model": "claude-sonnet-4-20250514"
      }
    }
  },
  "text-recognizers": [
    {
      "engine": "claude"
    }
  ]
}

Full Configuration

{
  "engines": {
    "claude": {
      "claude-vlm-parser": {
        "baseUrl": "https://api.anthropic.com",
        "model": "claude-sonnet-4-20250514",
        "prompt": "Extract all visible text from this image. Return the text in markdown format, preserving the original structure (headings, lists, tables, paragraphs). Do not describe the image. Only return the extracted text.",
        "maxTokens": 4096,
        "timeoutMillis": 300000,
        "apiKey": "YOUR_ANTHROPIC_API_KEY",
        "inlineContent": true,
        "skipOcr": false,
        "textRecognizer": true,
        "minFileSizeToOcr": 0,
        "maxFileSizeToOcr": 52428800,
        "maxImagePixels": 100000000,
        "allowRuntimePrompt": false
      }
    }
  },
  "text-recognizers": [
    {
      "engine": "claude"
    }
  ]
}

The ClaudeVLMParser uses the Anthropic Messages API. Authentication uses the x-api-key header (not Bearer). The required anthropic-version header is sent automatically.

Claude handles images and PDFs natively. For images, the content block type is image; for PDFs it is document. The parser detects the correct type from the input MIME type.

Google Gemini

Basic Configuration

{
  "engines": {
    "gemini": {
      "gemini-vlm-parser": {
        "apiKey": "your-gemini-api-key",
        "model": "gemini-2.5-flash"
      }
    }
  },
  "text-recognizers": [
    {
      "engine": "gemini"
    }
  ]
}

Full Configuration

{
  "engines": {
    "gemini": {
      "gemini-vlm-parser": {
        "baseUrl": "https://generativelanguage.googleapis.com",
        "model": "gemini-2.5-flash",
        "prompt": "Extract all visible text from this image. Return the text in markdown format, preserving the original structure (headings, lists, tables, paragraphs). Do not describe the image. Only return the extracted text.",
        "maxTokens": 4096,
        "timeoutMillis": 300000,
        "apiKey": "your-gemini-api-key",
        "inlineContent": true,
        "skipOcr": false,
        "textRecognizer": true,
        "minFileSizeToOcr": 0,
        "maxFileSizeToOcr": 52428800,
        "maxImagePixels": 100000000,
        "allowRuntimePrompt": false
      }
    }
  },
  "text-recognizers": [
    {
      "engine": "gemini"
    }
  ]
}

The GeminiVLMParser targets the Google Gemini generateContent endpoint. The API key is passed as a key query parameter.

Change baseUrl if you are using Vertex AI or a proxy.

Using a VLM parser for PDF parsing (deprecated)

Claude and Gemini can process entire PDFs with their vision capabilities. To route PDFs to a VLM parser instead of the default PDFParser, exclude the default and add the VLM parser:

This shape is deprecated since 4.1.0 and unsupported in 4.2.0, when an engine under parsers fails config load. The 4.2 way to put a PDF in front of a VLM is the PDF parser’s "pages": {"text": "OCR"} with the VLM named in text-recognizers: every page is rendered and transcribed, and the PDF’s own structure (attachments, bookmarks, forms) is still parsed. That sends page images, not the PDF file; there is no 4.2 shape yet for sending the file itself.
{
  "parsers": [
    {
      "default-parser": {
        "exclude": ["pdf-parser"]
      }
    },
    {
      "claude-vlm-parser": {
        "apiKey": "YOUR_ANTHROPIC_API_KEY",
        "model": "claude-sonnet-4-20250514",
        "prompt": "Extract all text from this document. Return the text in markdown format, preserving the original structure (headings, lists, tables, paragraphs). Do not describe the document. Only return the extracted text."
      }
    }
  ]
}

You can substitute gemini-vlm-parser for claude-vlm-parser above.

Named under parsers, the VLM is also the text recognizer for the image types it advertises, so embedded images in every other document type are sent to it too, and startup logs the deprecation WARN. Add a text-recognizers list to choose a different engine for those, or "text-recognizers": [] for none; see Configuration.

Configuration options reference

All three parsers share the same configuration POJO (VLMOCRConfig):

Property Default Description

baseUrl

varies by parser

Base URL of the API endpoint (no trailing slash).

model

varies by parser

Model identifier sent in the API request.

prompt

(markdown extraction prompt)

The text prompt sent alongside the image or document.

maxTokens

4096

Maximum tokens the model may generate.

timeoutMillis

120000

HTTP read timeout in milliseconds.

maxRetries

4

Retries of a 429, 502, 503 or 504 answer, with a jittered backoff inside the parse budget; 0 fails at once. A hosted service caps requests in flight per key, so see Hosted engines and concurrency.

apiKey

"" (empty)

API key. Format depends on the parser (Bearer header, x-api-key header, or query parameter). ${env:NAME} reads it from the environment (Secrets from the environment).

inlineContent

true

When parsing inline images (embedded resource type INLINE), write OCR text into the parent document’s content stream. Mirrors TesseractOCRParser inline behaviour.

skipOcr

false

Runtime kill-switch to disable the parser entirely.

textRecognizer

true

Whether the prompt transcribes the image, so the output is the document’s text. Set false for a captioning or tagging prompt: the parser still runs, but the PDF parser’s AUTO OCR never replaces a page’s extracted text with its output.

minFileSizeToOcr

0

Minimum input file size in bytes.

maxFileSizeToOcr

52428800 (50 MB)

Maximum input file size in bytes.

maxImagePixels

100000000 (100 megapixels)

Maximum decoded image area in pixels. Larger images are rejected before being sent to the model. Set to -1 to disable the limit. Guards against decompression-bomb inputs and runaway VLM cost on a single huge image.

allowRuntimePrompt

false

Whether a caller-supplied JSON config may override prompt. false pins the prompt to its initialization-time value and rejects the override. Security-relevant: a runtime-controllable prompt is a prompt-injection surface for anyone who can send per-request config, so enable it only for trusted callers. See Per-request configuration.

completionsPath

/v1/chat/completions (OpenAI/vLLM only)

HTTP path appended to baseUrl for the chat-completions endpoint. Used by the OpenAI-compatible parser only. Claude and Gemini hardcode their own API paths (/v1/messages and /v1beta/models/{model}:generateContent respectively) and ignore this field.

Markdown-to-XHTML conversion

The VLM’s text response is expected to be markdown. Tika parses it using commonmark-java and emits proper XHTML elements (<h1>, <p>, <table>, <b>, <i>, etc.) instead of dumping raw text. GFM tables and strikethrough are supported.

Document-OCR models such as jina-ocr-v1 and DeepSeek-OCR write tables as raw HTML inside the markdown, and a prompt does not change that. An HTML block is parsed leniently and re-emitted as the same structural elements (table, tr, th, td, p, lists, headings) with attributes dropped except colspan and rowspan; other elements contribute their text, and script/style contribute nothing. Inline HTML tags are dropped and their text kept, except <br>.

Per-request configuration

There are two paths, and they are trusted differently.

In-process Java. Set a VLMOCRConfig on the ParseContext. Nothing is restricted — code holding the ParseContext is trusted by construction.

VLMOCRConfig override = new VLMOCRConfig();
override.setModel("claude-opus-4-20250514");
override.setMaxTokens(8192);

ParseContext context = new ParseContext();
context.set(VLMOCRConfig.class, override);

Caller-supplied JSON — the tika-server /config endpoints and tika-grpc’s parse_context_json, each gated by its own allowPerRequestConfig. Here the parser applies a restricted config that rejects cost- and security-sensitive changes with an exception:

Field Runtime rule

baseUrl, apiKey

Always blocked; configure at initialization time.

model

Always blocked. Configure a separate parser instance for a different model.

maxTokens

May be lowered, never raised above the initialization-time value — that ceiling is what bounds cost on a paid endpoint.

maxImagePixels, maxRetries, allowRuntimePrompt

Always blocked.

prompt, textRecognizer

Blocked unless allowRuntimePrompt was set to true at initialization time; textRecognizer describes the prompt, so it changes only when the prompt can.