Class TikaInputStream

java.lang.Object
java.io.InputStream
java.io.FilterInputStream
org.apache.commons.io.input.ProxyInputStream
org.apache.commons.io.input.TaggedInputStream
org.apache.tika.io.TikaInputStream
All Implemented Interfaces:
Closeable, AutoCloseable

public class TikaInputStream extends org.apache.commons.io.input.TaggedInputStream
Input stream with extended capabilities for detection and parsing.

This implementation uses backing strategies to handle different input types:

  • ByteArraySource for byte[] inputs - no caching needed
  • FileSource for Path/File inputs - direct file access
  • CachingSource for InputStream inputs - passthrough by default; caches bytes only after enableRewind()
Since:
Apache Tika 0.8
  • Method Details

    • get

      public static TikaInputStream get(InputStream stream, TemporaryResources tmp, Metadata metadata)
    • get

      public static TikaInputStream get(org.apache.commons.io.function.IOSupplier<InputStream> opener, TemporaryResources tmp, Metadata metadata)
      Creates a TikaInputStream from a re-openable stream supplier. Unlike get(InputStream, TemporaryResources, Metadata) -- which caches a one-shot stream to memory/disk so it can be rewound -- the supplier is re-invoked to re-read the content, so rewinding (e.g. during digesting) never spills to disk. A temp file is created only if getPath() is later called (a parser/detector needing a File) or getSeekableByteChannel() is asked for content that does not fit in memory.
      Parameters:
      opener - supplies a fresh InputStream over the same content on each call
      tmp - temporary resources for any on-demand getPath() spill
      metadata - metadata used for extension/length hints; may be null
    • get

      public static TikaInputStream get(InputStream stream)
    • get

      public static TikaInputStream get(InputStream stream, Metadata metadata)
    • get

      public static TikaInputStream get(byte[] data)
    • get

      public static TikaInputStream get(byte[] data, Metadata metadata)
    • getPlaceholder

      public static TikaInputStream getPlaceholder()
      An empty stream standing in for content that is never extracted -- a metadata-only entry, a rendering carried as an open container. It reports an unknown length, so nothing mistakes the placeholder's size for the document's. Pair it with MetadataOnlyParse to register an entry without parsing it, unless an open container supplies the content.
    • get

      public static TikaInputStream get(Path path) throws IOException
      Throws:
      IOException
    • get

      public static TikaInputStream get(Path path, Metadata metadata) throws IOException
      Throws:
      IOException
    • get

      public static TikaInputStream get(Path path, Metadata metadata, TemporaryResources tmp) throws IOException
      Throws:
      IOException
    • get

      public static TikaInputStream get(File file) throws IOException
      Throws:
      IOException
    • get

      public static TikaInputStream get(File file, Metadata metadata) throws IOException
      Throws:
      IOException
    • get

      public static TikaInputStream get(Blob blob) throws SQLException, IOException
      Throws:
      SQLException
      IOException
    • get

      public static TikaInputStream get(Blob blob, Metadata metadata) throws SQLException, IOException
      Throws:
      SQLException
      IOException
    • get

      public static TikaInputStream get(URI uri) throws IOException
      Throws:
      IOException
    • get

      public static TikaInputStream get(URI uri, Metadata metadata) throws IOException
      Throws:
      IOException
    • get

      public static TikaInputStream get(URL url) throws IOException
      Throws:
      IOException
    • get

      public static TikaInputStream get(URL url, Metadata metadata) throws IOException
      Throws:
      IOException
    • getFromContainer

      public static TikaInputStream getFromContainer(Object openContainer, long length, Metadata metadata)
    • skip

      public long skip(long n) throws IOException
      Skips up to n bytes. Returns the actual number of bytes skipped, which may be less than requested if the end of stream is reached.

      This method does NOT throw EOFException if fewer bytes are available. Callers must check the return value to determine how many bytes were actually skipped.

      Overrides:
      skip in class org.apache.commons.io.input.ProxyInputStream
      Parameters:
      n - the number of bytes to skip
      Returns:
      the actual number of bytes skipped (may be less than n)
      Throws:
      IOException
    • mark

      public void mark(int readlimit)
      Overrides:
      mark in class org.apache.commons.io.input.ProxyInputStream
    • markSupported

      public boolean markSupported()
      Overrides:
      markSupported in class org.apache.commons.io.input.ProxyInputStream
    • reset

      public void reset() throws IOException
      Overrides:
      reset in class org.apache.commons.io.input.ProxyInputStream
      Throws:
      IOException
    • close

      public void close() throws IOException
      Specified by:
      close in interface AutoCloseable
      Specified by:
      close in interface Closeable
      Overrides:
      close in class org.apache.commons.io.input.ProxyInputStream
      Throws:
      IOException
    • afterRead

      protected void afterRead(int n) throws IOException
      Overrides:
      afterRead in class org.apache.commons.io.input.ProxyInputStream
      Throws:
      IOException
    • peek

      public int peek(byte[] buffer) throws IOException
      Throws:
      IOException
    • getOpenContainer

      public Object getOpenContainer()
    • setOpenContainer

      public void setOpenContainer(Object container)
    • addCloseableResource

      public void addCloseableResource(Closeable closeable)
    • hasFile

      public boolean hasFile()
      Whether the content is already on disk: a real file, or a stream cache that spilled. getPath() then returns that file without re-copying anything already written to it -- but it is not free, and it is not a getter: for a cache that spilled mid-stream it first drains the rest of the source into the file, switches this stream to reading from that file, and sets Content-Length on the Metadata this stream was created with. Use hasLength() if you only need the size.
    • getPath

      public Path getPath() throws IOException
      Throws:
      IOException
    • getFile

      public File getFile() throws IOException
      Throws:
      IOException
    • getFileChannel

      public FileChannel getFileChannel() throws IOException
      Throws:
      IOException
    • hasLength

      public boolean hasLength()
    • hasReliableLength

      public boolean hasReliableLength()
      True when getLength() would return a measured, ground-truth length (file, byte array, fully-drained cache, explicit override) without forcing a spool. False when the only length available is a caller-declared hint (Content-Length metadata, HTTP header, archive central directory), which may lie.
    • getLength

      public long getLength() throws IOException
      The stream length. For a stream-backed instance with no declared length this spools the entire remaining stream to a temporary file to measure it.
      Throws:
      IOException
    • getPosition

      public long getPosition()
    • setCloseShield

      public void setCloseShield()
    • removeCloseShield

      public void removeCloseShield()
    • isCloseShield

      public boolean isCloseShield()
    • rewind

      public void rewind() throws IOException
      Rewind the stream to the beginning.

      For streams created from byte arrays or files, this always works. For streams created from raw InputStreams, this requires enableRewind() to have been called first.

      Throws:
      IOException
    • enableRewind

      public void enableRewind() throws IOException
      Enables full rewind capability for this stream.

      For streams backed by byte arrays or files, this is a no-op since they are inherently rewindable. For streams backed by raw InputStreams, this switches from passthrough mode to caching mode, enabling subsequent rewind(), mark(int)/reset(), and random access.

      Must be called when position is 0 (before any reading), otherwise throws IOException.

      Use this method when you know you'll need to rewind the stream later (e.g., for detection followed by parsing, or digest calculation). For streaming-only operations (e.g., HTML parsing), skip this call to avoid unnecessary caching overhead.

      Throws:
      IOException - if bytes have already been read from the stream (position is not 0); rewind support cannot be enabled retroactively
    • enableRewind

      public void enableRewind(CacheMemoryBudget budget) throws IOException
      Like enableRewind(), but supplies a shared CacheMemoryBudget governing how much may be held in memory before spilling to disk (used only by stream-backed sources); null falls back to the per-object default.
      Parameters:
      budget - shared memory budget, or null
      Throws:
      IOException - if bytes have already been read (position is not 0)
    • getSeekableByteChannel

      public SeekableByteChannel getSeekableByteChannel() throws IOException
      Returns a read-only random-access SeekableByteChannel over this stream's full content. Unlike getPath()/getFile(), this never forces content that is already in memory onto disk: in-memory content is served from memory, file-backed or spilled content from a file channel, and unread stream content is drained through the cache which decides memory-vs-disk as it goes. Use this when random access is needed (e.g. reading a zip central directory); reserve getFile() for callers that truly need a File. Closed with the stream if the caller does not close it first. Does not disturb this stream's read position.
      Throws:
      IOException - if this stream has been partially read without rewind enabled
    • tryRetainInMemory

      public boolean tryRetainInMemory() throws IOException
      Asks the underlying source to hold its full content in memory, so a later rewind and re-read does not go back to the source. Worth calling before a pass that will be followed by another (digest then parse) when re-opening is expensive -- a zip entry has to be inflated again, where a file is served by the page cache.

      Advisory: sources that would have to spill, or cannot tell whether the content fits, return false and behave as before. Does not change the read position.

      Returns:
      true if the full content is now held in memory
      Throws:
      IOException
      See Also:
      • TikaInputSource.tryRetainInMemory()
    • inMemoryContent

      public static ByteBuffer inMemoryContent(SeekableByteChannel channel) throws IOException
      Zero-copy, read-only view of the content behind a channel from getSeekableByteChannel(), or null when that content is on disk. Lets a consumer that wants random access (PDFBox, metadata-extractor) read what is already in memory without a second copy; when this returns null the caller should use the file.

      The view aliases the cache's own array and is valid exactly while channel is open: the channel pins the array, and the content is fully drained before any channel is handed out. Keep the channel open for as long as the view is in use, then close it -- a view that outlives its channel still reads correctly but is no longer counted against the memory budget.

      Throws:
      IOException
    • toString

      public String toString()
      Overrides:
      toString in class Object