Class PDFMarkedContent2XHTML

java.lang.Object
org.apache.pdfbox.contentstream.PDFStreamEngine
org.apache.pdfbox.text.PDFTextStripper
org.apache.tika.parser.pdf.PDFMarkedContent2XHTML

public class PDFMarkedContent2XHTML extends org.apache.pdfbox.text.PDFTextStripper
Text extraction that follows a tagged PDF's structure tree.

The text stripper still does what it does in PDF2XHTML: it assembles words and lines from glyph positions, with the same spacing heuristics and config. This class only adds a marked-content stack over the content stream, so every glyph carries the MCID it was drawn under, and records the stripper's word and separator calls for the page instead of writing them. Once the page is assembled it either emits the words through the structure tree, or replays them exactly as PDF2XHTML would have written them. A page falls back when it has no tagged text, or under MarkedContentConfig.Strategy.AUTO when the tree claims too little of the page's text or references content the page does not have.

Text the tree does not reference is not lost: untagged text and /Artifact content (running headers, footers, page numbers) follow the tagged content on each page in their own divs.

Since:
1.24
  • Field Summary

    Fields
    Modifier and Type
    Field
    Description
    static final String
     
    static final String
     

    Fields inherited from class org.apache.pdfbox.text.PDFTextStripper

    charactersByArticle, document, LINE_SEPARATOR, output
  • Method Summary

    Modifier and Type
    Method
    Description
    void
    beginMarkedContentSequence(org.apache.pdfbox.cos.COSName tag, org.apache.pdfbox.cos.COSDictionary properties)
     
    protected float
    computeFontHeight(org.apache.pdfbox.pdmodel.font.PDFont arg0)
     
    protected void
    endDocument(org.apache.pdfbox.pdmodel.PDDocument pdf)
     
    void
     
    protected void
    endPage(org.apache.pdfbox.pdmodel.PDPage page)
     
    int
    we need to override this because we are overriding PDFTextStripper.processPages(PDPageTree)
    static void
    process(org.apache.pdfbox.pdmodel.PDDocument pdDocument, ContentHandler handler, ParseContext context, Metadata metadata, PDFParserConfig config, PagesConfig pages, org.apache.tika.parser.pdf.PageEmitter emitter, CompositeContentEnricher contentEnrichers)
    Converts the given PDF document (and related metadata) to a stream of XHTML SAX events sent to the given content handler.
    void
    processPage(org.apache.pdfbox.pdmodel.PDPage page)
     
    protected void
    processPages(org.apache.pdfbox.pdmodel.PDPageTree pages)
    See TIKA-2845 for why we need to override this.
    protected void
    processTextPosition(org.apache.pdfbox.text.TextPosition text)
     
    void
    setEndBookmark(org.apache.pdfbox.pdmodel.interactive.documentnavigation.outline.PDOutlineItem pdOutlineItem)
     
    void
    setStartBookmark(org.apache.pdfbox.pdmodel.interactive.documentnavigation.outline.PDOutlineItem pdOutlineItem)
     
    void
    showForm(org.apache.pdfbox.pdmodel.graphics.form.PDFormXObject form)
     
    protected void
    showGlyph(org.apache.pdfbox.util.Matrix textRenderingMatrix, org.apache.pdfbox.pdmodel.font.PDFont font, int code, org.apache.pdfbox.util.Vector displacement)
     
    void
    showTransparencyGroup(org.apache.pdfbox.pdmodel.graphics.form.PDTransparencyGroup form)
     
    protected void
    startDocument(org.apache.pdfbox.pdmodel.PDDocument pdf)
     
    protected void
    startPage(org.apache.pdfbox.pdmodel.PDPage page)
     
    protected void
    writeCharacters(org.apache.pdfbox.text.TextPosition text)
     
    protected void
     
    protected void
     
    protected void
     
    protected void
     
    protected void
     
    protected void
    writeString(String text, List<org.apache.pdfbox.text.TextPosition> textPositions)
     
    protected void
     

    Methods inherited from class org.apache.pdfbox.text.PDFTextStripper

    endArticle, getAddMoreFormatting, getArticleEnd, getArticleStart, getAverageCharTolerance, getCharactersByArticle, getDropThreshold, getEndBookmark, getEndPage, getIgnoreContentStreamSpaceGlyphs, getIndentThreshold, getLineSeparator, getListItemPatterns, getOutput, getPageEnd, getPageStart, getParagraphEnd, getParagraphStart, getSeparateByBeads, getSortByPosition, getSpacingTolerance, getStartBookmark, getStartPage, getSuppressDuplicateOverlappingText, getText, getWordSeparator, matchPattern, setAddMoreFormatting, setArticleEnd, setArticleStart, setAverageCharTolerance, setDropThreshold, setEndPage, setIgnoreContentStreamSpaceGlyphs, setIndentThreshold, setLineSeparator, setListItemPatterns, setPageEnd, setPageStart, setParagraphEnd, setParagraphStart, setShouldSeparateByBeads, setSortByPosition, setSpacingTolerance, setStartPage, setSuppressDuplicateOverlappingText, setWordSeparator, startArticle, startArticle, writePageEnd, writePageStart, writeParagraphSeparator, writeText

    Methods inherited from class org.apache.pdfbox.contentstream.PDFStreamEngine

    addOperator, applyTextAdjustment, beginText, decreaseLevel, endText, getAppearance, getCurrentPage, getGraphicsStackSize, getGraphicsState, getInitialMatrix, getLevel, getResources, getTextLineMatrix, getTextMatrix, increaseLevel, isShouldProcessColorOperators, markedContentPoint, operatorException, processAnnotation, processChildStream, processOperator, processOperator, processSoftMask, processTilingPattern, processTilingPattern, processTransparencyGroup, processType3Stream, restoreGraphicsStack, restoreGraphicsState, saveGraphicsStack, saveGraphicsState, setLineDashPattern, setTextLineMatrix, setTextMatrix, showAnnotation, showFontGlyph, showText, showTextString, showTextStrings, showType3Glyph, transformedPoint, transformWidth, unsupportedOperator

    Methods inherited from class java.lang.Object

    clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait
  • Field Details

  • Method Details

    • process

      public static void process(org.apache.pdfbox.pdmodel.PDDocument pdDocument, ContentHandler handler, ParseContext context, Metadata metadata, PDFParserConfig config, PagesConfig pages, org.apache.tika.parser.pdf.PageEmitter emitter, CompositeContentEnricher contentEnrichers) throws SAXException, TikaException
      Converts the given PDF document (and related metadata) to a stream of XHTML SAX events sent to the given content handler.
      Parameters:
      pdDocument - PDF document
      handler - SAX content handler
      context - parse context
      metadata - PDF metadata
      config - PDF parser config
      renderer - the renderer to use for rendering pages
      Throws:
      SAXException - if the content handler fails to process SAX events
      TikaException - if there was an exception outside of per page processing
    • beginMarkedContentSequence

      public void beginMarkedContentSequence(org.apache.pdfbox.cos.COSName tag, org.apache.pdfbox.cos.COSDictionary properties)
      Overrides:
      beginMarkedContentSequence in class org.apache.pdfbox.text.PDFTextStripper
    • endMarkedContentSequence

      public void endMarkedContentSequence()
      Overrides:
      endMarkedContentSequence in class org.apache.pdfbox.text.PDFTextStripper
    • showForm

      public void showForm(org.apache.pdfbox.pdmodel.graphics.form.PDFormXObject form) throws IOException
      Overrides:
      showForm in class org.apache.pdfbox.contentstream.PDFStreamEngine
      Throws:
      IOException
    • showTransparencyGroup

      public void showTransparencyGroup(org.apache.pdfbox.pdmodel.graphics.form.PDTransparencyGroup form) throws IOException
      Overrides:
      showTransparencyGroup in class org.apache.pdfbox.contentstream.PDFStreamEngine
      Throws:
      IOException
    • processTextPosition

      protected void processTextPosition(org.apache.pdfbox.text.TextPosition text)
      Overrides:
      processTextPosition in class org.apache.pdfbox.text.PDFTextStripper
    • startPage

      protected void startPage(org.apache.pdfbox.pdmodel.PDPage page) throws IOException
      Throws:
      IOException
    • writePage

      protected void writePage() throws IOException
      Overrides:
      writePage in class org.apache.pdfbox.text.PDFTextStripper
      Throws:
      IOException
    • endDocument

      protected void endDocument(org.apache.pdfbox.pdmodel.PDDocument pdf) throws IOException
      Throws:
      IOException
    • writeParagraphStart

      protected void writeParagraphStart() throws IOException
      Throws:
      IOException
    • writeParagraphEnd

      protected void writeParagraphEnd() throws IOException
      Throws:
      IOException
    • writeString

      protected void writeString(String text, List<org.apache.pdfbox.text.TextPosition> textPositions) throws IOException
      Overrides:
      writeString in class org.apache.pdfbox.text.PDFTextStripper
      Throws:
      IOException
    • writeWordSeparator

      protected void writeWordSeparator() throws IOException
      Throws:
      IOException
    • writeLineSeparator

      protected void writeLineSeparator() throws IOException
      Throws:
      IOException
    • processPage

      public void processPage(org.apache.pdfbox.pdmodel.PDPage page) throws IOException
      Overrides:
      processPage in class org.apache.pdfbox.text.PDFTextStripper
      Throws:
      IOException
    • endPage

      protected void endPage(org.apache.pdfbox.pdmodel.PDPage page) throws IOException
      Throws:
      IOException
    • writeString

      protected void writeString(String text) throws IOException
      Overrides:
      writeString in class org.apache.pdfbox.text.PDFTextStripper
      Throws:
      IOException
    • writeCharacters

      protected void writeCharacters(org.apache.pdfbox.text.TextPosition text) throws IOException
      Overrides:
      writeCharacters in class org.apache.pdfbox.text.PDFTextStripper
      Throws:
      IOException
    • startDocument

      protected void startDocument(org.apache.pdfbox.pdmodel.PDDocument pdf) throws IOException
      Overrides:
      startDocument in class org.apache.pdfbox.text.PDFTextStripper
      Throws:
      IOException
    • getCurrentPageNo

      public int getCurrentPageNo()
      we need to override this because we are overriding PDFTextStripper.processPages(PDPageTree)
      Overrides:
      getCurrentPageNo in class org.apache.pdfbox.text.PDFTextStripper
      Returns:
    • processPages

      protected void processPages(org.apache.pdfbox.pdmodel.PDPageTree pages) throws IOException
      See TIKA-2845 for why we need to override this.
      Overrides:
      processPages in class org.apache.pdfbox.text.PDFTextStripper
      Parameters:
      pages -
      Throws:
      IOException
    • setStartBookmark

      public void setStartBookmark(org.apache.pdfbox.pdmodel.interactive.documentnavigation.outline.PDOutlineItem pdOutlineItem)
      Overrides:
      setStartBookmark in class org.apache.pdfbox.text.PDFTextStripper
    • setEndBookmark

      public void setEndBookmark(org.apache.pdfbox.pdmodel.interactive.documentnavigation.outline.PDOutlineItem pdOutlineItem)
      Overrides:
      setEndBookmark in class org.apache.pdfbox.text.PDFTextStripper
    • showGlyph

      protected void showGlyph(org.apache.pdfbox.util.Matrix textRenderingMatrix, org.apache.pdfbox.pdmodel.font.PDFont font, int code, org.apache.pdfbox.util.Vector displacement) throws IOException
      Throws:
      IOException
    • computeFontHeight

      protected float computeFontHeight(org.apache.pdfbox.pdmodel.font.PDFont arg0) throws IOException
      Throws:
      IOException