Class PDFMarkedContent2XHTML
The text stripper still does what it does in PDF2XHTML: it assembles words and lines
from glyph positions, with the same spacing heuristics and config. This class only adds a
marked-content stack over the content stream, so every glyph carries the MCID it was drawn
under, and records the stripper's word and separator calls for the page instead of writing
them. Once the page is assembled it either emits the words through the structure tree, or
replays them exactly as PDF2XHTML would have written them. A page falls back when
it has no tagged text, or under MarkedContentConfig.Strategy.AUTO when the tree
claims too little of the page's text or references content the page does not have.
Text the tree does not reference is not lost: untagged text and /Artifact content (running headers, footers, page numbers) follow the tagged content on each page in their own divs.
- Since:
- 1.24
-
Field Summary
FieldsFields inherited from class org.apache.pdfbox.text.PDFTextStripper
charactersByArticle, document, LINE_SEPARATOR, output -
Method Summary
Modifier and TypeMethodDescriptionvoidbeginMarkedContentSequence(org.apache.pdfbox.cos.COSName tag, org.apache.pdfbox.cos.COSDictionary properties) protected floatcomputeFontHeight(org.apache.pdfbox.pdmodel.font.PDFont arg0) protected voidendDocument(org.apache.pdfbox.pdmodel.PDDocument pdf) voidprotected voidendPage(org.apache.pdfbox.pdmodel.PDPage page) intwe need to override this because we are overridingPDFTextStripper.processPages(PDPageTree)static voidprocess(org.apache.pdfbox.pdmodel.PDDocument pdDocument, ContentHandler handler, ParseContext context, Metadata metadata, PDFParserConfig config, PagesConfig pages, org.apache.tika.parser.pdf.PageEmitter emitter, CompositeContentEnricher contentEnrichers) Converts the given PDF document (and related metadata) to a stream of XHTML SAX events sent to the given content handler.voidprocessPage(org.apache.pdfbox.pdmodel.PDPage page) protected voidprocessPages(org.apache.pdfbox.pdmodel.PDPageTree pages) See TIKA-2845 for why we need to override this.protected voidprocessTextPosition(org.apache.pdfbox.text.TextPosition text) voidsetEndBookmark(org.apache.pdfbox.pdmodel.interactive.documentnavigation.outline.PDOutlineItem pdOutlineItem) voidsetStartBookmark(org.apache.pdfbox.pdmodel.interactive.documentnavigation.outline.PDOutlineItem pdOutlineItem) voidshowForm(org.apache.pdfbox.pdmodel.graphics.form.PDFormXObject form) protected voidshowGlyph(org.apache.pdfbox.util.Matrix textRenderingMatrix, org.apache.pdfbox.pdmodel.font.PDFont font, int code, org.apache.pdfbox.util.Vector displacement) voidshowTransparencyGroup(org.apache.pdfbox.pdmodel.graphics.form.PDTransparencyGroup form) protected voidstartDocument(org.apache.pdfbox.pdmodel.PDDocument pdf) protected voidstartPage(org.apache.pdfbox.pdmodel.PDPage page) protected voidwriteCharacters(org.apache.pdfbox.text.TextPosition text) protected voidprotected voidprotected voidprotected voidprotected voidwriteString(String text) protected voidwriteString(String text, List<org.apache.pdfbox.text.TextPosition> textPositions) protected voidMethods inherited from class org.apache.pdfbox.text.PDFTextStripper
endArticle, getAddMoreFormatting, getArticleEnd, getArticleStart, getAverageCharTolerance, getCharactersByArticle, getDropThreshold, getEndBookmark, getEndPage, getIgnoreContentStreamSpaceGlyphs, getIndentThreshold, getLineSeparator, getListItemPatterns, getOutput, getPageEnd, getPageStart, getParagraphEnd, getParagraphStart, getSeparateByBeads, getSortByPosition, getSpacingTolerance, getStartBookmark, getStartPage, getSuppressDuplicateOverlappingText, getText, getWordSeparator, matchPattern, setAddMoreFormatting, setArticleEnd, setArticleStart, setAverageCharTolerance, setDropThreshold, setEndPage, setIgnoreContentStreamSpaceGlyphs, setIndentThreshold, setLineSeparator, setListItemPatterns, setPageEnd, setPageStart, setParagraphEnd, setParagraphStart, setShouldSeparateByBeads, setSortByPosition, setSpacingTolerance, setStartPage, setSuppressDuplicateOverlappingText, setWordSeparator, startArticle, startArticle, writePageEnd, writePageStart, writeParagraphSeparator, writeTextMethods inherited from class org.apache.pdfbox.contentstream.PDFStreamEngine
addOperator, applyTextAdjustment, beginText, decreaseLevel, endText, getAppearance, getCurrentPage, getGraphicsStackSize, getGraphicsState, getInitialMatrix, getLevel, getResources, getTextLineMatrix, getTextMatrix, increaseLevel, isShouldProcessColorOperators, markedContentPoint, operatorException, processAnnotation, processChildStream, processOperator, processOperator, processSoftMask, processTilingPattern, processTilingPattern, processTransparencyGroup, processType3Stream, restoreGraphicsStack, restoreGraphicsState, saveGraphicsStack, saveGraphicsState, setLineDashPattern, setTextLineMatrix, setTextMatrix, showAnnotation, showFontGlyph, showText, showTextString, showTextStrings, showType3Glyph, transformedPoint, transformWidth, unsupportedOperator
-
Field Details
-
XMP_DOCUMENT_CATALOG_LOCATION
- See Also:
-
XMP_PAGE_LOCATION_PREFIX
- See Also:
-
-
Method Details
-
process
public static void process(org.apache.pdfbox.pdmodel.PDDocument pdDocument, ContentHandler handler, ParseContext context, Metadata metadata, PDFParserConfig config, PagesConfig pages, org.apache.tika.parser.pdf.PageEmitter emitter, CompositeContentEnricher contentEnrichers) throws SAXException, TikaException Converts the given PDF document (and related metadata) to a stream of XHTML SAX events sent to the given content handler.- Parameters:
pdDocument- PDF documenthandler- SAX content handlercontext- parse contextmetadata- PDF metadataconfig- PDF parser configrenderer- the renderer to use for rendering pages- Throws:
SAXException- if the content handler fails to process SAX eventsTikaException- if there was an exception outside of per page processing
-
beginMarkedContentSequence
public void beginMarkedContentSequence(org.apache.pdfbox.cos.COSName tag, org.apache.pdfbox.cos.COSDictionary properties) - Overrides:
beginMarkedContentSequencein classorg.apache.pdfbox.text.PDFTextStripper
-
endMarkedContentSequence
public void endMarkedContentSequence()- Overrides:
endMarkedContentSequencein classorg.apache.pdfbox.text.PDFTextStripper
-
showForm
- Overrides:
showFormin classorg.apache.pdfbox.contentstream.PDFStreamEngine- Throws:
IOException
-
showTransparencyGroup
public void showTransparencyGroup(org.apache.pdfbox.pdmodel.graphics.form.PDTransparencyGroup form) throws IOException - Overrides:
showTransparencyGroupin classorg.apache.pdfbox.contentstream.PDFStreamEngine- Throws:
IOException
-
processTextPosition
protected void processTextPosition(org.apache.pdfbox.text.TextPosition text) - Overrides:
processTextPositionin classorg.apache.pdfbox.text.PDFTextStripper
-
startPage
- Throws:
IOException
-
writePage
- Overrides:
writePagein classorg.apache.pdfbox.text.PDFTextStripper- Throws:
IOException
-
endDocument
- Throws:
IOException
-
writeParagraphStart
- Throws:
IOException
-
writeParagraphEnd
- Throws:
IOException
-
writeString
protected void writeString(String text, List<org.apache.pdfbox.text.TextPosition> textPositions) throws IOException - Overrides:
writeStringin classorg.apache.pdfbox.text.PDFTextStripper- Throws:
IOException
-
writeWordSeparator
- Throws:
IOException
-
writeLineSeparator
- Throws:
IOException
-
processPage
- Overrides:
processPagein classorg.apache.pdfbox.text.PDFTextStripper- Throws:
IOException
-
endPage
- Throws:
IOException
-
writeString
- Overrides:
writeStringin classorg.apache.pdfbox.text.PDFTextStripper- Throws:
IOException
-
writeCharacters
- Overrides:
writeCharactersin classorg.apache.pdfbox.text.PDFTextStripper- Throws:
IOException
-
startDocument
- Overrides:
startDocumentin classorg.apache.pdfbox.text.PDFTextStripper- Throws:
IOException
-
getCurrentPageNo
public int getCurrentPageNo()we need to override this because we are overridingPDFTextStripper.processPages(PDPageTree)- Overrides:
getCurrentPageNoin classorg.apache.pdfbox.text.PDFTextStripper- Returns:
-
processPages
See TIKA-2845 for why we need to override this.- Overrides:
processPagesin classorg.apache.pdfbox.text.PDFTextStripper- Parameters:
pages-- Throws:
IOException
-
showGlyph
protected void showGlyph(org.apache.pdfbox.util.Matrix textRenderingMatrix, org.apache.pdfbox.pdmodel.font.PDFont font, int code, org.apache.pdfbox.util.Vector displacement) throws IOException - Throws:
IOException
-
computeFontHeight
- Throws:
IOException
-