You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
T2.1 hardening: strict parser allowlist in PSTikaTextConvertor (Tika 2.9.4) #137
T2.1 of the parent epic #73 — complete the Tika hardening started in #92 by replacing AutoDetectParser with an explicit parser allowlist in PSTikaTextConvertor. The allowlist is a CompositeParser built from exactly seven individual parsers (text, HTML, XML, PDF, legacy Office, OOXML, RTF). For any media type not claimed by one of those parsers, CompositeParser falls back to EmptyParser and the document is parsed as empty.
This closes the file-type-confusion CVE class in tika-parsers-standard-package:2.9.4 (the dominant class behind the remaining tika-* matches after the #92 input cap). A document with a misleading Content-Type header is parsed by no parser rather than by an unexpected parser, and the addition of a new Tika parser in a future minor upgrade does not silently enter the trusted surface.
Background
The previous implementation used new AutoDetectParser(m_tikaConfig), which loads every parser Tika discovers on the classpath via the standard service-loader mechanism. That includes image OCR, audio transcription, video metadata parsers, font parsers, mail parsers, etc. — far more surface than the project actually needs. The existing system/config/tika-config.xml provides some exclusions (ExecutableParser, SQLite3Parser, image/jpeg, application/pdf, application/x-sqlite3) but is layered insideAutoDetectParser rather than gating the parser set itself, and the second AutoDetectParser declaration for application/pdf and application/msword re-enables autodetection for those types.
The T2.1 follow-up is the right level of hardening: gate the parser set in the Java code (visible, reviewable, and explicit) and let the existing tika-config.xml continue to gate which parser is used once a type is allowed.
OOXMLParser — modern Office (.docx, .xlsx, .pptx and templates)
RTFParser — application/rtf (note: Tika 2.x moved RTFParser to org.apache.tika.parser.microsoft.rtf, not org.apache.tika.parser.rtf — the import is updated accordingly)
Replaced new AutoDetectParser(m_tikaConfig) with getStrictParser().
Class Javadoc documents the security contract, the allowlist composition, the relationship to the existing tika-config.xml restrictions, and the deliberate decision to preserve recursive embedded-document parsing (the indexer needs the text from inside OOXML / OLE2 compound documents).
CompositeParser falls back to EmptyParser for any media type not claimed by the seven allowlist parsers, so a document with a Content-Type that is not text/html/xml/pdf/office/rtf (or any of Tika's recognized subtypes) is parsed as empty rather than being handed to an unexpected parser.
What this PR does NOT do
No version bump.tika-parsers-standard-package:2.9.4 is the current version; the Tika 2.x line is on Java 1.8-compatible bytecode. The CVE class is closed by the parser-set restriction, not by a version change.
No changes to the tika-config.xml. The existing mime-exclude and parser-exclude restrictions remain in effect and apply on top of the new allowlist.
No recursive-parsing disable. The indexer needs to extract text from inside OOXML / OLE2 containers, so a DefaultEmbeddedDocumentExtractor override that returns false from shouldParseEmbedded() would lose legitimate functionality. The [security] T2.1 hardening: cap Tika input stream size (closes 15 CVEs in commons-tika 2.9.x) #92 input cap and the new allowlist together are sufficient.
No maxStringLength / maxEntityExpansion change. The existing WriteOutContentHandler(writeLimit) (5M chars by default, configurable via the indexWriteLimit server property) handles the string-length cap, and the T2.12 Xerces factory hardening closes the XML entity-expansion CVE class.
Acceptance criteria
PSTikaTextConvertor uses getStrictParser() (a CompositeParser over the seven allowlist parsers) instead of AutoDetectParser.
EmptyParser is the fallback for unknown types.
The seven parsers are explicitly listed in the code; adding a new Tika parser requires a code change.
Public method signature of getConvertedText(InputStream, String) is unchanged; existing callers do not need to be touched.
Full reactor mvn clean install is green on Java 1.8.
Out of scope
Java 11+ migration. The main branch stays on Java 1.8.
Adding tika-parser-rtf-module as a separate dependency (not needed — tika-parser-microsoft-module provides RTFParser at org.apache.tika.parser.microsoft.rtf).
Tighter tika-config.xml restrictions (separate file, separate PR).
Summary
T2.1 of the parent epic #73 — complete the Tika hardening started in #92 by replacing
AutoDetectParserwith an explicit parser allowlist inPSTikaTextConvertor. The allowlist is aCompositeParserbuilt from exactly seven individual parsers (text, HTML, XML, PDF, legacy Office, OOXML, RTF). For any media type not claimed by one of those parsers,CompositeParserfalls back toEmptyParserand the document is parsed as empty.This closes the file-type-confusion CVE class in
tika-parsers-standard-package:2.9.4(the dominant class behind the remainingtika-*matches after the #92 input cap). A document with a misleadingContent-Typeheader is parsed by no parser rather than by an unexpected parser, and the addition of a new Tika parser in a future minor upgrade does not silently enter the trusted surface.Background
The previous implementation used
new AutoDetectParser(m_tikaConfig), which loads every parser Tika discovers on the classpath via the standard service-loader mechanism. That includes image OCR, audio transcription, video metadata parsers, font parsers, mail parsers, etc. — far more surface than the project actually needs. The existingsystem/config/tika-config.xmlprovides some exclusions (ExecutableParser,SQLite3Parser,image/jpeg,application/pdf,application/x-sqlite3) but is layered insideAutoDetectParserrather than gating the parser set itself, and the secondAutoDetectParserdeclaration forapplication/pdfandapplication/mswordre-enables autodetection for those types.The T2.1 follow-up is the right level of hardening: gate the parser set in the Java code (visible, reviewable, and explicit) and let the existing
tika-config.xmlcontinue to gate which parser is used once a type is allowed.What changes
system/src/main/java/com/percussion/search/lucene/textconverter/PSTikaTextConvertor.java— 1 file, 80 insertions, 2 deletions:m_strictParser: cached instance of the strictCompositeParser, built once per JVM.getStrictParser(): synchronously builds the seven-parser allowlist:TXTParser— text/*HtmlParser— text/html, application/xhtml+xmlXMLParser— text/xml, application/xml, application/rdf+xmlPDFParser— application/pdfOfficeParser— legacy Office (.doc, .xls, .ppt via OLE2 / POIFS)OOXMLParser— modern Office (.docx, .xlsx, .pptx and templates)RTFParser— application/rtf (note: Tika 2.x movedRTFParsertoorg.apache.tika.parser.microsoft.rtf, notorg.apache.tika.parser.rtf— the import is updated accordingly)new AutoDetectParser(m_tikaConfig)withgetStrictParser().tika-config.xmlrestrictions, and the deliberate decision to preserve recursive embedded-document parsing (the indexer needs the text from inside OOXML / OLE2 compound documents).CompositeParserfalls back toEmptyParserfor any media type not claimed by the seven allowlist parsers, so a document with aContent-Typethat is not text/html/xml/pdf/office/rtf (or any of Tika's recognized subtypes) is parsed as empty rather than being handed to an unexpected parser.What this PR does NOT do
tika-parsers-standard-package:2.9.4is the current version; the Tika 2.x line is on Java 1.8-compatible bytecode. The CVE class is closed by the parser-set restriction, not by a version change.tika-config.xml. The existingmime-excludeandparser-excluderestrictions remain in effect and apply on top of the new allowlist.DefaultEmbeddedDocumentExtractoroverride that returnsfalsefromshouldParseEmbedded()would lose legitimate functionality. The [security] T2.1 hardening: cap Tika input stream size (closes 15 CVEs in commons-tika 2.9.x) #92 input cap and the new allowlist together are sufficient.maxStringLength/maxEntityExpansionchange. The existingWriteOutContentHandler(writeLimit)(5M chars by default, configurable via theindexWriteLimitserver property) handles the string-length cap, and the T2.12 Xerces factory hardening closes the XML entity-expansion CVE class.Acceptance criteria
PSTikaTextConvertorusesgetStrictParser()(aCompositeParserover the seven allowlist parsers) instead ofAutoDetectParser.EmptyParseris the fallback for unknown types.getConvertedText(InputStream, String)is unchanged; existing callers do not need to be touched.mvn clean installis green on Java 1.8.Out of scope
mainbranch stays on Java 1.8.tika-parser-rtf-moduleas a separate dependency (not needed —tika-parser-microsoft-moduleprovidesRTFParseratorg.apache.tika.parser.microsoft.rtf).tika-config.xmlrestrictions (separate file, separate PR).References
tika-config.xml:system/config/tika-config.xmlCompositeParserAPI:org.apache.tika.parser.CompositeParser(MediaTypeRegistry, List<Parser>)— fallback toEmptyParseris built-indocs/ai-generated/tasks/PR#-DependencyVulnerabilityAnalysis/categorized-final.json(keyNO_JAVA8_UPGRADE, GAVstika-*)