Skip to content

T2.1 hardening: strict parser allowlist in PSTikaTextConvertor (Tika 2.9.4) #137

Description

@natechadwick-intsof

Summary

T2.1 of the parent epic #73 — complete the Tika hardening started in #92 by replacing AutoDetectParser with an explicit parser allowlist in PSTikaTextConvertor. The allowlist is a CompositeParser built from exactly seven individual parsers (text, HTML, XML, PDF, legacy Office, OOXML, RTF). For any media type not claimed by one of those parsers, CompositeParser falls back to EmptyParser and the document is parsed as empty.

This closes the file-type-confusion CVE class in tika-parsers-standard-package:2.9.4 (the dominant class behind the remaining tika-* matches after the #92 input cap). A document with a misleading Content-Type header is parsed by no parser rather than by an unexpected parser, and the addition of a new Tika parser in a future minor upgrade does not silently enter the trusted surface.

Background

The previous implementation used new AutoDetectParser(m_tikaConfig), which loads every parser Tika discovers on the classpath via the standard service-loader mechanism. That includes image OCR, audio transcription, video metadata parsers, font parsers, mail parsers, etc. — far more surface than the project actually needs. The existing system/config/tika-config.xml provides some exclusions (ExecutableParser, SQLite3Parser, image/jpeg, application/pdf, application/x-sqlite3) but is layered inside AutoDetectParser rather than gating the parser set itself, and the second AutoDetectParser declaration for application/pdf and application/msword re-enables autodetection for those types.

The T2.1 follow-up is the right level of hardening: gate the parser set in the Java code (visible, reviewable, and explicit) and let the existing tika-config.xml continue to gate which parser is used once a type is allowed.

What changes

system/src/main/java/com/percussion/search/lucene/textconverter/PSTikaTextConvertor.java — 1 file, 80 insertions, 2 deletions:

  • New static field m_strictParser: cached instance of the strict CompositeParser, built once per JVM.
  • New private method getStrictParser(): synchronously builds the seven-parser allowlist:
    • TXTParser — text/*
    • HtmlParser — text/html, application/xhtml+xml
    • XMLParser — text/xml, application/xml, application/rdf+xml
    • PDFParser — application/pdf
    • OfficeParser — legacy Office (.doc, .xls, .ppt via OLE2 / POIFS)
    • OOXMLParser — modern Office (.docx, .xlsx, .pptx and templates)
    • RTFParser — application/rtf (note: Tika 2.x moved RTFParser to org.apache.tika.parser.microsoft.rtf, not org.apache.tika.parser.rtf — the import is updated accordingly)
  • Replaced new AutoDetectParser(m_tikaConfig) with getStrictParser().
  • Class Javadoc documents the security contract, the allowlist composition, the relationship to the existing tika-config.xml restrictions, and the deliberate decision to preserve recursive embedded-document parsing (the indexer needs the text from inside OOXML / OLE2 compound documents).

CompositeParser falls back to EmptyParser for any media type not claimed by the seven allowlist parsers, so a document with a Content-Type that is not text/html/xml/pdf/office/rtf (or any of Tika's recognized subtypes) is parsed as empty rather than being handed to an unexpected parser.

What this PR does NOT do

  • No version bump. tika-parsers-standard-package:2.9.4 is the current version; the Tika 2.x line is on Java 1.8-compatible bytecode. The CVE class is closed by the parser-set restriction, not by a version change.
  • No changes to the tika-config.xml. The existing mime-exclude and parser-exclude restrictions remain in effect and apply on top of the new allowlist.
  • No recursive-parsing disable. The indexer needs to extract text from inside OOXML / OLE2 containers, so a DefaultEmbeddedDocumentExtractor override that returns false from shouldParseEmbedded() would lose legitimate functionality. The [security] T2.1 hardening: cap Tika input stream size (closes 15 CVEs in commons-tika 2.9.x) #92 input cap and the new allowlist together are sufficient.
  • No maxStringLength / maxEntityExpansion change. The existing WriteOutContentHandler(writeLimit) (5M chars by default, configurable via the indexWriteLimit server property) handles the string-length cap, and the T2.12 Xerces factory hardening closes the XML entity-expansion CVE class.

Acceptance criteria

  • PSTikaTextConvertor uses getStrictParser() (a CompositeParser over the seven allowlist parsers) instead of AutoDetectParser.
  • EmptyParser is the fallback for unknown types.
  • The seven parsers are explicitly listed in the code; adding a new Tika parser requires a code change.
  • Public method signature of getConvertedText(InputStream, String) is unchanged; existing callers do not need to be touched.
  • Full reactor mvn clean install is green on Java 1.8.

Out of scope

  • Java 11+ migration. The main branch stays on Java 1.8.
  • Adding tika-parser-rtf-module as a separate dependency (not needed — tika-parser-microsoft-module provides RTFParser at org.apache.tika.parser.microsoft.rtf).
  • Tighter tika-config.xml restrictions (separate file, separate PR).

References

Co-Authored by Mavis Mavis-Code using MiniMax-M3 with agent mavis.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions