XML External Entity (XXE) via Unsanitized Dictionary Parsing in Apache OpenNLP DictionaryEntryPersistor Versions Affected: before 2.5.9, before 3.0.0-M3 Description: The DictionaryEntryPersistor class initializes a static SAXParserFactory at class-load time without enabling FEATURE_SECURE_PROCESSING or disabling DTD processing. When create(InputStream, EntryInserter) is invoked, the only feature set on the XMLReader is namespace support — external entity resolution and DOCTYPE declarations remain fully enabled. An attacker who can supply a crafted dictionary file (e.g., a stop-word list or domain dictionary) containing a malicious DOCTYPE declaration can trigger local file disclosure via file:// entity references or server-side request forgery via http:// entity references during SAX parsing, before the application processes a single dictionary entry. This is inconsistent with the project's own XmlUtil.createSaxParser() helper, which correctly sets FEATURE_SECURE_PROCESSING and disallow-doctype-decl and is used by all other XML parsing paths in the codebase. The public Dictionary(InputStream) constructor delegates directly to this method and is the documented API for loading user-supplied dictionaries, making untrusted input a realistic scenario. Mitigation: 2.x users should upgrade to 2.5.9. 3.x users should upgrade to 3.0.0-M3. Users who cannot upgrade immediately should ensure that all dictionary files are sourced from trusted origins and should consider wrapping the Dictionary(InputStream) constructor with input validation that rejects any XML containing a DOCTYPE declaration before it reaches the parser.
The DictionaryEntryPersistor class creates a static SAXParserFactory instance during class loading without enabling the FEATURE_SECURE_PROCESSING flag or disabling DTD processing. The only feature set on the XMLReader object is namespace handling — external entity resolution and DOCTYPE declarations remain fully active. An attacker can embed a malicious DOCTYPE declaration in a dictionary file (e.g., a stop-words list or domain dictionary) containing references to external resources: file:// references allow reading files from the local system, and http:// references cause the server to initiate network requests (SSRF) — all of this occurs before any dictionary entry is processed. Other XML parsing paths in the project use the XmlUtil.createSaxParser() helper method, which correctly sets FEATURE_SECURE_PROCESSING and disallow-doctype-decl — the DictionaryEntryPersistor class is the sole exception.
An attacker can gain unauthorized access to sensitive files on the local operating system (e.g., /etc/passwd) or force the server to execute HTTP requests to internal network resources (SSRF), which may enable further infrastructure reconnaissance or attack escalation.
Users on the 2.x branch should update Apache OpenNLP to version 2.5.9, and users on the 3.x branch to version 3.0.0-M3. If immediate updating is not possible, ensure that all dictionary files come only from trusted sources, and consider implementing input validation that rejects XML documents containing DOCTYPE declarations before passing them to the parser.
Apache OpenNLP in versions before 2.5.9 (2.x branch) and before 3.0.0-M3 (3.x branch); the vulnerability affects any code using the public Dictionary(InputStream) constructor or the DictionaryEntryPersistor.create(InputStream, EntryInserter) method with untrusted dictionary files
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:NApache Opennlp
APPApache3.0.0< 2.5.9
Related vulnerabilities
Apache OpenNLP: arbitralne ładowanie klas przez manifest modelu (RCE-adjacent)
Apache OpenNLP — podatność XXE przy ładowaniu modeli i słowników XML
Apache OpenNLP: RCE przez niebezpieczną deserializację w SvmDoccatModel
OOM Denial of Service via Unbounded Array Allocation in Apache OpenNLP AbstractModelReader Versions Affected...
Arbitralne tworzenie instancji klasy przez Generator Descriptor XML i Format Name w Apache OpenNLP Wersje dot...