Repository navigation
fix: lenient metadata extraction for unbound XML prefixes (#2355) - #2375
Merged
Merged
Conversation
Pre-sanitize unknown XML-style prefixed elements/attributes (e.g. gcse:search without xmlns:gcse) before RDFa SAX parse, and catch residual unbound-prefix parse failures as WARN so DTS publish metadata does not ERROR the whole page. dcterms/RDFa elsewhere on the page continues to extract. Adds fixture + tests. > Co-Authored by Grok Build using grok-4.5 with agent main.
4 tasks
natechadwick
approved these changes
Aug 8, 2026
4 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #2355:
PSMetadataExtractorServiceno longer fails the whole page's metadata path when published HTML contains unbound XML-style prefixes such as Google Custom Search<gcse:search>withoutxmlns:gcse.Root cause
RDFa parse (Semargl/RDF4J SAX) treats vendor tags like
gcse:searchas namespaced elements. Without a matchingxmlns:declaration this throwsSAXParseException("prefix is not bound"), which was wrapped as a hardRuntimeExceptionand logged as ERROR byPSMetadataDeliveryHandlereven when file publish succeeded.Fix
xmlns:*(always allowxml/xmlns). Log WARN with page path + unbound prefix set.dcterms:*/ property values are attribute values, not unbound element names, so they still extract after strip.Operator: Grok: night-issue-prs (model grok-4.5)
Test plan
unbound-prefix-gcse.htmlwith<gcse:search>(no xmlns) + dcterms elsewheretestUnboundPrefixGcseSearch— no throw;dcterms:source/title/description/abstractextractedtestUnboundPrefixOnlyDoesNotThrow— minimal gcse-only HTML returns entry with typepagePSMetadataExtractorServiceTestsstill greenGates evidence
cd system && ../mvnw clean install→ BUILD SUCCESS2352d9cbc6598b47958ce84d3ca0a509341dd136Related
xmlnsfor clean publish metadata