Skip to content

fix: lenient metadata extraction for unbound XML prefixes (#2355) - #2375

Merged
natechadwick merged 1 commit into
mainfrom
fix/issue-2355-unbound-prefix-metadata
Aug 8, 2026
Merged

natechadwick merged 1 commit into
mainfrom
fix/issue-2355-unbound-prefix-metadata

Conversation

@natechadwick-intsof

Copy link
Copy Markdown
Collaborator

Summary

Fixes #2355: PSMetadataExtractorService no longer fails the whole page's metadata path when published HTML contains unbound XML-style prefixes such as Google Custom Search <gcse:search> without xmlns:gcse.

Root cause

RDFa parse (Semargl/RDF4J SAX) treats vendor tags like gcse:search as namespaced elements. Without a matching xmlns: declaration this throws SAXParseException ("prefix is not bound"), which was wrapped as a hard RuntimeException and logged as ERROR by PSMetadataDeliveryHandler even when file publish succeeded.

Fix

  1. Pre-sanitize after Jsoup load: unwrap elements and strip attributes that use XML-style prefixes not declared via xmlns:* (always allow xml/xmlns). Log WARN with page path + unbound prefix set.
  2. Defensive catch: if RDFa parse still reports an unbound-prefix SAX failure, log WARN (page path + prefix) and continue with non-RDFa fields instead of failing the item.
  3. Preserve RDFa: dcterms:* / property values are attribute values, not unbound element names, so they still extract after strip.

Operator: Grok: night-issue-prs (model grok-4.5)

Test plan

  • Unit fixture unbound-prefix-gcse.html with <gcse:search> (no xmlns) + dcterms elsewhere
  • testUnboundPrefixGcseSearch — no throw; dcterms:source/title/description/abstract extracted
  • testUnboundPrefixOnlyDoesNotThrow — minimal gcse-only HTML returns entry with type page
  • Existing PSMetadataExtractorServiceTests still green

Gates evidence

  • Erlang-style self-review: no new path I/O; sanitize is pure DOM; WARN includes page path + prefix; change-class companions = service + behavioral tests + fixture (no Spring/DI surface change)
  • Module: cd system && ../mvnw clean install → BUILD SUCCESS
  • Commit: 2352d9cbc6598b47958ce84d3ca0a509341dd136

Related

Co-Authored by Grok Build using grok-4.5 with agent main.

Pre-sanitize unknown XML-style prefixed elements/attributes (e.g. gcse:search
without xmlns:gcse) before RDFa SAX parse, and catch residual unbound-prefix
parse failures as WARN so DTS publish metadata does not ERROR the whole page.
dcterms/RDFa elsewhere on the page continues to extract. Adds fixture + tests.

> Co-Authored by Grok Build using grok-4.5 with agent main.
@natechadwick
natechadwick merged commit 4d621d5 into main Aug 8, 2026
6 checks passed
@natechadwick
natechadwick deleted the fix/issue-2355-unbound-prefix-metadata branch August 8, 2026 01:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model:grok-4.5 Grok 4.5 model operator:grok Changes authored by Grok operator:night-issue-prs night-issue-prs workflow

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Metadata extraction fails on unbound prefixes (e.g. gcse:search) during DTS publish

2 participants