Skip to content

Extraction-time constraints vs post-hoc validation #348

Description

@antejavor

Part of #344

Question

Do relation constraints apply during extraction, after the write, or both — and how does that square with ADR 0004?

The hygm plan answers this implicitly and only one way: enforce_relation_domain_range() checks what is already in Memgraph and stamps ontology_conformant = false on mismatches. Post-hoc, never rejecting. It also states "self._relation_schema construction is unaffected by RelationType gaining optional fields" — i.e. domain/range never reaches the extractor.

That line is what this map exists to revisit. JointSchema can apply head/tail typing, cardinality and self-loop rules during decoding, so invalid combinations are never produced. Those are genuinely different models, not two spellings of one:

  • Extraction-time only — cheapest graph, nothing wrong ever written. But the model silently withholds what it saw, there is no record of a suppressed relation, and it only works for backends that accept a schema (GLiNER2 today, not LightRAG).
  • Post-hoc only — the plan as written. Backend-agnostic, keeps everything, violations are visible and queryable. But the graph accumulates known-wrong edges carrying a flag every reader must remember to filter on.
  • Both — constrain where the backend supports it, validate everything afterwards regardless. More machinery, and it needs an answer for what a post-hoc check even means over relations that were constrained at extraction time.

Open:

  1. Which model, and why.
  2. The ADR 0004 collision. ADR 0004 is never-reject-entities-for-ontology-non-conformance — a non-conforming entity is always kept and stamped. Constrained decoding rejects by construction. Does ADR 0004's reasoning extend to relations, and if so does that rule out extraction-time constraints, or does the ADR need amending? If it needs amending, say so explicitly — it is a shipped decision.
  3. Observability of suppression. If the model declines to emit a relation because of a constraint, is that recorded anywhere? An ontology that silently suppresses real signal is hard to debug and hard to iterate on.
  4. enforce_ontology's scope. One flag gates entity-label promotion today. Does it also gate relation constraints, and does its meaning stay coherent if one half is extraction-time and the other post-hoc?

Blocked by #345: whether JointSchema enforces during decoding or merely filters afterwards decides whether "valid by construction" is real. If it is post-filtering, this decision mostly collapses.

Activity

self-assigned this
on Sep 21, 2026

antejavor commented on Sep 21, 2026

@antejavor
ContributorAuthor

Answer: both — one specification, two compilations

RelationType.start_labels/end_labels is authored once and compiled twice, to two mechanisms with different jobs:

extraction-time post-hoc
compiles to JointSchema head/tail typing → TypedEndpoints constraints Cypher check over Memgraph (enforce_relation_domain_range)
job quality — steer the decoder integrity — audit what landed
scope GLiNER2 only, per window every backend, whole graph
status best-effort, never a guarantee authoritative

This directly revises the plan's line "self._relation_schema construction is unaffected by RelationType gaining optional fields" — it is affected, and that was the premise this map was chartered to re-examine.

Why this is not belt-and-braces

Two verified findings make the pair non-redundant.

Constrained decoding is not a filter. allow_edge is called inside the beam expansion loop (gliner2/joint_ie/optimizers/beam.py:67), before expanded.append. The beam reallocates, so a constrained run can surface a valid edge an unconstrained run never emits. You cannot recover that by filtering an unconstrained run afterwards — which is why post-hoc-only leaves real quality on the table.

Constraints do not survive the merge. Per #345, long_text.py drops cross-window relations and never re-checks constraints when merging. So the post-hoc pass is the only mechanism that ever sees the merged graph. That answers the question this ticket raised against the "both" option — what does a post-hoc check even mean over relations that were constrained at extraction time? It means: the merged graph is a different artifact from any window, and nothing else validates it.

This is ADR 0002's architecture, applied to relations

ADR 0002 already settled where enforcement lives, and settled it post-hoc: "even GLiNER2's closed schema is enforced by the model's inference, not by Memgraph. Enforcement lives entirely in unstructured2graph's own Python code." Steering is explicitly not enforcement. Relations change only the strength of the steering — hard pruning inside beam search rather than inference bias — not its status.


Same spec in both compilations — no deliberate loosening

Both read the identical start_labels/end_labels. Rejected: loosening extraction-time to permissive endpoints (that reverts the quality argument above), and per-relation confidence thresholds (a #353 question in disguise, and no evidence exists to set a threshold from).

This buys a sharp invariant worth stating in the plan: for GLiNER2, a post-hoc domain/range violation can only come from a cross-window merge or a bug. Violations become diagnostic rather than routine noise. If derived domain/range proves too tight, that is a derivation defect to fix in #353 — not something to paper over with a looser compile.


ADR 0004 gets amended, explicitly, to cover relations

The collision is real but it already exists on main: gliner2_backend.py:184 builds _entity_schema from ontology.entity_types, so GLiNER2 today cannot emit an out-of-ontology entity type. That is rejection-by-construction, shipped, sitting beside ADR 0004 without anyone calling it a contradiction — because 0004's actual mechanism is "kept, visible, queryable" with label promotion withheld (CONTEXT.md: "Ontology Conformance never rejects an entity. Only controls label promotion.").

Considered and rejected: leaving 0004 alone with a scoping sentence, on the grounds that it already behaves this way. Rejected because it papers over a genuine asymmetry rather than recording it.

The asymmetry the amendment must state. ADR 0004's second leg is "a later ontology change can retroactively promote labels for entities that didn't conform under an older version, no re-running extraction." A relation suppressed during decoding was never written, so it cannot be re-projected — re-extraction is the only recovery. #347 made the ontology derived automatically with no human approval, which raises the odds that a wrong constraint is what is doing the suppressing. That is a real, accepted cost and it belongs in the record, not in a footnote.

The amendment therefore says: relations are in scope; post-hoc non-conformance is handled identically to entities (kept, flagged, never deleted); extraction-time suppression is a permitted exception to never-reject, whose cost is loss of re-projection, mitigated by the fallback below.


Suppression: retry unconstrained on infeasibility, write flagged

We must drive windows ourselves. JointResult.feasible (result.py:75) is the one suppression signal gliner2 exposes — "False when the decoder could not satisfy the declared hard constraints and fell back to an empty assignment (distinct from 'nothing to extract')". But long_text.py:77 constructs its merged result as JointResult(text, entities, relations, include_confidence, include_spans) — feasible is omitted and defaults to True. Calling extract_long_text destroys the only signal we have. This is a second, independent reason to drive windows ourselves, alongside #345's cross-window relation dropping. Feeds #352.

On feasible=False: re-run that window with permissive endpoints and write what comes back, stamped as a fallback. Plus a warning with the chunk hash and a per-run counter in the reconciliation summary. One extra local forward pass on a rare path; no private API. This is never-reject applied to relations at the one point suppression is observable — the concrete mitigation the ADR amendment points at.

Per-edge suppression stays unrecorded. Rejected: diffing problem.edges against solution.edges. It depends on JointIEEngine internals that can shift between releases, and it cannot attribute a drop to a constraint rather than to gain < 0.0, edge_conflicts, or beam-width truncation — so it would produce a noisy log that reads as authoritative. The per-edge view is an eval-time artifact instead: a deliberate constrained-vs-permissive diff run. Handed to #350.


enforce_ontology gates the post-hoc half only

One flag, meaning unchanged: "validate what is in the graph against the ontology and flag non-conformance." It now covers relation domain/range alongside entity-label promotion, as the plan already has it (enforce_relation_domain_range called under the same if enforce_ontology:).

Extraction-time steering stays where ADR 0003 put it — configured by handing the backend an ontology (GLiNER2Backend(ontology=...)), a separate call site coordinated by path. Rejected: gating both halves with one flag (threads it into the backend constructor, the exact call-site coupling ADR 0003 refused, and would silently change extraction behaviour rather than validation). Rejected: a separate enforce_relation_constraints flag — ADR 0005 split flags because promotion and vocabulary-restriction were different concerns for the same caller; entity-label promotion and relation domain/range are the same concern at the same call site.

ontology_conformant keeps one writer and one meaning. The infeasible-window fallback gets its own property, written by the backend regardless of the flag. Otherwise the flag would carry two different facts — "these endpoints violate declared domain/range" (post-hoc pass) and "this window was infeasible, so this edge came from an unconstrained retry" (backend) — and CONTEXT.md defines Ontology Conformance too narrowly to absorb that.


Follow-ups

antejavor commented on Sep 21, 2026

@antejavor
ContributorAuthor

#350 ran the prototype. Your core claim is confirmed; one mitigation is not reachable; and the post-merge invariant can be stated more sharply.

Confirmed — the beam reallocates rather than filters. Over 10 real sessions, 85 claims exist only in the constrained arm and they are real edges, not artifacts: purchased(User:'user' -> Product:"chef's knife"), prefers(User:'me' -> Product:'Shoeboxed'), practices(User:'me' -> Activity:'tracking expenses'). "Both compilations" over "post-hoc only" stands on evidence now, not only on reading beam.py. A third outcome also showed up that a text-keyed diff misses: 14 claims survive both arms with different endpoint typing (practices('user' -> 'routine') is (User,Topic) permissive, (User,Activity) constrained) — constrained decoding changes what an entity is, not just which edges exist.

Not reachable — the infeasible-window retry. 0/109 windows infeasible in both arms. feasible is False only when every candidate assignment fails validate_solution (optimizers/beam.py:96-109), and the empty solution always validates, so a prohibitive constraint can never force it. Probed on 12 windows:

schema infeasible relations kept
domain/range only 0/12 64
+ max_per_head=1, max_per_tail=1, no_self_loops, at_most(per_head=1) 0/12 20
+ symmetric/inverse companions against those caps 6/12 4

So the retry-permissive-and-flag path has no trigger under this spec shape. It becomes reachable only if a derived ontology declares symmetric/inverse relations — completeness constraints, whose companions are injected after screening — and then it fires on half the windows, each returning empty. The ReconciliationSummary counter you specified is still worth having (it is precisely the detector for a derived ontology growing completeness constraints), but it should be expected to read zero, and the fallback property will normally have no writer.

Row 2 is the uncomfortable one for the unrecorded-suppression decision: cardinality caps dropped 68% of relations (64 -> 20) while still reporting feasible=True. That is the quantity that goes unmeasured when a vocabulary carries caps.

Sharper invariant. You wrote that a GLiNER2 post-hoc violation means a cross-window merge or a bug. Measured: 0/1004 constrained edges violate the declared binding after the merge, against 194/672 permissive (29%). The reason is structural — typing is a per-edge predicate, so a set-union of per-window-valid edges is still type-valid; cardinality is a global predicate, so it is not (#345 saw max_per_head=1 end up with 2 edges on one head after merging). So for the domain/range half specifically, a post-hoc violation means a bug or an untyped backend, not a merge artifact — and the merge escape hatch only applies once cardinality enters the spec.

One more constraint on the spec shape, from building the schemas: a symmetric relation is rejected at build time unless its head and tail type sets are compatible — ValueError: symmetric relations require compatible head and tail types. In the plan's model that means a symmetric relation type forces start_labels == end_labels.

Full read and assets: #350 and branch prototype/gliner2-constrained-edges.

antejavor commented on Sep 28, 2026

@antejavor
ContributorAuthor

Revised by #355. The infeasible-window retry (permissive endpoints, written flagged) and its own fallback property leave the plan. #350 found feasible=False unreachable under domain/range and cardinality (0/109 windows), reachable only through symmetric/inverse, and #355 rules those out of derived ontologies (#353). A feasible=False is now treated as a bug and counted toward the integrity alarm: ontology_conformant count in ReconciliationSummary, asserted zero by an e2e test for GLiNER2. ontology_conformant keeps its single writer; the reason for the second property is gone.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions