Repository navigation
Extraction-time constraints vs post-hoc validation #348
Description
Activity
Answer: both — one specification, two compilations
RelationType.start_labels/end_labels is authored once and compiled twice, to two mechanisms with different jobs:
| extraction-time | post-hoc | |
|---|---|---|
| compiles to | JointSchema head/tail typing → TypedEndpoints constraints |
Cypher check over Memgraph (enforce_relation_domain_range) |
| job | quality — steer the decoder | integrity — audit what landed |
| scope | GLiNER2 only, per window | every backend, whole graph |
| status | best-effort, never a guarantee | authoritative |
This directly revises the plan's line "self._relation_schema construction is unaffected by RelationType gaining optional fields" — it is affected, and that was the premise this map was chartered to re-examine.
Why this is not belt-and-braces
Two verified findings make the pair non-redundant.
Constrained decoding is not a filter. allow_edge is called inside the beam expansion loop (gliner2/joint_ie/optimizers/beam.py:67), before expanded.append. The beam reallocates, so a constrained run can surface a valid edge an unconstrained run never emits. You cannot recover that by filtering an unconstrained run afterwards — which is why post-hoc-only leaves real quality on the table.
Constraints do not survive the merge. Per #345, long_text.py drops cross-window relations and never re-checks constraints when merging. So the post-hoc pass is the only mechanism that ever sees the merged graph. That answers the question this ticket raised against the "both" option — what does a post-hoc check even mean over relations that were constrained at extraction time? It means: the merged graph is a different artifact from any window, and nothing else validates it.
This is ADR 0002's architecture, applied to relations
ADR 0002 already settled where enforcement lives, and settled it post-hoc: "even GLiNER2's closed schema is enforced by the model's inference, not by Memgraph. Enforcement lives entirely in unstructured2graph's own Python code." Steering is explicitly not enforcement. Relations change only the strength of the steering — hard pruning inside beam search rather than inference bias — not its status.
Same spec in both compilations — no deliberate loosening
Both read the identical start_labels/end_labels. Rejected: loosening extraction-time to permissive endpoints (that reverts the quality argument above), and per-relation confidence thresholds (a #353 question in disguise, and no evidence exists to set a threshold from).
This buys a sharp invariant worth stating in the plan: for GLiNER2, a post-hoc domain/range violation can only come from a cross-window merge or a bug. Violations become diagnostic rather than routine noise. If derived domain/range proves too tight, that is a derivation defect to fix in #353 — not something to paper over with a looser compile.
ADR 0004 gets amended, explicitly, to cover relations
The collision is real but it already exists on main: gliner2_backend.py:184 builds _entity_schema from ontology.entity_types, so GLiNER2 today cannot emit an out-of-ontology entity type. That is rejection-by-construction, shipped, sitting beside ADR 0004 without anyone calling it a contradiction — because 0004's actual mechanism is "kept, visible, queryable" with label promotion withheld (CONTEXT.md: "Ontology Conformance never rejects an entity. Only controls label promotion.").
Considered and rejected: leaving 0004 alone with a scoping sentence, on the grounds that it already behaves this way. Rejected because it papers over a genuine asymmetry rather than recording it.
The asymmetry the amendment must state. ADR 0004's second leg is "a later ontology change can retroactively promote labels for entities that didn't conform under an older version, no re-running extraction." A relation suppressed during decoding was never written, so it cannot be re-projected — re-extraction is the only recovery. #347 made the ontology derived automatically with no human approval, which raises the odds that a wrong constraint is what is doing the suppressing. That is a real, accepted cost and it belongs in the record, not in a footnote.
The amendment therefore says: relations are in scope; post-hoc non-conformance is handled identically to entities (kept, flagged, never deleted); extraction-time suppression is a permitted exception to never-reject, whose cost is loss of re-projection, mitigated by the fallback below.
Suppression: retry unconstrained on infeasibility, write flagged
We must drive windows ourselves. JointResult.feasible (result.py:75) is the one suppression signal gliner2 exposes — "False when the decoder could not satisfy the declared hard constraints and fell back to an empty assignment (distinct from 'nothing to extract')". But long_text.py:77 constructs its merged result as JointResult(text, entities, relations, include_confidence, include_spans) — feasible is omitted and defaults to True. Calling extract_long_text destroys the only signal we have. This is a second, independent reason to drive windows ourselves, alongside #345's cross-window relation dropping. Feeds #352.
On feasible=False: re-run that window with permissive endpoints and write what comes back, stamped as a fallback. Plus a warning with the chunk hash and a per-run counter in the reconciliation summary. One extra local forward pass on a rare path; no private API. This is never-reject applied to relations at the one point suppression is observable — the concrete mitigation the ADR amendment points at.
Per-edge suppression stays unrecorded. Rejected: diffing problem.edges against solution.edges. It depends on JointIEEngine internals that can shift between releases, and it cannot attribute a drop to a constraint rather than to gain < 0.0, edge_conflicts, or beam-width truncation — so it would produce a noisy log that reads as authoritative. The per-edge view is an eval-time artifact instead: a deliberate constrained-vs-permissive diff run. Handed to #350.
enforce_ontology gates the post-hoc half only
One flag, meaning unchanged: "validate what is in the graph against the ontology and flag non-conformance." It now covers relation domain/range alongside entity-label promotion, as the plan already has it (enforce_relation_domain_range called under the same if enforce_ontology:).
Extraction-time steering stays where ADR 0003 put it — configured by handing the backend an ontology (GLiNER2Backend(ontology=...)), a separate call site coordinated by path. Rejected: gating both halves with one flag (threads it into the backend constructor, the exact call-site coupling ADR 0003 refused, and would silently change extraction behaviour rather than validation). Rejected: a separate enforce_relation_constraints flag — ADR 0005 split flags because promotion and vocabulary-restriction were different concerns for the same caller; entity-label promotion and relation domain/range are the same concern at the same call site.
ontology_conformant keeps one writer and one meaning. The infeasible-window fallback gets its own property, written by the backend regardless of the flag. Otherwise the flag would carry two different facts — "these endpoints violate declared domain/range" (post-hoc pass) and "this window was infeasible, so this edge came from an unconstrained retry" (backend) — and CONTEXT.md defines Ontology Conformance too narrowly to absorb that.
Follow-ups
- Does LightRAG consume domain/range too #349 narrowed: post-hoc is backend-agnostic, so LightRAG gets domain/range validation for free. What remains is only whether LightRAG gets extraction-time steering.
- Wire a constrained ontology and read the edges #350 gains a brief: run constrained vs permissive and diff the edge sets; exercise the infeasible-window fallback.
- Chunking and the durability of constraints #352 gains a second reason to drive windows ourselves —
feasibledoes not surviveextract_long_text. - New ticket on the read-time contract for flagged edges: the "both" model deliberately accumulates known-wrong edges behind a flag, and nothing today says who filters it.
- Local
mainwas 6 commits stale while this was worked; ADR 0004 and ADR 0005 are shipped (context-graph/unstructured2graph: publish ADR/CONTEXT docs, consolidate numbering, close code/doc drift #337, merged 2026-09-18). Fetched, and all citations above are againstorigin/main.
#350 ran the prototype. Your core claim is confirmed; one mitigation is not reachable; and the post-merge invariant can be stated more sharply.
Confirmed — the beam reallocates rather than filters. Over 10 real sessions, 85 claims exist only in the constrained arm and they are real edges, not artifacts: purchased(User:'user' -> Product:"chef's knife"), prefers(User:'me' -> Product:'Shoeboxed'), practices(User:'me' -> Activity:'tracking expenses'). "Both compilations" over "post-hoc only" stands on evidence now, not only on reading beam.py. A third outcome also showed up that a text-keyed diff misses: 14 claims survive both arms with different endpoint typing (practices('user' -> 'routine') is (User,Topic) permissive, (User,Activity) constrained) — constrained decoding changes what an entity is, not just which edges exist.
Not reachable — the infeasible-window retry. 0/109 windows infeasible in both arms. feasible is False only when every candidate assignment fails validate_solution (optimizers/beam.py:96-109), and the empty solution always validates, so a prohibitive constraint can never force it. Probed on 12 windows:
| schema | infeasible | relations kept |
|---|---|---|
| domain/range only | 0/12 | 64 |
+ max_per_head=1, max_per_tail=1, no_self_loops, at_most(per_head=1) |
0/12 | 20 |
+ symmetric/inverse companions against those caps |
6/12 | 4 |
So the retry-permissive-and-flag path has no trigger under this spec shape. It becomes reachable only if a derived ontology declares symmetric/inverse relations — completeness constraints, whose companions are injected after screening — and then it fires on half the windows, each returning empty. The ReconciliationSummary counter you specified is still worth having (it is precisely the detector for a derived ontology growing completeness constraints), but it should be expected to read zero, and the fallback property will normally have no writer.
Row 2 is the uncomfortable one for the unrecorded-suppression decision: cardinality caps dropped 68% of relations (64 -> 20) while still reporting feasible=True. That is the quantity that goes unmeasured when a vocabulary carries caps.
Sharper invariant. You wrote that a GLiNER2 post-hoc violation means a cross-window merge or a bug. Measured: 0/1004 constrained edges violate the declared binding after the merge, against 194/672 permissive (29%). The reason is structural — typing is a per-edge predicate, so a set-union of per-window-valid edges is still type-valid; cardinality is a global predicate, so it is not (#345 saw max_per_head=1 end up with 2 edges on one head after merging). So for the domain/range half specifically, a post-hoc violation means a bug or an untyped backend, not a merge artifact — and the merge escape hatch only applies once cardinality enters the spec.
One more constraint on the spec shape, from building the schemas: a symmetric relation is rejected at build time unless its head and tail type sets are compatible — ValueError: symmetric relations require compatible head and tail types. In the plan's model that means a symmetric relation type forces start_labels == end_labels.
Full read and assets: #350 and branch prototype/gliner2-constrained-edges.
Revised by #355. The infeasible-window retry (permissive endpoints, written flagged) and its own fallback property leave the plan. #350 found feasible=False unreachable under domain/range and cardinality (0/109 windows), reachable only through symmetric/inverse, and #355 rules those out of derived ontologies (#353). A feasible=False is now treated as a bug and counted toward the integrity alarm: ontology_conformant count in ReconciliationSummary, asserted zero by an e2e test for GLiNER2. ontology_conformant keeps its single writer; the reason for the second property is gone.
Part of #344
Question
Do relation constraints apply during extraction, after the write, or both — and how does that square with ADR 0004?
The
hygmplan answers this implicitly and only one way:enforce_relation_domain_range()checks what is already in Memgraph and stampsontology_conformant = falseon mismatches. Post-hoc, never rejecting. It also states "self._relation_schemaconstruction is unaffected byRelationTypegaining optional fields" — i.e. domain/range never reaches the extractor.That line is what this map exists to revisit.
JointSchemacan apply head/tail typing, cardinality and self-loop rules during decoding, so invalid combinations are never produced. Those are genuinely different models, not two spellings of one:Open:
enforce_ontology's scope. One flag gates entity-label promotion today. Does it also gate relation constraints, and does its meaning stay coherent if one half is extraction-time and the other post-hoc?Blocked by #345: whether
JointSchemaenforces during decoding or merely filters afterwards decides whether "valid by construction" is real. If it is post-filtering, this decision mostly collapses.