You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Map: Typed relation model for context-graph extraction #344
The hygm shared-ontology plan, revised until every open decision is settled and it can be implemented without further design work.
The plan under revision lives at ~/.claude/plans/cuddly-bubbling-dijkstra.md (local to @antejavor). Its structural core — RelationType gaining start_labels/end_labels, three production strategies (manual YAML / OWL import / LLM recommendation), unstructured2graph/ontology.py as a thin adapter — is not up for re-decision here. This map resolves what that plan left open or got wrong.
Why this map exists
A 100-question GLiNER2 eval run produced a graph with zero semantic edges. Diagnosis found three nested causes:
context_graph_eval.reconcile builds GLiNER2Backend() with no ontology, so it falls back to DEFAULT_ONTOLOGY — 11 entity types, zero relation types. _extract_sync's if self._relation_schema: guard never fires, so relation extraction never runs at all.
Even when relation types are configured, GLiNER2Backend calls gliner2's coarse Schema.relations(dict) builder, whose source always writes {"head": "", "tail": ""} — no type constraint, no cardinality. That is the pre-2.5 mode fastino's own write-up contrasts against: "relation extraction returned independently thresholded triples and consistency was the user's responsibility."
The richer API already ships in the installed gliner2==2.0.0, unused: gliner2.joint_ie.schema.JointSchema, with .relation(name, head=(...), tail=(...), max_per_head=, max_per_tail=, symmetric=, inverse=, allow_self=) and .no_self_loops().
unstructured2graph.ontology.RelationType is (label, description) — there is nowhere to put a head/tail constraint even if we wanted to. That is the same gap blocking the hygm/OWL unification, which is why this is one effort and not two.
Convergent evidence for the shape: the plan chose start_labels/end_labels because OWL's rdfs:domain/rdfs:range map onto them; sql2graph's HyGM independently arrived at start_node_labels/end_node_labels; GLiNER2's RelationSpec(head=, tail=) is the same concept a third time.
Notes
Domain: ontology / graph modeling across unstructured2graph, the planned hygm package, and the context-graph family.
Skills: consult /grilling and /domain-modeling each session. Read the plan file in full before working any ticket.
Decision criterion: human judgment on sampled edges — not eval score deltas. Per eval: the judge scores identical answers differently, which is what makes the passing set unstable #324, the eval cannot resolve a small delta: three repeat runs against a frozen graph gave 15% / 10% / 15% coverage with zero questions passing in all three, because the retrieval agent emits different Cypher every run even at temperature=0. The eval is a confirmation step after implementation, never the gate on a modeling decision.
Primary backend: GLiNER2 is the backend this effort targets and the one the codebase is being adapted toward. LightRAGBackend is a retained, non-interchangeable alternative pipeline — interchangeable in process, never in results — kept rather than deleted, but not invested in. Established while resolving Does LightRAG consume domain/range too #349; it is why the typed relation model is GLiNER2-only by design rather than by omission.
Plan, don't do: this map produces decisions. Implementation happens after the revised plan is handed off.
Decisions so far
What JointSchema actually enforces — Constraints ARE enforced during decoding (optimizers/beam.py:67, inside beam expansion), and typing even earlier at candidate construction — but per window only: long_text.py drops any relation whose endpoints came from different chunks and never re-checks constraints at the merge. Entry point is JointIE/JointIEEngine, notAutoExtractor.extract(); compile_schema().build() silently drops all typing. Output is JointResult with entity-id endpoints — replaces the whole span-matching apparatus. "Unconstrained" is not expressible: empty head/tail raises, so the plan's start_labels=() compat rule must be translated to "all declared types" in our adapter.
The starter relation vocabulary — There is no hand-authored vocabulary. Measured the corpus: 42% of questions turn on time/change, 98/100 are first-person. So relations are time-anchored (valid_at from the session, stamped by our writer, not extracted); superseded facts are retained, timestamps only (a real question asks for the initial value, so "newest wins" would delete the answer); and the ontology is derived per corpus by LlmRecommendationStrategy, fully automatic with no human approval, gated by hygm.validation self-validation (hard) plus a stability check (reports, never blocks). User becomes an extracted entity type (verified: GLiNER2 resolves "I" to it), and processing — never collection — collapses all User mentions onto the session's (:User) node, synthesising anon:<session_id> when collection recorded none.
Extraction-time constraints vs post-hoc validation — Both: one specification, two compilations.start_labels/end_labels is authored once and compiled to JointSchema typing at extraction time (job: quality, steer the decoder; GLiNER2-only, best-effort) and to a Cypher check over Memgraph post-hoc (job: integrity; every backend, authoritative). Non-redundant for two verified reasons: allow_edge runs inside beam expansion so constrained decoding reallocates rather than filters, and per What JointSchema actually enforces #345 constraints never survive the window merge — so post-hoc is the only thing that sees the merged graph. Same spec in both, no deliberate loosening, which buys the invariant that a GLiNER2 post-hoc violation means a cross-window merge or a bug. This revises the plan's "self._relation_schema construction is unaffected". ADR 0004 is amended to cover relations explicitly, recording that decode-time suppression forfeits re-projection — the cost The starter relation vocabulary #347's automatic derivation makes likelier. Mitigation: long_text.py:77 drops feasible, so we drive windows ourselves, and an infeasible window is retried with permissive endpoints and written flagged; per-edge suppression stays unrecorded (private API, unattributable) and moves to Wire a constrained ontology and read the edges #350 as an eval-time diff. enforce_ontology gates the post-hoc half only, meaning unchanged; the fallback gets its own property so ontology_conformant keeps one writer.
Entity identity across chunks — Exact normalized text within type, per-type opt-in, merged at write time._entity_id drops chunk_hash for opted-in types and the existing MERGE does the work — no second pass. Rejected embeddings (non-deterministic across model versions, breaks eval reproducibility) and uniform merging (12,583 Concept vs 3,809 Person: merging every identically-worded Concept builds hubs that discriminate nothing). The file_path blocker in _entity_id's docstring is shallower than documented — ON CREATE SET means a re-merged node keeps the first chunk's file_path, so the fix is writing MENTIONED_IN at ingest, which also retires a full-graph cartesian product for this backend. Because every mention keeps its edge, merging destroys nothing and ADR 0004's instinct holds rather than collides — unlike Extraction-time constraints vs post-hoc validation #348's decode-time suppression. No confidence flag or cap; a reporting-only highest-degree signal catches supernodes. Identity stays per-backend (LightRAG keeps LLM consolidation), with only the per-type mergeability flag shared. Prerequisite filed as unstructured2graph: entity_id is MERGEd on but never indexed or constrained #357: entity_id is MERGEd on but never indexed or constrained.
Does LightRAG consume domain/range too — No, at neither end — by design, not omission. LightRAG's extraction record has no relation type field (relation|source|target|keywords|description) and its prompt orders relationships treated as undirected, so domain/range is not expressible; storage writes one untyped MERGE (a)-[r:DIRECTED]-(b). Extraction-time steering is therefore impossible, and post-hoc validation — which keys on the type label — is vacuous, not free, revising Extraction-time constraints vs post-hoc validation #348's follow-up note and the plan's §2 claim (the action, no code change, stands; the claim of coverage does not). The backends are interchangeable in process, never in results, so the asymmetry is stated rather than reconciled: no capability flag, no Protocol change. What ships instead is an observational coverage signal in enforce_relation_domain_range reading the graph rather than the backend — discriminating "N relationships, none of a declared type" (untyped backend) from "zero relationships at all" (nothing extracted), plus one aggregated issue naming declared relation types with zero instances, the reporting-only counterpart to The starter relation vocabulary #347's hard self-validation gate and the only detector for a derived type that never materializes. Surfaced in ReconciliationSummary beside Extraction-time constraints vs post-hoc validation #348's fallback counter, not as a log line, so the eval path sees it. Per-type half's consumer is The derivation contract for LlmRecommendationStrategy #353.
Wire a constrained ontology and read the edges — The model stands; its vocabulary does not. 10 real evidence sessions, one per question type, run twice through a typed JointSchema (assets on prototype/gliner2-constrained-edges). Extraction-time constraints vs post-hoc validation #348's claim holds: 119 constrained-only claims (85 as first reported; corrected after a gliner2 schema-cache bug contaminated 3 of 10 permissive runs — see Wire a constrained ontology and read the edges #350's correction) the permissive run never emitted, and they are real — the beam reallocates rather than filters. Its cost is not junk but an answer-bearing edge: visited(User -> Organization:'Museum of Modern Art') is suppressed at decode time because the declared range was (Location,) — one label too narrow, in the regime The starter relation vocabulary #347 made automatic and unreviewed (now How wide should a relation's declared range be #359). A third outcome besides kept/suppressed: 21 claims survive both arms retyped. feasible=False is unreachable under this spec shape — 0/109 windows in both arms; domain/range and cardinality are prohibitive, so the empty solution always validates. Only symmetric/inverse completeness constraints reach it (6/12 windows, each returning empty), which revises Extraction-time constraints vs post-hoc validation #348's infeasible-window retry to a path expected to stay cold; cardinality meanwhile drops 68% of relations silently while still reporting feasible. Also: a symmetric relation is rejected at build unless start_labels == end_labels. Reading the output found three vocabulary problems, each now a ticket: 14% of User endpoints are not the user (visited(User:'kahlo' -> Location:'Paris')), so The starter relation vocabulary #347's unconditional collapse would write false memories (Which User mentions are the user #358); 14% of pairs carry several relation labels — practices/prefers/studied on one pair (Near-synonymous relation labels all fire on the same pair #360); and 3 of 6 question types have no expressible answer because the answer is a value, not a link (Values, not links: do entity attributes belong in the model #361, graduated from the fog). Identity: 2253 mentions → 2067 span-scoped → 903 merged nodes, hubs at User:'user' (364) and User:'i' (132) — the same entity, unjoinable by Entity identity across chunks #346's rule. The GLiNER2-vs-LightRAG comparison Entity identity across chunks #346 handed here was not run: the eval graph is empty and the LightRAG arm needs a fresh LLM-backed reconciliation.
What gliner2 2.0.0 actually offers for entity attributes — No attribute slot exists on the joint path at all, and JointSchema.from_dict silently ignores unknown keys, so an attributes: block would load and do nothing. The mechanism is the legacy API — entity_attributes (closed label sets, no offsets), structure().field() (real value spans, but document-scoped, attached to nothing), classification (doc-level) — and all three compose with entities and relations in one call, but that call is the legacy runtime What JointSchema actually enforces #345 retired, so the trade is values in one pass, or typed relations, not both; keeping JointSchema costs a second call per window (+90%) with no entity ids, resurrecting the span-matching apparatus What JointSchema actually enforces #345 deleted. Record mode is the only per-entity construct and its anchors do match joint entity ids by (start,end), so coupling is recoverable. Nothing is enforced during decoding (no search, hence no feasible analogue), validators= is a provable post-hoc filter, a single-label group can never abstain (argmax), and choices=hallucinates a value absent from the text at conf 0.99. Cost is linear, not quadratic, but a document-scoped structure conflicts across windows — same text gave two purchase dates, and personal_best flipped by window size alone (-> Chunking and the durability of constraints #352). No version bump exists or is needed: 2.0.0 is latest, "2.5" is the checkpoint generation we already run, transformers floor untouched. Beyond the brief: Wire a constrained ontology and read the edges #350's "no expressible answer" was its vocabulary, not the API — value-shaped entity types (Duration, Quantity, TimeWindow, Money) with relations into them express every value fact on the joint path, ids and constraints intact, though 2 of 4 values bound wrong in a sample of one. Record mode's occurrence_policy sweep is unrun and folded into Values, not links: do entity attributes belong in the model #361.
Which User mentions are the user — Collapse iff first-person surface AND user-role turn, plus 'you' in an assistant turn, which inverts. Mapping Wire a constrained ontology and read the edges #350's document-coordinate spans back to their turns showed the two halves catch disjoint errors — surface catches the 95 endpoints that are not first person at all, role catches the 10 that are first person in the assistant's mouth — so The starter relation vocabulary #347's 14% is really 17.5% and the conjunction is the only gate scoring zero in both directions. It executes at write time, as identity: a caller-supplied mention resolver runs inside GLiNER2Backend before _entity_id, so Entity identity across chunks #346's MERGE performs the collapse and the extracted User mention never becomes a node (the ticket's question 4, answered) — not optional, since a merged User:'i' node holds both roles' mentions and can no longer be gated. Not on the ExtractionBackend Protocol, per Does LightRAG consume domain/range too #349. Failures split three ways: third parties re-type to Person (measured to merge into a node the same run already built, frida x103), while the literal 'assistant' and first-person-in-an-assistant-turn are dropped under the rule a first-person mention in an assistant turn is not the user asserting anything — which survives the ventriloquism case ("Did I wear this shirt frequently?", the assistant scripting the user's own voice) that "the assistant is talking about itself" does not. Re-typing is made safe by a new hard rule in hygm.validation — User in a relation's endpoint labels requires Person there too — keeping Extraction-time constraints vs post-hoc validation #348's invariant intact and constraining The derivation contract for LlmRecommendationStrategy #353. The role prefix stays: an ablation over the same 10 sessions found removing it gives 47% more raw edges but 8% fewer distinct claims, nearly doubles the User type's noise (17.5% -> 32%), and puts third parties inside user turns for the first time (0 -> 13), the case only the surface half catches. Its measured job is anchoring the User type, not context-graph: sessions are distilled with a document chunker, and the memory tier ends up 14:1 assistant output #328's role attribution. Resolves The starter relation vocabulary #347's User:'user'(364)/User:'i'(132) fragmentation — both resolve to the same user_id at write time, the join Entity identity across chunks #346's rule could never make.
How wide should a relation's declared range be — v1 is observed then pruned; the loop sees what it suppresses through a shadow sample. Of 266 edges the constrained ranges suppress (clean Wire a constrained ontology and read the edges #350 arm), 29 are real, 43 label confusion and 194 junk, so ~6.7 junk kept out per real loss. 19 of the 29 sit on Organization, and the model's typing there is consistent, not noisy (brooklyn museum is Organization ×24, never Location): the range was wrong for this model, not the model wrong. Neither mechanical rule separates real from junk: junk groups are the larger frequency share (knows -> Topic 31% vs visited -> Organization 11%) and type ambiguity passes 17% real vs 20% junk, so the call is semantic. The asymmetry that decides it under the propose → ingest → revise loop: over-wide junk is visible in the graph, but over-narrow suppression hides its own evidence. So: Extraction-time constraints vs post-hoc validation #348's same-spec invariant stands; a periodic permissive shadow sample reports suppressed endpoint pairs as widening candidates, reporting-only, the input the revise step consumes and the partial-loss detector Does LightRAG consume domain/range too #349's zero-instance signal can't be; and v1's width comes from one permissive pass over the derivation sample that The derivation contract for LlmRecommendationStrategy #353's proposer prunes semantically. Item 3 dissolves: ranges follow the model's typing rather than re-cutting entity types.
Near-synonymous relation labels all fire on the same pair — The graph carries every label. 33 of 244 pairs are multi-label, but 21 are owns + purchased, co-true rather than synonyms, and black jeans carries three labels, all correct; confidence picks wrong (false prefers 0.94, true 0.75). So no write-time arbitration, no decode-time exclusivity (buildable: JointSchema.constraint() takes any Constraint, and UniqueRelationPair is per-label, constraints.py:139), no disjointness rule, no cardinality caps. The real defect is one label's precision: co-fired prefers is right 6/17, firing on mention rather than preference. That's accepted as a known limit with a reporting-only co-firing signal (prefers 63%) for the loop, beside How wide should a relation's declared range be #359's widening candidates. Beyond the brief: the joint compiler drops relation descriptions (compiler.py:27, relations compile to {head:"", tail:""}; only entity descriptions pass). A sharper prefers description gave a byte-identical edge list, so every relation description in Wire a constrained ontology and read the edges #350/Values, not links: do entity attributes belong in the model #361 was inert. The name is the only lever, and it trades recall for precision (likes fires 2.4×; says_they_like halves true and false and heads on 'assistant') while shifting the other relations. For The derivation contract for LlmRecommendationStrategy #353: relation names are the steering surface, descriptions are documentation.
The derivation contract for LlmRecommendationStrategy — A fixed core plus a derived domain, through one propose → observe → prune pass. Every ontology carries User/Person (load-bearing for Which User mentions are the user #358) and the universal value types Duration/Quantity/Money/Date/TimeWindow (span). Derivation proposes the domain types, every relation, and optional corpus-specific value types: the first non-corpus-derived part of the model, a deliberate exception to The starter relation vocabulary #347. Cold start defers extraction, not the session: episodic summaries from day one, semantic extraction from the first derivation, backlog included (text retained, GLiNER2 local), so reconciliation splits. The pipeline is 2 LLM calls plus 1 local extraction. Propose is told names are all the extractor sees (Near-synonymous relation labels all fire on the same pair #360); one permissive GLiNER2 observe pass; prune narrows ranges (How wide should a relation's declared range be #359), sets identity from observation stats (not casing: name-likeness tracks the type roughly, Person 75% vs Activity 9%, but belongs to the mention, and coding corpora are lowercase-but-name-like), and drops zero-instance types/relations (beam tax). It may not rename, since a rename shifts every relation. Two time-stratified samples: whole sessions under a prompt budget for propose, a far larger one (the whole backlog at cold start) for observe. Stability is behavioural: both vocabularies over one sample, types aligned by span co-assignment, reporting-only; zero-instance is a separate axis. The LLM is caller-supplied, with provenance recorded (model, sample ids, observation tables, stability report). hygm defines an Observer protocol that unstructured2graph implements with GLiNER2, which avoids a dependency cycle; create_model gains observer. The re-derivation trigger goes to the evolution loop's map, with its inputs named. The open risk, that nobody has run it, is Run the derivation contract once and read the vocabulary #366.
Chunking and the durability of constraints — One turn per window. Most of the ticket was settled downstream: domain/range is a per-edge predicate, so the union of per-window-valid edges stays valid (0/1004), and nothing needs re-applying after the merge (Extraction-time constraints vs post-hoc validation #348's Cypher remains the backstop); identity across windows is Entity identity across chunks #346/Which User mentions are the user #358's; What gliner2 2.0.0 actually offers for entity attributes #362's value arbitration dissolved with Values, not links: do entity attributes belong in the model #361. The sweep (value vocabulary, 10 sessions) found what nobody had anticipated: under today's 384-word windows 45% of edges cross turns, and 35% of all edges put an assistant-turn tail on a user-turn head. Those pass Which User mentions are the user #358's gate, which checks only the head, onto (:User); 96 are about things the user never wrote (purchased(user -> tops/bottoms/shoes) from the assistant's closet list). One turn per window (split only turns over 384 words) takes that to 0%, yields more distinct claims (393 vs 342), costs no answer, and gains How wide should a relation's declared range be #359's MoMA visited edge. The session stays one Chunk, and the windows reuse Which User mentions are the user #358's turn spans. Windows never exceed ~384 words: past the encoder's 512 positions, mentions and distinct claims halve. Cross-turn coreference is a known limit. Tables are too, and they raise n-ary facts: Admon's miss wasn't a window boundary but a table. Generic linearization binds nothing; a verb template binds but puts Admon on the wrong shift and drops the day. (Admon, Sunday, 8 am – 4 pm) is three-way.
The read-time contract for non-conformant edges — An integrity alarm, not a read filter. Nothing reads ontology_conformant today, and after this map nothing writes it on a GLiNER2 graph either: entity types are the ontology's list (0 by construction), relation domain/range survives the merge (0/1004), the infeasible fallback is unreachable (0/109). So a flag means a bug or vintage, and readers never filter: the count goes into ReconciliationSummary and an e2e test asserts zero for GLiNER2. Extraction-time constraints vs post-hoc validation #348's infeasible retry and its property leave the plan, because derivation never emits symmetric/inverse (they empty windows); feasible=False becomes a bug. Self-loops are dropped at write, counted, when resolved head and tail ids match, which is what allow_self can't see after our identity runs. Wire a constrained ontology and read the edges #350's 12 self-loops were text-only: 11 span two types, so the real count is 1. The asymmetry stays (LightRAG entities are already structural, its relations vacuous). Vintage flags go to the evolution-loop map.
Resolving relative dates against valid_at — No resolver in v1: an event's date is its edge's valid_at, made computable.Values, not links: do entity attributes belong in the model #361's premise failed measurement: LongMemEval phrases every evidence event as "today"/"just" in its session, so the gold answer equals the difference of the session dates in 8/8 two-session "between" questions, and valid_at alone serves 17 of 18 date questions. The one miss is a "yesterday", off by a day. That is a benchmark artifact, so revisit with the coding corpus. Date keeps raw text and stays span-scoped. What ships: valid_at as a Memgraph temporal value (today every timestamp is a string, started_at the corpus's raw '2023/05/30 (Tue) 17:27'), making date arithmetic a deterministic duration.between, not LLM arithmetic; and valid_at from the source turn's Action.timestamp, not the session's (free after Chunking and the durability of constraints #352, right for resumed sessions; revises The starter relation vocabulary #347), which needs inject.py to stop leaving turns at ingest time. Found on the way: the retrieval agent is never given question_date, so the 8 "ago" questions are unanswerable in the eval whatever the graph holds (The retrieval agent is never given question_date #367, map Map: Context-graph emergence pipeline — reconciliation, retrieval, eval #297).
Run the derivation contract once and read the vocabulary — The contract works once its input is clean, but one ~10-session propose sample decides which domains exist. Two derivations from disjoint samples, read on the 10 evidence sessions: B roughly matches the hand vocabulary (842 raw / 313 distinct vs 895 / 393, MoMA visited at 0.94-0.97 vs 0.66-0.78); A, whose sample had no travel, proposed no Location and yields zero answer-bearing edges. Prune can't add, so a missing domain is unrecoverable (How wide should a relation's declared range be #359's asymmetry). Prune itself is good: junk relations dropped, identity calls right (Organization/Location/CreativeWork global, generic nouns chunk). Revises The derivation contract for LlmRecommendationStrategy #353: batched propose with one consolidate call merging synonyms (JobRole/Occupation share 31/32 spans); the proposer must propose relations into core value types (neither did, so 25:50 was lost); no tense variants (plans_to_visit co-fires with visited_location 21/30, contradicting it; exception to Near-synonymous relation labels all fire on the same pair #360); stability reports coverage as well as agreement (84% agreement hid 876 spans only B typed). ~$0.55 of LLM per derivation; observe (~1 h / 60 sessions) dominates. Found on the way: unstructured2graph: gliner2 cuts relation candidates alphabetically by type on ties #371: gliner2 cuts relation candidates with ties broken alphabetically by type, so a permissive pass past 11 types never heads an edge on User (the first run's prune kept 1 relation of 12), and constrained extraction loses ~4% of User edges; observe and production set both caps to 4096. The revised contract's run is Run the revised derivation contract: batched propose and consolidate #372.
Observe value relations with their intended tail — Observe fixed, typing not. With their tails constrained, value relations fire into values, and prune keeps 7 in S1 and 9 in S2. Sharded observe drops from 3.7 h to 8–9 min. But 25:50 is still lost: in larger derived vocabularies GLiNER2 doesn't type it Duration (S1V misses it, S2V types it Food). Open, moved out of this map with the rest of derivation.
Outcome
Destination reached. The plan (cuddly-bubbling-dijkstra.md) was rewritten with every decision above. It is implemented on a hand-written vocabulary (ManualStrategy) in PR #373, which closes this map:
hygm (types, validation, Manual and OWL strategies);
per-type identity, where lowercase generic nouns stay per chunk;
one turn per window;
valid_at from the source turn;
session reconciliation from segmented documents.
A port check on #350's evidence sessions matched the prototype: 2,525 mentions against 2,509, all three answer edges on the user's node, and 0 non-conformant.
Still unspecified: n-ary facts, per-tenant identity, and a companion ADR. ADR 0004's amendment covers the two-compilation model, so no companion ADR is needed. N-ary facts and per-tenant identity wait for the second corpus and for multi-tenancy.
Whether the ADR 0004 amendment needs a companion ADR for the two-compilation model itself, or whether the amended 0004 plus the plan carry it. Decide when the amendment is actually drafted — the shape of what 0004 ends up saying settles it.
Whether global entity identity should be scoped per user or tenant. Entity identity across chunks #346 made identity global across sessions; merging "Admon" across two different users' histories is a different proposition from merging within one, and nothing in the plan says the graph is single-tenant. Cannot be phrased sharply until multi-tenancy is settled, which sits well outside this map.
The revised plan's implementation sequencing — downstream of everything else here.
The versioned schema-evolution loop (schema vN → data vN+1 → schema vN+1, driven by usage shape). Wanted — and The starter relation vocabulary #347's fully-automatic derivation leans on it as its error-correction path — but it needs ontology versioning, a re-extraction/migration story, and a mixed-vintage-data policy, none of which exist today. Beyond this map's destination; needs its own map.
Typing LightRAG's edges. Ruled out while resolving Does LightRAG consume domain/range too #349: LightRAG emits untyped, explicitly-undirected :DIRECTED edges and exposes no relation-type slot or prompt hook, so making it carry the typed relation model means changing what LightRAG extracts — a separate effort against a backend this map has designated a retained, non-interchangeable alternative. Not deferred, not pursued.
Destination reached. The revised plan is implemented and merged in #373 (b3ab503), on a hand-written vocabulary:
hygm: types, validation, Manual and OWL strategies;
the GLiNER2 joint backend: held schema, 4096 candidate caps, one turn per window, the Which User mentions are the user #358 resolver extended to pronoun Person mentions, per-type identity with lowercase generic nouns kept per chunk, valid_at from the source turn, and self-loops dropped;
post-hoc domain/range checking and ontology_report;
post-processing scoped to the ingested chunks;
sessions-graph reconciliation from segmented documents.
A port check on #350's evidence sessions matched the prototype: 2,525 mentions against 2,509, all three answer edges on the user's node, and 0 non-conformant.
Stays open, tracked elsewhere (see Outcome above):
Destination
The
hygmshared-ontology plan, revised until every open decision is settled and it can be implemented without further design work.The plan under revision lives at
~/.claude/plans/cuddly-bubbling-dijkstra.md(local to @antejavor). Its structural core —RelationTypegainingstart_labels/end_labels, three production strategies (manual YAML / OWL import / LLM recommendation),unstructured2graph/ontology.pyas a thin adapter — is not up for re-decision here. This map resolves what that plan left open or got wrong.Why this map exists
A 100-question GLiNER2 eval run produced a graph with zero semantic edges. Diagnosis found three nested causes:
context_graph_eval.reconcilebuildsGLiNER2Backend()with no ontology, so it falls back toDEFAULT_ONTOLOGY— 11 entity types, zero relation types._extract_sync'sif self._relation_schema:guard never fires, so relation extraction never runs at all.GLiNER2Backendcallsgliner2's coarseSchema.relations(dict)builder, whose source always writes{"head": "", "tail": ""}— no type constraint, no cardinality. That is the pre-2.5 mode fastino's own write-up contrasts against: "relation extraction returned independently thresholded triples and consistency was the user's responsibility."gliner2==2.0.0, unused:gliner2.joint_ie.schema.JointSchema, with.relation(name, head=(...), tail=(...), max_per_head=, max_per_tail=, symmetric=, inverse=, allow_self=)and.no_self_loops().unstructured2graph.ontology.RelationTypeis(label, description)— there is nowhere to put a head/tail constraint even if we wanted to. That is the same gap blocking thehygm/OWL unification, which is why this is one effort and not two.Convergent evidence for the shape: the plan chose
start_labels/end_labelsbecause OWL'srdfs:domain/rdfs:rangemap onto them; sql2graph's HyGM independently arrived atstart_node_labels/end_node_labels; GLiNER2'sRelationSpec(head=, tail=)is the same concept a third time.Notes
unstructured2graph, the plannedhygmpackage, and thecontext-graphfamily./grillingand/domain-modelingeach session. Read the plan file in full before working any ticket.temperature=0. The eval is a confirmation step after implementation, never the gate on a modeling decision.LightRAGBackendis a retained, non-interchangeable alternative pipeline — interchangeable in process, never in results — kept rather than deleted, but not invested in. Established while resolving Does LightRAG consume domain/range too #349; it is why the typed relation model is GLiNER2-only by design rather than by omission.Decisions so far
What JointSchema actually enforces — Constraints ARE enforced during decoding (
optimizers/beam.py:67, inside beam expansion), and typing even earlier at candidate construction — but per window only:long_text.pydrops any relation whose endpoints came from different chunks and never re-checks constraints at the merge. Entry point isJointIE/JointIEEngine, notAutoExtractor.extract();compile_schema().build()silently drops all typing. Output isJointResultwith entity-id endpoints — replaces the whole span-matching apparatus. "Unconstrained" is not expressible: emptyhead/tailraises, so the plan'sstart_labels=()compat rule must be translated to "all declared types" in our adapter.The starter relation vocabulary — There is no hand-authored vocabulary. Measured the corpus: 42% of questions turn on time/change, 98/100 are first-person. So relations are time-anchored (
valid_atfrom the session, stamped by our writer, not extracted); superseded facts are retained, timestamps only (a real question asks for the initial value, so "newest wins" would delete the answer); and the ontology is derived per corpus byLlmRecommendationStrategy, fully automatic with no human approval, gated byhygm.validationself-validation (hard) plus a stability check (reports, never blocks).Userbecomes an extracted entity type (verified: GLiNER2 resolves "I" to it), and processing — never collection — collapses allUsermentions onto the session's(:User)node, synthesisinganon:<session_id>when collection recorded none.Extraction-time constraints vs post-hoc validation — Both: one specification, two compilations.
start_labels/end_labelsis authored once and compiled toJointSchematyping at extraction time (job: quality, steer the decoder; GLiNER2-only, best-effort) and to a Cypher check over Memgraph post-hoc (job: integrity; every backend, authoritative). Non-redundant for two verified reasons:allow_edgeruns inside beam expansion so constrained decoding reallocates rather than filters, and per What JointSchema actually enforces #345 constraints never survive the window merge — so post-hoc is the only thing that sees the merged graph. Same spec in both, no deliberate loosening, which buys the invariant that a GLiNER2 post-hoc violation means a cross-window merge or a bug. This revises the plan's "self._relation_schemaconstruction is unaffected". ADR 0004 is amended to cover relations explicitly, recording that decode-time suppression forfeits re-projection — the cost The starter relation vocabulary #347's automatic derivation makes likelier. Mitigation:long_text.py:77dropsfeasible, so we drive windows ourselves, and an infeasible window is retried with permissive endpoints and written flagged; per-edge suppression stays unrecorded (private API, unattributable) and moves to Wire a constrained ontology and read the edges #350 as an eval-time diff.enforce_ontologygates the post-hoc half only, meaning unchanged; the fallback gets its own property soontology_conformantkeeps one writer.Entity identity across chunks — Exact normalized text within type, per-type opt-in, merged at write time.
_entity_iddropschunk_hashfor opted-in types and the existingMERGEdoes the work — no second pass. Rejected embeddings (non-deterministic across model versions, breaks eval reproducibility) and uniform merging (12,583Conceptvs 3,809Person: merging every identically-wordedConceptbuilds hubs that discriminate nothing). Thefile_pathblocker in_entity_id's docstring is shallower than documented —ON CREATE SETmeans a re-merged node keeps the first chunk'sfile_path, so the fix is writingMENTIONED_INat ingest, which also retires a full-graph cartesian product for this backend. Because every mention keeps its edge, merging destroys nothing and ADR 0004's instinct holds rather than collides — unlike Extraction-time constraints vs post-hoc validation #348's decode-time suppression. No confidence flag or cap; a reporting-only highest-degree signal catches supernodes. Identity stays per-backend (LightRAG keeps LLM consolidation), with only the per-type mergeability flag shared. Prerequisite filed as unstructured2graph: entity_id is MERGEd on but never indexed or constrained #357:entity_idis MERGEd on but never indexed or constrained.Does LightRAG consume domain/range too — No, at neither end — by design, not omission. LightRAG's extraction record has no relation type field (
relation|source|target|keywords|description) and its prompt orders relationships treated as undirected, so domain/range is not expressible; storage writes one untypedMERGE (a)-[r:DIRECTED]-(b). Extraction-time steering is therefore impossible, and post-hoc validation — which keys on the type label — is vacuous, not free, revising Extraction-time constraints vs post-hoc validation #348's follow-up note and the plan's §2 claim (the action, no code change, stands; the claim of coverage does not). The backends are interchangeable in process, never in results, so the asymmetry is stated rather than reconciled: no capability flag, no Protocol change. What ships instead is an observational coverage signal inenforce_relation_domain_rangereading the graph rather than the backend — discriminating "N relationships, none of a declared type" (untyped backend) from "zero relationships at all" (nothing extracted), plus one aggregated issue naming declared relation types with zero instances, the reporting-only counterpart to The starter relation vocabulary #347's hard self-validation gate and the only detector for a derived type that never materializes. Surfaced inReconciliationSummarybeside Extraction-time constraints vs post-hoc validation #348's fallback counter, not as a log line, so the eval path sees it. Per-type half's consumer is The derivation contract for LlmRecommendationStrategy #353.Wire a constrained ontology and read the edges — The model stands; its vocabulary does not. 10 real evidence sessions, one per question type, run twice through a typed
JointSchema(assets onprototype/gliner2-constrained-edges). Extraction-time constraints vs post-hoc validation #348's claim holds: 119 constrained-only claims (85 as first reported; corrected after a gliner2 schema-cache bug contaminated 3 of 10 permissive runs — see Wire a constrained ontology and read the edges #350's correction) the permissive run never emitted, and they are real — the beam reallocates rather than filters. Its cost is not junk but an answer-bearing edge:visited(User -> Organization:'Museum of Modern Art')is suppressed at decode time because the declared range was(Location,)— one label too narrow, in the regime The starter relation vocabulary #347 made automatic and unreviewed (now How wide should a relation's declared range be #359). A third outcome besides kept/suppressed: 21 claims survive both arms retyped.feasible=Falseis unreachable under this spec shape — 0/109 windows in both arms; domain/range and cardinality are prohibitive, so the empty solution always validates. Onlysymmetric/inversecompleteness constraints reach it (6/12 windows, each returning empty), which revises Extraction-time constraints vs post-hoc validation #348's infeasible-window retry to a path expected to stay cold; cardinality meanwhile drops 68% of relations silently while still reporting feasible. Also: asymmetricrelation is rejected at build unlessstart_labels == end_labels. Reading the output found three vocabulary problems, each now a ticket: 14% ofUserendpoints are not the user (visited(User:'kahlo' -> Location:'Paris')), so The starter relation vocabulary #347's unconditional collapse would write false memories (WhichUsermentions are the user #358); 14% of pairs carry several relation labels —practices/prefers/studiedon one pair (Near-synonymous relation labels all fire on the same pair #360); and 3 of 6 question types have no expressible answer because the answer is a value, not a link (Values, not links: do entity attributes belong in the model #361, graduated from the fog). Identity: 2253 mentions → 2067 span-scoped → 903 merged nodes, hubs atUser:'user'(364) andUser:'i'(132) — the same entity, unjoinable by Entity identity across chunks #346's rule. The GLiNER2-vs-LightRAG comparison Entity identity across chunks #346 handed here was not run: the eval graph is empty and the LightRAG arm needs a fresh LLM-backed reconciliation.What gliner2 2.0.0 actually offers for entity attributes — No attribute slot exists on the joint path at all, and
JointSchema.from_dictsilently ignores unknown keys, so anattributes:block would load and do nothing. The mechanism is the legacy API —entity_attributes(closed label sets, no offsets),structure().field()(real value spans, but document-scoped, attached to nothing),classification(doc-level) — and all three compose with entities and relations in one call, but that call is the legacy runtime What JointSchema actually enforces #345 retired, so the trade is values in one pass, or typed relations, not both; keepingJointSchemacosts a second call per window (+90%) with no entity ids, resurrecting the span-matching apparatus What JointSchema actually enforces #345 deleted. Record mode is the only per-entity construct and its anchors do match joint entity ids by(start,end), so coupling is recoverable. Nothing is enforced during decoding (no search, hence nofeasibleanalogue),validators=is a provable post-hoc filter, a single-label group can never abstain (argmax), andchoices=hallucinates a value absent from the text at conf 0.99. Cost is linear, not quadratic, but a document-scoped structure conflicts across windows — same text gave two purchase dates, andpersonal_bestflipped by window size alone (-> Chunking and the durability of constraints #352). No version bump exists or is needed: 2.0.0 is latest, "2.5" is the checkpoint generation we already run, transformers floor untouched. Beyond the brief: Wire a constrained ontology and read the edges #350's "no expressible answer" was its vocabulary, not the API — value-shaped entity types (Duration,Quantity,TimeWindow,Money) with relations into them express every value fact on the joint path, ids and constraints intact, though 2 of 4 values bound wrong in a sample of one. Record mode'soccurrence_policysweep is unrun and folded into Values, not links: do entity attributes belong in the model #361.Which
Usermentions are the user — Collapse iff first-person surface AND user-role turn, plus'you'in an assistant turn, which inverts. Mapping Wire a constrained ontology and read the edges #350's document-coordinate spans back to their turns showed the two halves catch disjoint errors — surface catches the 95 endpoints that are not first person at all, role catches the 10 that are first person in the assistant's mouth — so The starter relation vocabulary #347's 14% is really 17.5% and the conjunction is the only gate scoring zero in both directions. It executes at write time, as identity: a caller-supplied mention resolver runs insideGLiNER2Backendbefore_entity_id, so Entity identity across chunks #346'sMERGEperforms the collapse and the extractedUsermention never becomes a node (the ticket's question 4, answered) — not optional, since a mergedUser:'i'node holds both roles' mentions and can no longer be gated. Not on theExtractionBackendProtocol, per Does LightRAG consume domain/range too #349. Failures split three ways: third parties re-type toPerson(measured to merge into a node the same run already built,fridax103), while the literal'assistant'and first-person-in-an-assistant-turn are dropped under the rule a first-person mention in an assistant turn is not the user asserting anything — which survives the ventriloquism case ("Did I wear this shirt frequently?", the assistant scripting the user's own voice) that "the assistant is talking about itself" does not. Re-typing is made safe by a new hard rule inhygm.validation—Userin a relation's endpoint labels requiresPersonthere too — keeping Extraction-time constraints vs post-hoc validation #348's invariant intact and constraining The derivation contract for LlmRecommendationStrategy #353. The role prefix stays: an ablation over the same 10 sessions found removing it gives 47% more raw edges but 8% fewer distinct claims, nearly doubles theUsertype's noise (17.5% -> 32%), and puts third parties inside user turns for the first time (0 -> 13), the case only the surface half catches. Its measured job is anchoring theUsertype, not context-graph: sessions are distilled with a document chunker, and the memory tier ends up 14:1 assistant output #328's role attribution. Resolves The starter relation vocabulary #347'sUser:'user'(364)/User:'i'(132) fragmentation — both resolve to the sameuser_idat write time, the join Entity identity across chunks #346's rule could never make.Values, not links: do entity attributes belong in the model — Values are entity types, not attributes. Measured over all 100 goldens: 61% of answerable questions want a value, above The starter relation vocabulary #347's own 42% — but only 48% need a value slot, since 12 are counts of edges (relations + Entity identity across chunks #346's identity, nothing new), and 18 are date questions (dates in session text are 17 relative to 6 absolute). (Corrected by Resolving relative dates against valid_at #364: the claim that 10 "between X and Y" questions need an event date
valid_atis not was wrong;valid_atserves 17 of the 18.)Duration/Quantity/Money/Date/TimeWindowride the joint path (ids, decode-time constraints,feasibleand Extraction-time constraints vs post-hoc validation #348's two compilations unchanged), which retires What gliner2 2.0.0 actually offers for entity attributes #362'soccurrence_policyunknown rather than resolving it; the legacy structure path costs typed relations and record mode costs +90% on an unswept binding. Value types get span-scoped identity — under today's(chunk_hash, type, text)every3in a session is one node, and 35% of value keys measured cover several spans — sohygm.typesgains nothing new: Entity identity across chunks #346'smergeable: boolwidens toidentity: global | chunk | span, which The derivation contract for LlmRecommendationStrategy #353 now emits. Reified assertions are subsumed: a value fact is an ordinary edge inheritingvalid_atandMENTIONED_IN. Validated on the 10 evidence sessions:personal_best(User -> Duration:'25:50')lands where the baseline could not express it; Admon's8 am - 4 pmis extracted but never bound (a table — Chunking and the durability of constraints #352); no MoMA event date comes out. Two costs carried to The derivation contract for LlmRecommendationStrategy #353: adding types taxes the existing relations (−8% distinct claims) and roughly half the value edges are junk (Money:'premium features'). Found on the way: gliner2 caches compiled schemas under a memory address (unstructured2graph: gliner2's compiled-schema cache is keyed on a memory address #365), which had contaminated 3 of Wire a constrained ontology and read the edges #350's 10 permissive runs — corrected there.How wide should a relation's declared range be — v1 is observed then pruned; the loop sees what it suppresses through a shadow sample. Of 266 edges the constrained ranges suppress (clean Wire a constrained ontology and read the edges #350 arm), 29 are real, 43 label confusion and 194 junk, so ~6.7 junk kept out per real loss. 19 of the 29 sit on
Organization, and the model's typing there is consistent, not noisy (brooklyn museumisOrganization×24, neverLocation): the range was wrong for this model, not the model wrong. Neither mechanical rule separates real from junk: junk groups are the larger frequency share (knows -> Topic31% vsvisited -> Organization11%) and type ambiguity passes 17% real vs 20% junk, so the call is semantic. The asymmetry that decides it under the propose → ingest → revise loop: over-wide junk is visible in the graph, but over-narrow suppression hides its own evidence. So: Extraction-time constraints vs post-hoc validation #348's same-spec invariant stands; a periodic permissive shadow sample reports suppressed endpoint pairs as widening candidates, reporting-only, the input the revise step consumes and the partial-loss detector Does LightRAG consume domain/range too #349's zero-instance signal can't be; and v1's width comes from one permissive pass over the derivation sample that The derivation contract for LlmRecommendationStrategy #353's proposer prunes semantically. Item 3 dissolves: ranges follow the model's typing rather than re-cutting entity types.Near-synonymous relation labels all fire on the same pair — The graph carries every label. 33 of 244 pairs are multi-label, but 21 are
owns+purchased, co-true rather than synonyms, andblack jeanscarries three labels, all correct; confidence picks wrong (falseprefers0.94, true 0.75). So no write-time arbitration, no decode-time exclusivity (buildable:JointSchema.constraint()takes anyConstraint, andUniqueRelationPairis per-label,constraints.py:139), no disjointness rule, no cardinality caps. The real defect is one label's precision: co-firedprefersis right 6/17, firing on mention rather than preference. That's accepted as a known limit with a reporting-only co-firing signal (prefers63%) for the loop, beside How wide should a relation's declared range be #359's widening candidates. Beyond the brief: the joint compiler drops relation descriptions (compiler.py:27, relations compile to{head:"", tail:""}; only entity descriptions pass). A sharperprefersdescription gave a byte-identical edge list, so every relation description in Wire a constrained ontology and read the edges #350/Values, not links: do entity attributes belong in the model #361 was inert. The name is the only lever, and it trades recall for precision (likesfires 2.4×;says_they_likehalves true and false and heads on'assistant') while shifting the other relations. For The derivation contract for LlmRecommendationStrategy #353: relation names are the steering surface, descriptions are documentation.The derivation contract for LlmRecommendationStrategy — A fixed core plus a derived domain, through one propose → observe → prune pass. Every ontology carries
User/Person(load-bearing for WhichUsermentions are the user #358) and the universal value typesDuration/Quantity/Money/Date/TimeWindow(span). Derivation proposes the domain types, every relation, and optional corpus-specific value types: the first non-corpus-derived part of the model, a deliberate exception to The starter relation vocabulary #347. Cold start defers extraction, not the session: episodic summaries from day one, semantic extraction from the first derivation, backlog included (text retained, GLiNER2 local), so reconciliation splits. The pipeline is 2 LLM calls plus 1 local extraction. Propose is told names are all the extractor sees (Near-synonymous relation labels all fire on the same pair #360); one permissive GLiNER2 observe pass; prune narrows ranges (How wide should a relation's declared range be #359), setsidentityfrom observation stats (not casing: name-likeness tracks the type roughly,Person75% vsActivity9%, but belongs to the mention, and coding corpora are lowercase-but-name-like), and drops zero-instance types/relations (beam tax). It may not rename, since a rename shifts every relation. Two time-stratified samples: whole sessions under a prompt budget for propose, a far larger one (the whole backlog at cold start) for observe. Stability is behavioural: both vocabularies over one sample, types aligned by span co-assignment, reporting-only; zero-instance is a separate axis. The LLM is caller-supplied, with provenance recorded (model, sample ids, observation tables, stability report).hygmdefines anObserverprotocol thatunstructured2graphimplements with GLiNER2, which avoids a dependency cycle;create_modelgainsobserver. The re-derivation trigger goes to the evolution loop's map, with its inputs named. The open risk, that nobody has run it, is Run the derivation contract once and read the vocabulary #366.Chunking and the durability of constraints — One turn per window. Most of the ticket was settled downstream: domain/range is a per-edge predicate, so the union of per-window-valid edges stays valid (0/1004), and nothing needs re-applying after the merge (Extraction-time constraints vs post-hoc validation #348's Cypher remains the backstop); identity across windows is Entity identity across chunks #346/Which
Usermentions are the user #358's; What gliner2 2.0.0 actually offers for entity attributes #362's value arbitration dissolved with Values, not links: do entity attributes belong in the model #361. The sweep (value vocabulary, 10 sessions) found what nobody had anticipated: under today's 384-word windows 45% of edges cross turns, and 35% of all edges put an assistant-turn tail on a user-turn head. Those pass WhichUsermentions are the user #358's gate, which checks only the head, onto(:User); 96 are about things the user never wrote (purchased(user -> tops/bottoms/shoes)from the assistant's closet list). One turn per window (split only turns over 384 words) takes that to 0%, yields more distinct claims (393 vs 342), costs no answer, and gains How wide should a relation's declared range be #359's MoMAvisitededge. The session stays oneChunk, and the windows reuse WhichUsermentions are the user #358's turn spans. Windows never exceed ~384 words: past the encoder's 512 positions, mentions and distinct claims halve. Cross-turn coreference is a known limit. Tables are too, and they raise n-ary facts: Admon's miss wasn't a window boundary but a table. Generic linearization binds nothing; a verb template binds but puts Admon on the wrong shift and drops the day. (Admon, Sunday, 8 am – 4 pm) is three-way.The read-time contract for non-conformant edges — An integrity alarm, not a read filter. Nothing reads
ontology_conformanttoday, and after this map nothing writes it on a GLiNER2 graph either: entity types are the ontology's list (0 by construction), relation domain/range survives the merge (0/1004), the infeasible fallback is unreachable (0/109). So a flag means a bug or vintage, and readers never filter: the count goes intoReconciliationSummaryand an e2e test asserts zero for GLiNER2. Extraction-time constraints vs post-hoc validation #348's infeasible retry and its property leave the plan, because derivation never emitssymmetric/inverse(they empty windows);feasible=Falsebecomes a bug. Self-loops are dropped at write, counted, when resolved head and tail ids match, which is whatallow_selfcan't see after our identity runs. Wire a constrained ontology and read the edges #350's 12 self-loops were text-only: 11 span two types, so the real count is 1. The asymmetry stays (LightRAG entities are already structural, its relations vacuous). Vintage flags go to the evolution-loop map.Resolving relative dates against valid_at — No resolver in v1: an event's date is its edge's
valid_at, made computable. Values, not links: do entity attributes belong in the model #361's premise failed measurement: LongMemEval phrases every evidence event as "today"/"just" in its session, so the gold answer equals the difference of the session dates in 8/8 two-session "between" questions, andvalid_atalone serves 17 of 18 date questions. The one miss is a "yesterday", off by a day. That is a benchmark artifact, so revisit with the coding corpus.Datekeeps raw text and stays span-scoped. What ships:valid_atas a Memgraph temporal value (today every timestamp is a string,started_atthe corpus's raw'2023/05/30 (Tue) 17:27'), making date arithmetic a deterministicduration.between, not LLM arithmetic; andvalid_atfrom the source turn'sAction.timestamp, not the session's (free after Chunking and the durability of constraints #352, right for resumed sessions; revises The starter relation vocabulary #347), which needsinject.pyto stop leaving turns at ingest time. Found on the way: the retrieval agent is never givenquestion_date, so the 8 "ago" questions are unanswerable in the eval whatever the graph holds (The retrieval agent is never given question_date #367, map Map: Context-graph emergence pipeline — reconciliation, retrieval, eval #297).Run the derivation contract once and read the vocabulary — The contract works once its input is clean, but one ~10-session propose sample decides which domains exist. Two derivations from disjoint samples, read on the 10 evidence sessions: B roughly matches the hand vocabulary (842 raw / 313 distinct vs 895 / 393, MoMA
visitedat 0.94-0.97 vs 0.66-0.78); A, whose sample had no travel, proposed noLocationand yields zero answer-bearing edges. Prune can't add, so a missing domain is unrecoverable (How wide should a relation's declared range be #359's asymmetry). Prune itself is good: junk relations dropped, identity calls right (Organization/Location/CreativeWorkglobal, generic nouns chunk). Revises The derivation contract for LlmRecommendationStrategy #353: batched propose with one consolidate call merging synonyms (JobRole/Occupationshare 31/32 spans); the proposer must propose relations into core value types (neither did, so25:50was lost); no tense variants (plans_to_visitco-fires withvisited_location21/30, contradicting it; exception to Near-synonymous relation labels all fire on the same pair #360); stability reports coverage as well as agreement (84% agreement hid 876 spans only B typed). ~$0.55 of LLM per derivation; observe (~1 h / 60 sessions) dominates. Found on the way: unstructured2graph: gliner2 cuts relation candidates alphabetically by type on ties #371: gliner2 cuts relation candidates with ties broken alphabetically by type, so a permissive pass past 11 types never heads an edge onUser(the first run's prune kept 1 relation of 12), and constrained extraction loses ~4% ofUseredges; observe and production set both caps to 4096. The revised contract's run is Run the revised derivation contract: batched propose and consolidate #372.Run the revised derivation contract — Not accepted: places recovered, value facts still lost. Both seeds recover
Placeand MoMAvisited, and neither has a tense pair. But both lose25:50: the value relations proposed per Run the derivation contract once and read the vocabulary #366 never fire into a value under the permissive observe pass (improved_personal_best_by405 edges, 0 into Duration), so prune drops them. S2 also has no education domain (it losesBusiness Administration). Consolidate roughly doubles the relation count (47/51 → 34/38 after prune), and observe cost scales with it (~3.7 h per derivation on one core). Decision: observe value relations with their intended value-type tail (Observe value relations with their intended tail during derivation #386). Implementation of the rest of the map landed in PR context-graph/unstructured2graph: typed relation model on a shared hygm ontology (map #344) #373 on a hand vocabulary.Observe value relations with their intended tail — Observe fixed, typing not. With their tails constrained, value relations fire into values, and prune keeps 7 in S1 and 9 in S2. Sharded observe drops from 3.7 h to 8–9 min. But
25:50is still lost: in larger derived vocabularies GLiNER2 doesn't type itDuration(S1V misses it, S2V types itFood). Open, moved out of this map with the rest of derivation.Outcome
Destination reached. The plan (
cuddly-bubbling-dijkstra.md) was rewritten with every decision above. It is implemented on a hand-written vocabulary (ManualStrategy) in PR #373, which closes this map:hygm(types, validation, Manual and OWL strategies);Usermentions are the user #358 resolver, with pronounPersonmentions routed through it;valid_atfrom the source turn;A port check on #350's evidence sessions matched the prototype: 2,525 mentions against 2,509, all three answer edges on the user's node, and 0 non-conformant.
Handed off:
LlmRecommendationStrategy): Run the revised derivation contract: batched propose and consolidate #372 recovered places but lost value facts, and Observe value relations with their intended tail during derivation #386 moved the loss from observe to entity typing. Observe value relations with their intended tail during derivation #386 continues outside this map, with the evolution-loop work.Not yet specified
Out of scope
model.batch_extract()for the GLiNER2 reconciliation path). Pure throughput mechanics, no bearing on the graph model. Stays on unstructured2graph: GLiNER2Backend has no cross-chunk coreference and no batched reconciliation path (LightRAG has both) #339.agents/sql2graph's own HyGM — already deferred by the plan (§5), has its own independent packaging blocker.— moved INTO scope by The starter relation vocabulary #347. The plan deferred it, but deciding the vocabulary is derived rather than hand-authored makes it the only thing that produces one. Its contract is now The derivation contract for LlmRecommendationStrategy #353.LlmRecommendationStrategy's real implementation:DIRECTEDedges and exposes no relation-type slot or prompt hook, so making it carry the typed relation model means changing what LightRAG extracts — a separate effort against a backend this map has designated a retained, non-interchangeable alternative. Not deferred, not pursued.