Repository navigation
context-graph/unstructured2graph: typed relation model on a shared hygm ontology (map #344) - #373
Merged
Merged
Conversation
…gm ontology (map #344) Implements map #344's decisions with hand-written ontologies (ManualStrategy). LlmRecommendationStrategy stays an interface until its evidence run (#372). - hygm (new package): NodeType with identity (global/chunk/span), RelationType with start_labels/end_labels, validate_model (incl. User requires Person), ManualStrategy (YAML), OwlImportStrategy (rdflib, hygm[owl]). - unstructured2graph: ontology loads through hygm; post-hoc domain/range check flags, never deletes (ADR 0004 amended); ontology_report for integrity and coverage; from_documents ingests verbatim text with turn segments; Chunk.hash and entity_id are unique AND indexed (a Memgraph unique constraint does not index). - GLiNER2Backend rewritten on gliner2's joint path: schema compiled once and held (#365), candidate caps 4096 (#371), one window per turn (#352), the user-mention resolver (#358), per-type identity with MENTIONED_IN per chunk (#346), valid_at as a datetime from the source turn (#364), self-loops dropped (#355). - sessions-graph: reconcile_session sends a segmented Document, synthesizes anon-<session_id> when a session has no user (#347, #354), and reports integrity counts in ReconciliationSummary. - eval: injected turns carry the session date; GLiNER2 extracts against a bundled LongMemEval ontology (#361's hand vocabulary).
antejavor
added a commit
that referenced
this pull request
Sep 29, 2026
… and the #373 port check derive_batched.py runs #353's contract as revised by #366 (4 propose batches, one consolidate call, value-relation and no-tense rules, observe at 4096 caps) from two disjoint seeds; read_derived.py reads them on the evidence sessions. Both seeds recover Place and MoMA visited with no tense pairs, but both lose 25:50: value relations never fire into a value type under the permissive observe pass, so prune drops them. port_check.py runs PR #373's GLiNER2Backend over the same sessions (2525 mentions vs the prototype's 2509, all three answer edges on the user's node, 0 non-conformant).
antejavor
added a commit
that referenced
this pull request
Sep 29, 2026
… typed GLiNER2 graph eval_diagnosis.py traces each of PR #373's 100-question/5-session eval answers through its evidence sessions (stated, extracted as an entity, on an edge, on a user edge) and scores two oracle retrievals with the eval's own answer prompt and judge: every typed edge of the evidence sessions scores 21/100, their full text 48/100 (text search 33, the graph agent 12-15). Of 27 span answers, 24 are extracted as entities but only 11 end up on an edge off the user. Claude-Session: https://claude.ai/code/session_01BczgXHGBrKsHSQUGB8P8te
…ngested connect_chunks_to_entities, promote_entity_types_to_labels, enforce_relation_domain_range and ontology_report ran over the whole workspace after every ingest, so each session rescanned every entity and relationship: measured 9 s/session at 8k entities rising to 10 s at 16k on a 4,600-session eval run, heading for ~400k. They now take the ingested chunk hashes and start from those chunks' entities (MENTIONED_IN), with a workspace file_path index for the chunk lookup. Unscoped calls keep the old whole-workspace behaviour for re-projection after an ontology change. Claude-Session: https://claude.ai/code/session_01BczgXHGBrKsHSQUGB8P8te
…eneric nouns per chunk
Found reading the typed graph built over a 500-session eval batch:
- The model types speaker pronouns Person as well as User, so the mention
resolver now applies its user rule to them too. Person:'you'/'I'/'assistant'
had merged, under global identity, into hubs spanning up to 110 unrelated
sessions and headed the user's own facts.
- An all-lowercase mention of a global type is a generic noun ('home', 'area',
'city'), so it stays per chunk instead of joining ~50 sessions.
Claude-Session: https://claude.ai/code/session_01BczgXHGBrKsHSQUGB8P8te
antejavor
force-pushed
the
feat/typed-relation-model
branch
from
September 30, 2026 07:52
9640916 to
f9d4f07
Compare
This was referenced Sep 30, 2026
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #344. Closes #357. Closes #365. Closes #371.
ensure_lookupadds both a unique constraint and an index. In Memgraph the constraint alone doesn't index.Implements map #344 (typed relation model) with hand-written ontologies through
ManualStrategy. The derived vocabulary (LlmRecommendationStrategy) stays an interface until its evidence run, #372, passes. The revised plan puts that second, then a full eval as confirmation.What changes
hygm(new workspace package, pure Python)NodeType(label, description, identity), whereidentityisglobal,chunkorspan(Entity identity across chunks #346, Values, not links: do entity attributes belong in the model #361).RelationType(label, description, start_labels, end_labels). Empty endpoints mean any declared label.validate_modelis the hard gate: labels unique, endpoints declared, andUserrequiresPerson(WhichUsermentions are the user #358).ManualStrategy(YAML);OwlImportStrategy(owl:Class/owl:ObjectProperty,rdfs:domain/rdfs:range, includingowl:unionOf; thehygm[owl]extra);LlmRecommendationStrategyplus anObserverprotocol, interface only.unstructured2graph
Ontologyloads throughManualStrategy. The YAML gainsidentity,start_labelsandend_labels.enforce_relation_domain_rangeis the post-hoc half of the one-spec, two-compilation model (Extraction-time constraints vs post-hoc validation #348). It runs underenforce_ontologyand flags a mismatch, never deleting it. ADR 0004 is amended to cover relations.ontology_reportreturns integrity counts and Does LightRAG consume domain/range too #349's coverage signal: untyped relationships versus none at all, and declared types with zero instances.from_documents/Document/Segment: text stored verbatim, oneChunkper document, with turn segments.unstructured's partitioner rewrites text (it drops list bullets and folds tables), which would break turn offsets.ensure_lookupadds a unique constraint and an index onChunk.hashand<workspace>.entity_id(unstructured2graph: entity_id is MERGEd on but never indexed or constrained #357). Verified withEXPLAIN: in Memgraph a uniqueness constraint alone does not index, so every lookup was a label scan.connect_chunks_to_entitiesis an indexed lookup instead of a cartesian product.GLiNER2Backend, rewritten on
gliner2.joint_ieJointSchemais compiled once and held (unstructured2graph: gliner2's compiled-schema cache is keyed on a memory address #365). Both candidate caps are 4096 (unstructured2graph: gliner2 cuts relation candidates alphabetically by type on ties #371).Usermentions are the user #358):(:User {user_id});Person;MENTIONED_INwritten for every chunk a node is mentioned in (Entity identity across chunks #346).valid_at(a Memgraphdatetime, from the turn, Resolving relative dates against valid_at #364) andconfidence. Superseded facts are kept.sessions-graph
reconcile_sessionsends one segmentedDocument, with the speaker andAction.timestampper source.anon-<session_id>user (The starter relation vocabulary #347, context-graph-eval: injector never sets a user, so reconciled graphs have no (:User) nodes #354). It isanon-becausevalidate_user_idrejects:.ReconciliationSummarygainsnonconformant_entitiesandnonconformant_relations.eval
ontologies/longmemeval.yaml, Values, not links: do entity attributes belong in the model #361's hand vocabulary. A non-conformant session is reported as an error.Follow-ups from running this on a 500-session eval batch:
25d7349). Linking, label promotion, the domain/range check andontology_reportrescanned the whole workspace after every session: 9 s per session at 8,000 entities, 10 s at 16,000, on the way to about 400,000.Persongo through the user rule too (f9d4f07). Without this,Person:'you'/'I'/'assistant'merged, under global identity, into hubs spanning up to 110 sessions.homeorareamerged across about 50 sessions.The eval work built on this (hybrid retrieval and friends) is a separate stacked PR. The judge-scoring fix (#387) is also separate, off
main.Found on the way
WHERE (a:X OR b:X)is planned asFilter (a:X), (b:X), an AND. Every edge with only one endpoint in the workspace, e.g. from(:User), silently escaped the check. The workaround is$workspace IN labels(a) OR ..., covered by a regression test. Worth reporting upstream.duration.betweendoesn't exist in Memgraph. Date arithmetic is datetime subtraction.Verification
ruffandtyare clean.personal_best → 25:50,visited → Museum of Modern Art,studied → Business Administration;Not in this PR
question_datefor the retrieval agent, map Map: Context-graph emergence pipeline — reconciliation, retrieval, eval #297).Usermentions in them are dropped.hygmis not on PyPI yet. The unstructured2graph CI job installs it from the workspace, and a release needshygmpublished first.https://claude.ai/code/session_01BczgXHGBrKsHSQUGB8P8te