Skip to content

unstructured2graph/sessions-graph: write turn provenance at extraction time - #400

Merged
antejavor merged 1 commit into
mainfrom
feat/392-turn-provenance
Oct 1, 2026
Merged

antejavor merged 1 commit into
mainfrom
feat/392-turn-provenance

Conversation

@antejavor

@antejavor antejavor commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Closes #392. Part of map #390. The decisions are recorded on #392.

Why

Hybrid retrieval needs to know which turn each entity and fact came from. Until now the GLiNER2 backend dropped that at write time: a session is a single chunk, so entities linked only to the whole session. The eval then guessed the turn afterwards (link_turns in #389), by timestamp or substring match, which links a short or common entity name to every turn that contains the string. The backend already knows the exact turn, because it extracts one window per segment, and a segment is a turn.

What changes

  • Segment.source_id: an optional, opaque id that unstructured2graph never interprets. sessions-graph fills it with the turn's action_id. When identical turns are merged into one segment, it takes the first turn's id.
  • MENTIONED_IN.sources: the turn ids that mention each entity in the chunk, merged by union, so re-ingesting is idempotent.
  • One edge per turn. upsert_extracted_relationships keys on (type, head, tail, chunk, source_id). A fact said in two turns becomes two edges, each with its own valid_at. Before, they collapsed into one edge, and SET r.valid_at kept whichever turn was written last.
  • Every edge carries source_id, role (user or assistant) and text: the sentence or sentences covering both endpoint spans, at most 300 characters. The full turn is one hop away through source_id. role tells the two speakers apart; it isn't a filter.
  • Without a source_id, e.g. plain document ingestion, keys and behaviour are unchanged.

No extra model calls: all of it comes from spans GLiNER2 already returns.

Tests

All run against a real Memgraph with the GLiNER2 fakes:

  • exact sources for an entity mentioned in two turns;
  • an edge's turn, speaker and sentence;
  • the same fact in two turns becomes two edges with their own times;
  • re-ingesting with provenance is idempotent;
  • the sentence span and the 300-character cap.

The sessions-graph reconciliation e2e test now checks that the edge's source_id is the user message's action_id, and that sources lists both turns.

Results, run locally:

  • unstructured2graph: 142 passed, 10 skipped (they need gliner2 or LightRAG);
  • the real-GLiNER2 e2e tests, run separately in a gliner2 environment: 3 passed;
  • sessions-graph: 71 passed;
  • ruff and ty are clean.

Follow-ups

…n time

Segment gains an optional, opaque source_id, which sessions-graph fills with
the turn's action_id. The GLiNER2 backend writes exact provenance from the
spans the model already returns, with no extra model calls:

- MENTIONED_IN.sources lists the turns that mention each entity, merged by
  union so re-ingesting is idempotent;
- each extracted relationship is keyed on (type, head, tail, chunk,
  source_id): a fact said in two turns is two edges, each with its own
  valid_at, which also ends the last-write-wins SET of valid_at within a
  session;
- each edge carries source_id, role (the speaker) and text (the sentence or
  sentences covering both endpoints, capped at 300 characters).

Without a source_id, keys and behaviour are unchanged.

Closes #392

Claude-Session: https://claude.ai/code/session_01BczgXHGBrKsHSQUGB8P8te
@antejavor
antejavor merged commit dea36a3 into main Oct 1, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Write turn links and edge source sentences at extraction time

1 participant