Skip to content

Run the derivation contract once and read the vocabulary #366

Description

@antejavor

Part of #344

Question

Does #353's derivation contract produce a usable vocabulary when it actually runs? Nobody has run it. Whether an LLM proposes a sensible domain vocabulary from session text, and prunes a GLiNER2 observation table sensibly, is the same kind of question #350 answered for extraction: by reading real output, not by argument.

What to build (throwaway, branch only, same convention as #350)

The #353 pipeline, end to end, once:

  1. Propose. Whole sessions, time-stratified, up to a prompt budget, go to an LLM. It returns domain entity types (label plus description), relation names, and any corpus-specific value types, on top of the fixed core (User, Person, Duration, Quantity, Money, Date, TimeWindow). The prompt states that relation names are all the extractor sees (Near-synonymous relation labels all fire on the same pair #360: relation descriptions are dropped by joint_ie/compiler.py:27).
  2. Observe. One permissive GLiNER2 pass over a larger sample: per-relation endpoint type tables with example edges, and per-type mention stats. Hold every compiled schema for the whole run (unstructured2graph: gliner2's compiled-schema cache is keyed on a memory address #365).
  3. Prune. The LLM keeps sensible endpoint pairs, sets identity per derived type, and drops zero-instance types and relations. It doesn't rename or add.
  4. Validate. Run hygm.validation's hard gate, including Which User mentions are the user #358's rule that User in a range requires Person there too.
  5. Stability. Derive a second time from a disjoint propose sample, then report behavioural agreement: both vocabularies over one shared sample, types aligned by span co-assignment.

Run it on the eval corpus (LongMemEval) and, if feasible, on a sample of real Claude Code sessions. #347 said each corpus gets its own vocabulary, and production is coding sessions.

What to read, and judge by human reading of sampled output (the map's decision criterion)

  • Is the derived vocabulary better or worse than Wire a constrained ontology and read the edges #350's hand-written stand-in on the same 10 sessions? Compare with the committed baseline (1004 raw edges; prototype/gliner2-constrained-edges).
  • Does prune recover visited -> Organization (How wide should a relation's declared range be #359's MoMA edge) and drop knows -> Topic?
  • Are the per-type identity calls right? Person/Organization-like types should be global, generic ones chunk.
  • Does the proposer produce near-overlapping relation names (Near-synonymous relation labels all fire on the same pair #360), and if so, what's the co-firing rate?
  • How much do the two derivations agree behaviourally, and does the disagreement land where a human would expect ambiguity?
  • Cost: LLM tokens per derivation, and wall clock for the observe pass.

Not in scope

The evolution loop (revision, versioning, migration). This ticket runs v1 once, plus the stability derivation.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions