Skip to content

Near-synonymous relation labels all fire on the same pair #360

Description

@antejavor

Part of #344

Question

The model asserts every plausible relation label rather than choosing one. Does the vocabulary have to be label-disjoint, does the writer arbitrate, or does the graph just carry all of them?

Measured in #350 over 10 real sessions: 14% of distinct (head, tail) pairs in the constrained arm carry more than one relation label (21% permissive).

'user' -> 'meal prep':     practices, prefers, studied
'user' -> 'yoga pants':    owns, prefers, purchased
'user' -> 'tennis racket': owns, prefers, purchased
'frida' -> 'retablos':     owns, prefers, purchased, studied

Every one of those is individually plausible and collectively noise. It matters most for the label the corpus most needs: #350's single-session-preference question wants what the user prefers, and prefers is indistinguishable from owns and purchased on the same pair.

This is a vocabulary-shape problem, not a bug: the labels were hand-written for the prototype, but #347 made vocabularies derived automatically, and a proposer asked for relation types will happily propose practices, prefers and studied for the same corpus.

Open:

  1. Disjointness as a validation rule. Should hygm.validation reject (or warn on) a model whose relation types overlap in meaning — and is that even checkable? Structural checks (dangling labels, duplicates) are mechanical; "these two labels mean nearly the same thing" is not.
  2. Or accept multi-labelling as the graph's shape. Then it is a retrieval concern: three edges where one is true inflates the payload and gives the retrieval agent three different Cyphers that all look right. Map Map: Context-graph emergence pipeline — reconciliation, retrieval, eval #297 owns retrieval quality, but the graph contract is decided here.
  3. Or arbitrate at write time. Keep the highest-confidence label per (head, tail), or per (head, tail, window). That is a deliberate information loss, so it collides with ADR 0004's instinct to keep everything — and unlike Entity identity across chunks #346's merging (where every mention keeps its MENTIONED_IN edge, so nothing is destroyed), this one genuinely discards a claim the model made.
  4. Does GLiNER2 already have a mechanism? Every compiled schema carries UniqueRelationPair and UniqueRelationSlot (What JointSchema actually enforces #345). Does UniqueRelationPair mean "one edge per (head, tail)" — and if so, why does it not fire here? The likely answer is that each label is a separate relation type and the constraint is per type, which if confirmed means the mechanism exists but is scoped wrong for this problem.
  5. Cardinality is the blunt instrument and it is expensive. Wire a constrained ontology and read the edges #350 measured max_per_head=1/max_per_tail=1/at_most(per_head=1) dropping 68% of relations (64 -> 20) while still reporting feasible=True — no signal that anything was lost.

Interacts with #353 (derivation's output contract), #355 (a reader deciding which edges to trust), and #359 (the other half of "what shape must a derived vocabulary have").

Evidence: prototype/gliner2-constrained-edges, section C2 of the prototype README.

Activity

added a commit that references this issue on Sep 28, 2026

antejavor commented on Sep 28, 2026

@antejavor
ContributorAuthor

Resolution

The ticket asked whether the vocabulary must be label-disjoint, whether the writer must arbitrate, or whether the graph carries every label. It carries every label. The premise that the extra labels are noise mostly doesn't hold; where it does hold, the cause is a single label's precision, not overlap between labels; and nothing available at decode or write time separates true labels from false ones.

Evidence: section H of the prototype README (32d1928): multilabel.py / .log and relation_lever.py / .log. Clean constrained arm (#365).

The multi-label pairs, read

33 of 244 pairs (14%) carry more than one label. 21 of them are owns + purchased, which are co-true rather than synonyms: you bought the tennis racket and you own it. 'user' -> 'black jeans' carries owns, purchased and prefers, and all three hold ("my new black jeans from Levi's, which I'm really loving").

Confidence can't arbitrate. A false prefers scores 0.94 (tennis racket) while a true one scores 0.75 (black jeans), and tennis racket's owns 0.98 would outvote a true purchased 0.96.

The defect is one label's precision. Co-fired prefers is right on 6 of 17. It's true where the text says really loving, especially interested, particularly fascinated, particularly drawn to. It's false on tennis racket, yoga pants, tops, the sarcophagus's gold, and mint (an app the assistant listed): prefers fires on mention, not on preference. prefers co-fires on 17 of its 27 pairs (63%).

Item 4: why UniqueRelationPair doesn't fire

It's scoped per label: it skips any existing relation with a different label (joint_ie/constraints.py:139), so it stops a duplicate owns and never stops owns + prefers. There's no built-in cross-label constraint, but JointSchema.constraint() accepts any Constraint subclass (schema.py:95), applied inside beam expansion. So a decode-time "mutually exclusive labels" rule is buildable. It's rejected on the evidence above, because it would destroy the co-true labels.

A finding that reaches #353: the joint compiler drops relation descriptions

The first lever I tested for prefers' precision was its description, sharpened to "explicitly says they like, love, enjoy or favour it — not merely owning, buying, using or asking about it". The edge list came back byte-identical. The reason is joint_ie/compiler.py:27-28: relations compile to {name: {"head": "", "tail": ""}} with empty json_descriptions, and only entity descriptions reach the model. Every relation description in #350's and #361's vocabularies was inert, and a derived vocabulary's relation descriptions will be too.

The name is the only lever, and it trades recall for precision rather than separating them:

label raw true kept (of 6) false kept (of 11)
prefers 111 6 11 baseline
likes 261 6 10 fires 2.4× as often, same precision
says_they_like 223 3 5 halves true and false alike, and heads on 'assistant' ×26 (the one who "says")

Renaming one label also moves the rest of the graph (other relations 893 → 876 → 858 raw), the same beam tax #361 measured when adding types.

This matters less than the ticket feared for the preference questions. #361 measured them as 6/6 prose, where an LLM judges a summary. They need candidate evidence of interest, and there recall matters more than precision; no single prefers edge is the answer.

Decisions

  1. The graph carries every label (items 1–3, 5). Each label is its own claim and its own edge, with valid_at and MENTIONED_IN. Rejected:
    • write-time arbitration (keep highest confidence per pair): measured to delete true facts, and it discards claims outright, against ADR 0004;
    • decode-time exclusive-label groups: the same loss;
    • a disjointness rule in hygm.validation: owns/purchased/prefers are distinct concepts that happen to co-hold, so a correct rule rejects nothing, and "overlap in meaning" isn't mechanically checkable anyway;
    • cardinality caps: Wire a constrained ontology and read the edges #350 measured 68% of relations lost while feasible still reads true.
  2. Per-label precision is accepted as a known limit, with a reporting-only co-firing signal. Per relation, report the share of its pairs where it co-fires with another label (prefers 63%). A label that almost never fires alone isn't discriminating, and the revise loop can rename, merge or drop it. That's the same shape as How wide should a relation's declared range be #359's widening-candidate report, and nothing gates on it. Rejected:
    • a hard co-firing gate: it deletes black jeans' true prefers;
    • a lexical-evidence gate: brittle against drawn to and fascinated by, and it rebuilds the per-label rules derivation was meant to avoid;
    • dropping prefers from extraction: it removes the candidate evidence retrieval searches, e.g. prefers(User -> Topic:'advanced settings') anchors the Premiere Pro answer.

Consequences

antejavor commented on Sep 28, 2026

@antejavor
ContributorAuthor

An exception from #366: "carry every label" assumed co-firing labels are co-true (owns + purchased). Derived vocabularies also produce modal/tense variants of one relation, which contradict: plans_to_visit(user -> Museum of Modern Art) fired on the same visit as visited_location, and visited_location co-fires with a variant on 21 of its 30 pairs (plans_to_sell with owns_item 12/14). Since the name is the only lever and it fails to separate them, #353's proposer may not emit such variants; the episodic summary still carries plans. Co-true multi-labeling stands as decided here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions