Skip to content

Eval: one user per question, retrieval scoped to it #396

Description

@antejavor

Part of #390

Problem

inject.py gives every injected session EVAL_USER_ID = "longmemeval-user", and hybrid retrieval searches the whole graph for every question. In LongMemEval each question's haystack belongs to a different person. In the #391 run (100 questions, 501 sessions):

  • only 31% of retrieved turns and 17% of retrieved facts came from the question's own sessions;
  • answers counted other personas' items: 4 weddings for 3, 6 furniture items for 4, another persona's $1,200 bag added to the right $800;
  • one User holds 897 purchased and 291 lives_in edges, so the user_facts lane gathers 100 people's facts;
  • cross-persona rows crowd the question's own evidence out of the top results. That drives most of the 25 retrieval and partial-retrieval misses.

What to do

  • Each question gets its own User; its sessions hang off that user.
  • Every hybrid lane is scoped to the asking user's sessions.
  • If two questions share a haystack session, inject it under each question's user, or link the shared session to both users. Decide in the PR.
  • Land it in the same eval-graph rebuild as Write turn links and edge source sentences at extraction time #392.

Acceptance

Retrieved context comes only from the question's own sessions. The baseline is re-measured on the rebuilt graph with one full-hybrid run.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions