You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Eval: one user per question, retrieval scoped to it #396
inject.py gives every injected session EVAL_USER_ID = "longmemeval-user", and hybrid retrieval searches the whole graph for every question. In LongMemEval each question's haystack belongs to a different person. In the #391 run (100 questions, 501 sessions):
only 31% of retrieved turns and 17% of retrieved facts came from the question's own sessions;
answers counted other personas' items: 4 weddings for 3, 6 furniture items for 4, another persona's $1,200 bag added to the right $800;
one User holds 897 purchased and 291 lives_in edges, so the user_facts lane gathers 100 people's facts;
cross-persona rows crowd the question's own evidence out of the top results. That drives most of the 25 retrieval and partial-retrieval misses.
What to do
Each question gets its own User; its sessions hang off that user.
Every hybrid lane is scoped to the asking user's sessions.
If two questions share a haystack session, inject it under each question's user, or link the shared session to both users. Decide in the PR.
Part of #390
Problem
inject.pygives every injected sessionEVAL_USER_ID = "longmemeval-user", and hybrid retrieval searches the whole graph for every question. In LongMemEval each question's haystack belongs to a different person. In the #391 run (100 questions, 501 sessions):Userholds 897purchasedand 291lives_inedges, so theuser_factslane gathers 100 people's facts;What to do
User; its sessions hang off that user.Acceptance
Retrieved context comes only from the question's own sessions. The baseline is re-measured on the rebuilt graph with one full-hybrid run.