Skip to content

context-graph: document GQLAlchemy baseline comparison and pilot findings - #482

Draft
antejavor wants to merge 4 commits into
mainfrom
docs/gqlalchemy-pilot-findings
Draft

antejavor wants to merge 4 commits into
mainfrom
docs/gqlalchemy-pilot-findings

Conversation

@antejavor

@antejavor antejavor commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Context Graph's GQLAlchemy pilot demonstrated fresh-session recall, but the first attempts could not compare it fairly with ordinary Codex context. This PR records the infrastructure failures and completes a small retrospective baseline comparison, using four new GPT-6 Luna sessions and the existing transport-valid graph pair.

All three candidates passed the multigraph issue check. Only ordinary context passed every automated quality gate: the notes-allowed candidate had a failing added test, and the graph candidate regressed default edge properties. The graph fresh session was faster (28.21s versus 35.82s), but approximately 103s of serial memory processing outweighed that saving in this case. No practical advantage or causal quality effect is established; the notes-allowed agent never wrote notes, so active file-note recall remains untested.

The reports separate test-qualified candidates from accepted resolutions, keep the earlier failed attempts in the usage ledger, and state the limitations: one case/replicate, earlier graph collection, cooperative isolation, extraction integrity warnings and pending blinded human review. They propose focused follow-ups for independent embedding scheduling and opt-in live capture/recall diagnostics. Package behavior is unchanged.

Validation: matching model/settings, frozen resource and evaluator hashes, and normalized prompts verified programmatically. Final issue acceptance, property compatibility and translator tests ran independently against real Memgraph for every candidate. Four process-supervisor tests, Ruff and targeted type checks pass. Account-level weekly usage reported 20% before and 21% afterward, with quota checks before every new session; no paid credits or resets were requested. New containers are stopped, copied credentials removed, and the original GQLAlchemy checkout is unchanged. Raw traces remain local; this is a report, not a published benchmark dataset.

Related maps: #297 and #374. Draft for review of the evidence and proposed implementation scope.

@antejavor antejavor changed the title context-graph: document GQLAlchemy pilot failures and rerun gates context-graph: document GQLAlchemy baseline comparison and pilot findings Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant