Repository navigation
unstructured2graph: extraction-quality eval, GLiNER2 vs LightRAG - #332
Merged
Merged
Conversation
antejavor
added a commit
that referenced
this pull request
Sep 15, 2026
Prompted by review feedback on #332 (the extraction-quality eval), which was reaching through backend.wrapper.afinalize() to finalize a LightRAGBackend between repeated runs -- an ExtractionBackend consumer shouldn't need to know LightRAGBackend wraps a MemgraphLightRAGWrapper at all. Exposed as a LightRAGBackend-specific method, not added to the ExtractionBackend protocol: GLiNER2Backend has no equivalent persistent resource needing teardown, so this isn't a contract every backend must satisfy.
antejavor
added a commit
that referenced
this pull request
Sep 15, 2026
Standards, hard violations: - Cypher injection risk. _read_entities/_read_ontology_conformance/ _read_relations now validate workspace/text_property/relation_types via _require_valid_identifier (reused from unstructured2graph.memgraph, same helper #330's upsert_typed_relationships review fix added) before interpolating them, instead of trusting the caller. - Incomplete docstrings. Added Attributes/Args/Returns/Raises to GoldEntity, GoldRelation, GoldRecord, PRF1 (+ its properties), BackendReport, run_backend(), run(), _run_repeated(), print_report(), BackendSpec. Standards, judgment calls: - afinalize() reach-through. Eliminated by #330's new LightRAGBackend.afinalize() passthrough (merged in) -- run() now calls backend.afinalize() instead of backend.wrapper.afinalize(). - Duplicated repeat-loop orchestration. LightRAG's and GLiNER2's near- identical "build, run_backend, maybe finalize, print" loops collapsed into one shared _run_repeated(), parameterized by a BackendBuilder callable (LightRAG rebuilds + resets counters fresh per repeat; GLiNER2 reuses one loaded model via _reuse_gliner2_builder) and an optional finalize hook. - Configuration as loose strings. name/text_property/relation_types/ llm_model bundled into a new frozen BackendSpec (LIGHTRAG_SPEC/ GLINER2_SPEC) instead of traveling run_backend()'s call chain as four separate parameters. Spec: - [High] Zero F1 reported as n/a. PRF1.f1's `if not p or not r` treated a real, computed 0.0 the same as "undefined" (None), since 0.0 is falsy. Fixed to `if p is None or r is None`, with an explicit 0.0 for the precision=recall=0 case (would otherwise divide by zero). Verified: zero- overlap now scores f1=0.0; genuine "nothing gold and nothing predicted" still scores None; perfect match still scores 1.0. - [High] LightRAG cost undercounted. _counting_llm_wrapper now tokenizes every message in history_messages (each billed as real provider input, same as prompt/system_prompt), not just the two it previously counted. Verified: a call with two history messages now counts ~5x the prompt tokens of an otherwise-identical call with none. - [High] Backends could target different Memgraph instances. New _pin_memgraph_env() sets MEMGRAPH_URL *and* MEMGRAPH_URI (+ empty MEMGRAPH_USER/MEMGRAPH_USERNAME/MEMGRAPH_PASSWORD defaults) unconditionally before any backend is built, instead of relying on lightrag_memgraph.core's _bridge_lightrag_env_names(), which only mirrors MEMGRAPH_URL onto MEMGRAPH_URI when MEMGRAPH_URI isn't already set -- meaning a stale ambient MEMGRAPH_URI would previously have silently won over --memgraph-url for LightRAG's storage backend while this eval's own client used the correct one. Verified: a pre-set stale MEMGRAPH_URI is now overridden. - [Medium] Gold span scoring discards spans. Not changed to offset-based matching (would reintroduce the exact cross-backend brittleness normalized- text matching was chosen to avoid -- LightRAG's LLM-normalized text doesn't char-align with source spans the way GLiNER2's does). Instead documented explicitly and precisely: GoldEntity/GoldRelation's docstrings now state start/end exist only for build_gold_corpus.py's own verbatim-text assertion and are never read back for scoring, and that matching is by unique normalized text per chunk (a set, collapsing repeated identical mentions), not span position or per-occurrence count. print_report() prints the same caveat. Verified: ruff check/format clean, ty check clean, 71 unstructured2graph unit tests pass (was 70 -- LightRAGBackend gained one afinalize() test in #330). Ran the full rewritten script end to end against both backends (real GLiNER2 model + real LightRAG/OpenAI) on a live dedicated Memgraph -- both produce full reports with the new cost/token-breakdown and gold-span-caveat lines. Independently verified each of the four Spec fixes in isolation (PRF1.f1's three cases, Cypher-identifier rejection, history_messages token delta, MEMGRAPH_URI override) rather than relying on the end-to-end run alone to prove them.
antejavor
added this pull request to stack #334
September 16, 2026 07:14
antejavor
added a commit
that referenced
this pull request
Sep 16, 2026
* unstructured2graph: add pluggable extraction backends, GLiNER2Backend Entity/relation extraction is now behind an ExtractionBackend protocol (workspace_label + async aingest_chunk) instead of being hardwired to LightRAG, so from_texts()/from_unstructured() can drive different extractors through one seam. LightRAGBackend adapts the existing MemgraphLightRAGWrapper; GLiNER2Backend is a new local, LLM-free extractor using GLiNER2 (https://github.com/fastino-ai/GLiNER2) for joint entity+relation extraction with no API key or network cost. - extraction_backend.py: ExtractionBackend protocol + LightRAGBackend - gliner2_backend.py: GLiNER2Backend, entity identity scoped to (chunk, entity_type, normalized text) since GLiNER2 does no cross-chunk coreference, relations written as per-label Cypher edge types (e.g. :works_for) drawn from Ontology.relation_types - ontology.py: Ontology gains an optional relation_types field, read directly by GLiNER2Backend as its extraction schema - memgraph.py: new upsert_typed_relationships() helper - loaders.py: lightrag_wrapper param renamed to extraction_backend (pre-1.0 package, no compat shim; the one external call site in sessions-graph/core.py is updated) - pyproject.toml: new optional `gliner2` extra Also lowers the workspace's transformers floor from >=5.0.0rc3 to >=4.57.6 (latest patched 4.x): gliner2[local] hard-pins transformers<5 on every published release, which is incompatible with >=5.0.0rc3 (the fix for CVE-2026-1839, an RCE in Trainer's checkpoint loading, with no 4.x backport). Confirmed via `uv lock`/`uv sync --all-extras` that no version satisfies both constraints -- this tradeoff was discussed and confirmed explicitly before making the change. Verified against a live Memgraph with the real GLiNER2 model (not just the fake-model unit tests): clean entity typing and a real typed relation edge. * unstructured2graph: address review feedback on #330 Standards, blocking: - Security floor weakened. Reverted -- transformers>=5.0.0rc3 restored in root pyproject.toml. gliner2 is no longer a pyproject.toml extra of unstructured2graph at all (gliner2[local] hard-pins transformers<5, which is unconditionally incompatible with the floor under uv's workspace-wide resolution -- confirmed again via `uv lock`). Isolated instead: manual `pip install 'gliner2[local]>=2.0.0'` in your own environment, documented in gliner2_backend.py's module docstring, unstructured2graph/README.md, the example, and test_e2e_gliner2.py (whose skip in CI is now permanent and documented as such, not a coverage gap). uv.lock regenerated -- gliner2/peft removed, transformers back to 5.17.0. - Cypher injection risk. upsert_typed_relationships() now validates node_label, match_key, every relation type, and every edge-property key via a new shared _require_valid_identifier() (memgraph.py), not just relation_type. GLiNER2Backend.__init__ now validates `workspace` itself (raises ValueError) -- the actual source the review flagged as reaching these queries arbitrarily; fixing it there closes every downstream Cypher use of workspace (create_nodes_from_list, connect_chunks_to_entities, promote_*_to_labels, upsert_typed_relationships) in one place rather than requiring the same check duplicated across pre-existing functions this PR didn't otherwise touch. New tests for both layers (test_memgraph.py, test_e2e.py, test_gliner2_backend.py). - AGENTS.md stale. Updated the unstructured2graph row to describe the pluggable ExtractionBackend architecture. - Public API documentation incomplete. Added Args/Raises to RelationType (+ EntityType, for consistency), GLiNER2Backend.__init__, _resolve_entity_workspace, and upsert_typed_relationships. Standards, judgment calls: - Typed records. GLiNER2Backend._extract_sync now returns list[ExtractedEntity]/list[ExtractedRelation] (frozen dataclasses) instead of list[dict[str, Any]]. - Four-object reach for workspace. Added MemgraphLightRAGWrapper.workspace (lightrag_memgraph/core.py) so LightRAGBackend.workspace_label is one attribute access instead of reaching through get_lightrag(). chunk_entity_relation_graph. New tests in lightrag-memgraph's own suite. - GLiNER2 e2e never runs in CI. Now explained rather than left implicit: test_e2e_gliner2.py's docstring states this is permanent, not pending, given the security-floor isolation above. Spec: - [P1] Root quick start crashes. README.md's unstructured2graph quick start now uses extraction_backend=LightRAGBackend(lightrag). - [P2] Advertised joint extraction wasn't joint. _extract_sync now makes one model.create_schema().entities(...).relations(...) + model.extract() call instead of two separate extract_entities()/extract_relations() calls -- confirmed against a real gliner2==2.0.0 install that this returns {"entities": ..., "relation_extraction": ...} in one shape. Verified live this fixes the specific failure mode raised (independent inference passes disagreeing on a span for the same entity between the two calls) for entities the model recognizes in both its entity and relation output. Also verified live that it does NOT make every relation matchable -- GLiNER2's relation task can name a head/tail its entity task never surfaces under the configured schema at all, joint call or not -- and corrected _extract_sync's docstring to say so accurately rather than overclaim. aingest_chunk() still correctly skips-and-logs that case: no extracted entity_type exists to write such an endpoint under. - [P2] Workspace override silently broke linking. _resolve_entity_workspace now raises ValueError when entity_workspace is given and doesn't match extraction_backend.workspace_label, instead of silently using the mismatched value. Updated/added tests in test_loaders.py. Verified: ruff check/format clean; ty check clean (gliner2's own unresolved lazy import is a narrow, commented `ty: ignore[unresolved-import]` -- see gliner2_backend.py, since it's permanently outside the managed dependency graph now, not a normal optional extra ty could resolve via --all-extras). 86 unstructured2graph tests (was 62) + 37 lightrag-memgraph tests (was 35) + 55 sessions-graph tests all pass, including real-model GLiNER2 e2e tests run against a manually-installed gliner2 (uv pip install --no-config, to simulate the documented manual-install path without touching the workspace lock) and a live dedicated Memgraph. * unstructured2graph: add LightRAGBackend.afinalize() passthrough Prompted by review feedback on #332 (the extraction-quality eval), which was reaching through backend.wrapper.afinalize() to finalize a LightRAGBackend between repeated runs -- an ExtractionBackend consumer shouldn't need to know LightRAGBackend wraps a MemgraphLightRAGWrapper at all. Exposed as a LightRAGBackend-specific method, not added to the ExtractionBackend protocol: GLiNER2Backend has no equivalent persistent resource needing teardown, so this isn't a contract every backend must satisfy.
Deterministic entity/relation-level comparison of the two ExtractionBackend implementations (#330), separate from context-graph-eval's end-to-end memory recall loop -- that harness's Golden corpus has no entity/relation gold labels and reconcile_batch() is hardwired to lightrag_wrapper, neither of which fits this question. No LLM judge: precision/recall/F1 against gold spans is a countable metric, not a quality judgment. - evals/build_gold_corpus.py: generates gold_corpus.jsonl (41 chunks) from reviewable Python -- 36 sampled from context-graph/eval's LongMemEval corpus (real text, but carries almost no named-entity relations), 5 authored specifically to exercise works_for/founded/located_in relation scoring - evals/extraction_quality.py: runs both backends through the real from_texts() pipeline against a dedicated Memgraph instance, reads back what each wrote, scores type-agnostic and type-sensitive entity P/R/F1 plus relation P/R/F1 (GLiNER2 only -- LightRAG has no relation_types concept), tracks latency and (for LightRAG) LLM call/token cost - evals/README.md: how to run it, what each column means, and a documented gold-corpus caveat found by actually running it -- gold marks only salient entities, not everything the ontology's broad categories (Concept, Event, Artifact...) could arguably cover, so precision is a floor, not a measured ceiling Verified live: GLiNER2 end-to-end against a real dedicated Memgraph instance (port 7691, isolated from context-graph-eval's own 7689). LightRAG needs OPENAI_API_KEY, not available in this environment -- run --backend lightrag yourself before trusting any LightRAG-vs-GLiNER2 delta.
unstructured2graph (and context-graph-eval) had no .env loading anywhere, unlike integrations/langchain-memgraph and integrations/mcp-memgraph, which already do a bare load_dotenv() in their own tests -- surfaced concretely when the extraction-quality eval's lightrag backend silently skipped because the repo-root .env was never read into the process environment. main() now does the same bare load_dotenv() call (soft dependency via the test extra), which finds the repo-root .env through python-dotenv's own upward search regardless of whether the script runs from unstructured2graph/ or the repo root. Verified with OPENAI_API_KEY unset in the shell and only present in .env: both backends now run end to end. First real side-by-side numbers (41-chunk gold corpus, single run each): lightrag ent F1 65.3% (P 49.4/R 96.3), 4.83s/chunk median, 83 calls/167k tokens gliner2 ent F1 52.0% (P 36.0/R 93.9), 0.27s/chunk median, 0 calls, $0 (relation scoring is gliner2-only; ontology conformance 93.2% vs 100%) Both precision numbers are understated by the same gold-corpus partial- annotation caveat evals/README.md already documents.
Answers the actual question a token count doesn't: how much did this run cost. MODEL_PRICING_PER_1M_TOKENS is a small, pinned (point-in-time, not fetched live) per-model price table; estimated_cost_usd() multiplies it by the approximate prompt/completion token split already tracked. A model missing from the table reports "unpriced", never a silent $0 -- the distinction matters since GLiNER2 genuinely is $0 (n/a, no LLM calls at all) and an unpriced LightRAG model is a gap in the table, not a finding. Repeat run against the same 41-chunk gold corpus, same OPENAI_API_KEY-via- .env path verified last commit: 83 calls, 152k prompt + 16k completion tokens, ~$0.03 at gpt-4o-mini pricing. Entity F1 moved from 65.3% to 57.7% between the two runs (both LightRAG, nothing else changed) -- exactly the LLM-extraction run-to-run variance the README already warns about, now visible in the wild rather than theoretical.
Follows #330's review-feedback fix: gliner2 is no longer a pyproject.toml extra (isolated from the workspace's transformers>=5.0.0rc3 floor instead of weakening it). Updates this eval's own README/module docstring off the now-removed `pip install unstructured2graph[gliner2]` to the manual `pip install 'gliner2[local]>=2.0.0'` path.
Standards, hard violations: - Cypher injection risk. _read_entities/_read_ontology_conformance/ _read_relations now validate workspace/text_property/relation_types via _require_valid_identifier (reused from unstructured2graph.memgraph, same helper #330's upsert_typed_relationships review fix added) before interpolating them, instead of trusting the caller. - Incomplete docstrings. Added Attributes/Args/Returns/Raises to GoldEntity, GoldRelation, GoldRecord, PRF1 (+ its properties), BackendReport, run_backend(), run(), _run_repeated(), print_report(), BackendSpec. Standards, judgment calls: - afinalize() reach-through. Eliminated by #330's new LightRAGBackend.afinalize() passthrough (merged in) -- run() now calls backend.afinalize() instead of backend.wrapper.afinalize(). - Duplicated repeat-loop orchestration. LightRAG's and GLiNER2's near- identical "build, run_backend, maybe finalize, print" loops collapsed into one shared _run_repeated(), parameterized by a BackendBuilder callable (LightRAG rebuilds + resets counters fresh per repeat; GLiNER2 reuses one loaded model via _reuse_gliner2_builder) and an optional finalize hook. - Configuration as loose strings. name/text_property/relation_types/ llm_model bundled into a new frozen BackendSpec (LIGHTRAG_SPEC/ GLINER2_SPEC) instead of traveling run_backend()'s call chain as four separate parameters. Spec: - [High] Zero F1 reported as n/a. PRF1.f1's `if not p or not r` treated a real, computed 0.0 the same as "undefined" (None), since 0.0 is falsy. Fixed to `if p is None or r is None`, with an explicit 0.0 for the precision=recall=0 case (would otherwise divide by zero). Verified: zero- overlap now scores f1=0.0; genuine "nothing gold and nothing predicted" still scores None; perfect match still scores 1.0. - [High] LightRAG cost undercounted. _counting_llm_wrapper now tokenizes every message in history_messages (each billed as real provider input, same as prompt/system_prompt), not just the two it previously counted. Verified: a call with two history messages now counts ~5x the prompt tokens of an otherwise-identical call with none. - [High] Backends could target different Memgraph instances. New _pin_memgraph_env() sets MEMGRAPH_URL *and* MEMGRAPH_URI (+ empty MEMGRAPH_USER/MEMGRAPH_USERNAME/MEMGRAPH_PASSWORD defaults) unconditionally before any backend is built, instead of relying on lightrag_memgraph.core's _bridge_lightrag_env_names(), which only mirrors MEMGRAPH_URL onto MEMGRAPH_URI when MEMGRAPH_URI isn't already set -- meaning a stale ambient MEMGRAPH_URI would previously have silently won over --memgraph-url for LightRAG's storage backend while this eval's own client used the correct one. Verified: a pre-set stale MEMGRAPH_URI is now overridden. - [Medium] Gold span scoring discards spans. Not changed to offset-based matching (would reintroduce the exact cross-backend brittleness normalized- text matching was chosen to avoid -- LightRAG's LLM-normalized text doesn't char-align with source spans the way GLiNER2's does). Instead documented explicitly and precisely: GoldEntity/GoldRelation's docstrings now state start/end exist only for build_gold_corpus.py's own verbatim-text assertion and are never read back for scoring, and that matching is by unique normalized text per chunk (a set, collapsing repeated identical mentions), not span position or per-occurrence count. print_report() prints the same caveat. Verified: ruff check/format clean, ty check clean, 71 unstructured2graph unit tests pass (was 70 -- LightRAGBackend gained one afinalize() test in #330). Ran the full rewritten script end to end against both backends (real GLiNER2 model + real LightRAG/OpenAI) on a live dedicated Memgraph -- both produce full reports with the new cost/token-breakdown and gold-span-caveat lines. Independently verified each of the four Spec fixes in isolation (PRF1.f1's three cases, Cypher-identifier rejection, history_messages token delta, MEMGRAPH_URI override) rather than relying on the end-to-end run alone to prove them.
antejavor
force-pushed
the
unstructured2graph-gliner2-vs-lightrag-eval
branch
from
September 16, 2026 07:16
95de9a4 to
d0d00db
Compare
antejavor
added a commit
that referenced
this pull request
Sep 16, 2026
main's GLiNER2-backend work (#332) renamed every loaders.py function's lightrag_wrapper: MemgraphLightRAGWrapper param to extraction_backend: ExtractionBackend, and dropped the now-unused MemgraphLightRAGWrapper import along with it. enqueue_texts and process_enqueued_and_finalize didn't exist on main yet, so the merge left them calling the old MemgraphLightRAGWrapper-typed _resolve_entity_workspace signature that no longer exists -- an undefined name at lint time, and a broken .workspace_label lookup on a raw wrapper at runtime. Both functions drive LightRAG's document queue directly (not part of the ExtractionBackend Protocol, which only exposes aingest_chunk), so they stay typed on MemgraphLightRAGWrapper rather than joining the Protocol-based rename. process_enqueued_and_finalize now resolves its own entity workspace inline instead of through the incompatible shared helper. Verified: ruff check/format clean, full-workspace ty check clean, unstructured2graph 98 passed/9 skipped, sessions-graph 69 passed/1 skipped, context-graph-eval 175 passed/5 skipped (all skips are the expected OPENAI_API_KEY/gliner2/sample-data gates).
Merged
5 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #330 (base =
unstructured2graph-gliner2-backend, notmain-- retarget once that merges).A direct, deterministic comparison of the two
ExtractionBackendimplementations from #330 -- given the same text, how doLightRAGBackendandGLiNER2Backendcompare on entity/relation precision/recall/F1, latency, and cost?This is deliberately not built on
context-graph-eval's existing harness: that eval's Golden corpus is Q&A pairs with no entity/relation-level gold labels, and itsreconcile_batch()is hardwired tolightrag_wrapper(sessions-graph isn't backend-pluggable). No LLM judge either -- entity/relation matching against gold spans is a countable metric, not a quality judgment, so this is plain deterministic scoring.unstructured2graph/evals/build_gold_corpus.py-- generatesgold_corpus.jsonl(41 chunks) from reviewable Python: 36 sampled fromcontext-graph/eval's LongMemEval corpus (real conversational text -- though it turns out to carry almost no explicit named-entity relations), 5 authored specifically to exerciseworks_for/founded/located_inrelation scoringunstructured2graph/evals/extraction_quality.py-- runs both backends through the realfrom_texts()pipeline into a dedicated Memgraph instance, reads back what each wrote, scores type-agnostic + type-sensitive entity P/R/F1 and relation P/R/F1 (GLiNER2 only), tracks latency and LightRAG's LLM call/token costunstructured2graph/evals/README.md-- how to run it, what each column means, and a caveat found by actually running it live: gold marks only salient entities, not everything the ontology's broad categories (Concept,Event,Artifact...) could arguably cover, so precision numbers are a floor, not a measured ceilingDeliberately out of scope: wiring
GLiNER2Backendintosessions_graph.reconcile_session/context_graph_eval.reconcile_batchso the existing Tier 1run/comparemachinery could measure end-to-end recall coverage under each backend. Real, separate work -- touchessessions-graph's public API and the Episode-summary LLM call is unrelated to which extraction backend is used. Worth a follow-up once this direct comparison gives a first read.Test plan
ruff check/ruff format --checkcleanty checkclean (worktree-local venv viauv sync --all-packages --all-extras)context-graph-eval's own 7689) with the real GLiNER2 model: entity recall 93.9%, relation recall 90.9%, 100% ontology conformance -- numbers make sense given the documented gold-corpus caveat (partial, salient-entity-only annotation deflates precision)OPENAI_API_KEY, not available in this environment -- run--backend lightrag(or--backend both) yourself before trusting any LightRAG-vs-GLiNER2 delta