Skip to content

unstructured2graph: extraction-quality eval, GLiNER2 vs LightRAG - #332

Merged
antejavor merged 5 commits into
mainfrom
unstructured2graph-gliner2-vs-lightrag-eval
Sep 16, 2026
Merged

antejavor merged 5 commits into
mainfrom
unstructured2graph-gliner2-vs-lightrag-eval

Conversation

@antejavor

@antejavor antejavor commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Stacked on #330 (base = unstructured2graph-gliner2-backend, not main -- retarget once that merges).

A direct, deterministic comparison of the two ExtractionBackend implementations from #330 -- given the same text, how do LightRAGBackend and GLiNER2Backend compare on entity/relation precision/recall/F1, latency, and cost?

This is deliberately not built on context-graph-eval's existing harness: that eval's Golden corpus is Q&A pairs with no entity/relation-level gold labels, and its reconcile_batch() is hardwired to lightrag_wrapper (sessions-graph isn't backend-pluggable). No LLM judge either -- entity/relation matching against gold spans is a countable metric, not a quality judgment, so this is plain deterministic scoring.

  • unstructured2graph/evals/build_gold_corpus.py -- generates gold_corpus.jsonl (41 chunks) from reviewable Python: 36 sampled from context-graph/eval's LongMemEval corpus (real conversational text -- though it turns out to carry almost no explicit named-entity relations), 5 authored specifically to exercise works_for/founded/located_in relation scoring
  • unstructured2graph/evals/extraction_quality.py -- runs both backends through the real from_texts() pipeline into a dedicated Memgraph instance, reads back what each wrote, scores type-agnostic + type-sensitive entity P/R/F1 and relation P/R/F1 (GLiNER2 only), tracks latency and LightRAG's LLM call/token cost
  • unstructured2graph/evals/README.md -- how to run it, what each column means, and a caveat found by actually running it live: gold marks only salient entities, not everything the ontology's broad categories (Concept, Event, Artifact...) could arguably cover, so precision numbers are a floor, not a measured ceiling

Deliberately out of scope: wiring GLiNER2Backend into sessions_graph.reconcile_session/context_graph_eval.reconcile_batch so the existing Tier 1 run/compare machinery could measure end-to-end recall coverage under each backend. Real, separate work -- touches sessions-graph's public API and the Episode-summary LLM call is unrelated to which extraction backend is used. Worth a follow-up once this direct comparison gives a first read.

Test plan

  • ruff check / ruff format --check clean
  • ty check clean (worktree-local venv via uv sync --all-packages --all-extras)
  • Ran the eval for real against a dedicated local Memgraph (port 7691, isolated from context-graph-eval's own 7689) with the real GLiNER2 model: entity recall 93.9%, relation recall 90.9%, 100% ontology conformance -- numbers make sense given the documented gold-corpus caveat (partial, salient-entity-only annotation deflates precision)
  • LightRAG side needs OPENAI_API_KEY, not available in this environment -- run --backend lightrag (or --backend both) yourself before trusting any LightRAG-vs-GLiNER2 delta

antejavor added a commit that referenced this pull request Sep 15, 2026
Prompted by review feedback on #332 (the extraction-quality eval), which
was reaching through backend.wrapper.afinalize() to finalize a
LightRAGBackend between repeated runs -- an ExtractionBackend consumer
shouldn't need to know LightRAGBackend wraps a MemgraphLightRAGWrapper at
all. Exposed as a LightRAGBackend-specific method, not added to the
ExtractionBackend protocol: GLiNER2Backend has no equivalent persistent
resource needing teardown, so this isn't a contract every backend must
satisfy.
antejavor added a commit that referenced this pull request Sep 15, 2026
Standards, hard violations:

- Cypher injection risk. _read_entities/_read_ontology_conformance/
  _read_relations now validate workspace/text_property/relation_types via
  _require_valid_identifier (reused from unstructured2graph.memgraph, same
  helper #330's upsert_typed_relationships review fix added) before
  interpolating them, instead of trusting the caller.

- Incomplete docstrings. Added Attributes/Args/Returns/Raises to GoldEntity,
  GoldRelation, GoldRecord, PRF1 (+ its properties), BackendReport,
  run_backend(), run(), _run_repeated(), print_report(), BackendSpec.

Standards, judgment calls:

- afinalize() reach-through. Eliminated by #330's new
  LightRAGBackend.afinalize() passthrough (merged in) -- run() now calls
  backend.afinalize() instead of backend.wrapper.afinalize().

- Duplicated repeat-loop orchestration. LightRAG's and GLiNER2's near-
  identical "build, run_backend, maybe finalize, print" loops collapsed into
  one shared _run_repeated(), parameterized by a BackendBuilder callable
  (LightRAG rebuilds + resets counters fresh per repeat; GLiNER2 reuses one
  loaded model via _reuse_gliner2_builder) and an optional finalize hook.

- Configuration as loose strings. name/text_property/relation_types/
  llm_model bundled into a new frozen BackendSpec (LIGHTRAG_SPEC/
  GLINER2_SPEC) instead of traveling run_backend()'s call chain as four
  separate parameters.

Spec:

- [High] Zero F1 reported as n/a. PRF1.f1's `if not p or not r` treated a
  real, computed 0.0 the same as "undefined" (None), since 0.0 is falsy.
  Fixed to `if p is None or r is None`, with an explicit 0.0 for the
  precision=recall=0 case (would otherwise divide by zero). Verified: zero-
  overlap now scores f1=0.0; genuine "nothing gold and nothing predicted"
  still scores None; perfect match still scores 1.0.

- [High] LightRAG cost undercounted. _counting_llm_wrapper now tokenizes
  every message in history_messages (each billed as real provider input,
  same as prompt/system_prompt), not just the two it previously counted.
  Verified: a call with two history messages now counts ~5x the prompt
  tokens of an otherwise-identical call with none.

- [High] Backends could target different Memgraph instances. New
  _pin_memgraph_env() sets MEMGRAPH_URL *and* MEMGRAPH_URI (+ empty
  MEMGRAPH_USER/MEMGRAPH_USERNAME/MEMGRAPH_PASSWORD defaults) unconditionally
  before any backend is built, instead of relying on
  lightrag_memgraph.core's _bridge_lightrag_env_names(), which only mirrors
  MEMGRAPH_URL onto MEMGRAPH_URI when MEMGRAPH_URI isn't already set --
  meaning a stale ambient MEMGRAPH_URI would previously have silently won
  over --memgraph-url for LightRAG's storage backend while this eval's own
  client used the correct one. Verified: a pre-set stale MEMGRAPH_URI is now
  overridden.

- [Medium] Gold span scoring discards spans. Not changed to offset-based
  matching (would reintroduce the exact cross-backend brittleness normalized-
  text matching was chosen to avoid -- LightRAG's LLM-normalized text
  doesn't char-align with source spans the way GLiNER2's does). Instead
  documented explicitly and precisely: GoldEntity/GoldRelation's docstrings
  now state start/end exist only for build_gold_corpus.py's own verbatim-text
  assertion and are never read back for scoring, and that matching is by
  unique normalized text per chunk (a set, collapsing repeated identical
  mentions), not span position or per-occurrence count. print_report() prints
  the same caveat.

Verified: ruff check/format clean, ty check clean, 71 unstructured2graph
unit tests pass (was 70 -- LightRAGBackend gained one afinalize() test in
#330). Ran the full rewritten script end to end against both backends (real
GLiNER2 model + real LightRAG/OpenAI) on a live dedicated Memgraph -- both
produce full reports with the new cost/token-breakdown and gold-span-caveat
lines. Independently verified each of the four Spec fixes in isolation
(PRF1.f1's three cases, Cypher-identifier rejection, history_messages token
delta, MEMGRAPH_URI override) rather than relying on the end-to-end run alone
to prove them.
@antejavor
antejavor added this pull request to stack #334 September 16, 2026 07:14
Base automatically changed from unstructured2graph-gliner2-backend to main September 16, 2026 07:16
antejavor added a commit that referenced this pull request Sep 16, 2026
* unstructured2graph: add pluggable extraction backends, GLiNER2Backend

Entity/relation extraction is now behind an ExtractionBackend protocol
(workspace_label + async aingest_chunk) instead of being hardwired to
LightRAG, so from_texts()/from_unstructured() can drive different
extractors through one seam. LightRAGBackend adapts the existing
MemgraphLightRAGWrapper; GLiNER2Backend is a new local, LLM-free
extractor using GLiNER2 (https://github.com/fastino-ai/GLiNER2) for
joint entity+relation extraction with no API key or network cost.

- extraction_backend.py: ExtractionBackend protocol + LightRAGBackend
- gliner2_backend.py: GLiNER2Backend, entity identity scoped to
  (chunk, entity_type, normalized text) since GLiNER2 does no
  cross-chunk coreference, relations written as per-label Cypher edge
  types (e.g. :works_for) drawn from Ontology.relation_types
- ontology.py: Ontology gains an optional relation_types field, read
  directly by GLiNER2Backend as its extraction schema
- memgraph.py: new upsert_typed_relationships() helper
- loaders.py: lightrag_wrapper param renamed to extraction_backend
  (pre-1.0 package, no compat shim; the one external call site in
  sessions-graph/core.py is updated)
- pyproject.toml: new optional `gliner2` extra

Also lowers the workspace's transformers floor from >=5.0.0rc3 to
>=4.57.6 (latest patched 4.x): gliner2[local] hard-pins transformers<5
on every published release, which is incompatible with >=5.0.0rc3 (the
fix for CVE-2026-1839, an RCE in Trainer's checkpoint loading, with no
4.x backport). Confirmed via `uv lock`/`uv sync --all-extras` that no
version satisfies both constraints -- this tradeoff was discussed and
confirmed explicitly before making the change.

Verified against a live Memgraph with the real GLiNER2 model (not just
the fake-model unit tests): clean entity typing and a real typed
relation edge.

* unstructured2graph: address review feedback on #330

Standards, blocking:

- Security floor weakened. Reverted -- transformers>=5.0.0rc3 restored in
  root pyproject.toml. gliner2 is no longer a pyproject.toml extra of
  unstructured2graph at all (gliner2[local] hard-pins transformers<5, which
  is unconditionally incompatible with the floor under uv's workspace-wide
  resolution -- confirmed again via `uv lock`). Isolated instead: manual
  `pip install 'gliner2[local]>=2.0.0'` in your own environment, documented
  in gliner2_backend.py's module docstring, unstructured2graph/README.md,
  the example, and test_e2e_gliner2.py (whose skip in CI is now permanent
  and documented as such, not a coverage gap). uv.lock regenerated --
  gliner2/peft removed, transformers back to 5.17.0.

- Cypher injection risk. upsert_typed_relationships() now validates
  node_label, match_key, every relation type, and every edge-property key
  via a new shared _require_valid_identifier() (memgraph.py), not just
  relation_type. GLiNER2Backend.__init__ now validates `workspace` itself
  (raises ValueError) -- the actual source the review flagged as reaching
  these queries arbitrarily; fixing it there closes every downstream Cypher
  use of workspace (create_nodes_from_list, connect_chunks_to_entities,
  promote_*_to_labels, upsert_typed_relationships) in one place rather than
  requiring the same check duplicated across pre-existing functions this PR
  didn't otherwise touch. New tests for both layers (test_memgraph.py,
  test_e2e.py, test_gliner2_backend.py).

- AGENTS.md stale. Updated the unstructured2graph row to describe the
  pluggable ExtractionBackend architecture.

- Public API documentation incomplete. Added Args/Raises to RelationType
  (+ EntityType, for consistency), GLiNER2Backend.__init__,
  _resolve_entity_workspace, and upsert_typed_relationships.

Standards, judgment calls:

- Typed records. GLiNER2Backend._extract_sync now returns
  list[ExtractedEntity]/list[ExtractedRelation] (frozen dataclasses) instead
  of list[dict[str, Any]].

- Four-object reach for workspace. Added MemgraphLightRAGWrapper.workspace
  (lightrag_memgraph/core.py) so LightRAGBackend.workspace_label is one
  attribute access instead of reaching through get_lightrag().
  chunk_entity_relation_graph. New tests in lightrag-memgraph's own suite.

- GLiNER2 e2e never runs in CI. Now explained rather than left implicit:
  test_e2e_gliner2.py's docstring states this is permanent, not pending,
  given the security-floor isolation above.

Spec:

- [P1] Root quick start crashes. README.md's unstructured2graph quick start
  now uses extraction_backend=LightRAGBackend(lightrag).

- [P2] Advertised joint extraction wasn't joint. _extract_sync now makes one
  model.create_schema().entities(...).relations(...) + model.extract() call
  instead of two separate extract_entities()/extract_relations() calls --
  confirmed against a real gliner2==2.0.0 install that this returns
  {"entities": ..., "relation_extraction": ...} in one shape. Verified live
  this fixes the specific failure mode raised (independent inference passes
  disagreeing on a span for the same entity between the two calls) for
  entities the model recognizes in both its entity and relation output.
  Also verified live that it does NOT make every relation matchable --
  GLiNER2's relation task can name a head/tail its entity task never
  surfaces under the configured schema at all, joint call or not -- and
  corrected _extract_sync's docstring to say so accurately rather than
  overclaim. aingest_chunk() still correctly skips-and-logs that case: no
  extracted entity_type exists to write such an endpoint under.

- [P2] Workspace override silently broke linking. _resolve_entity_workspace
  now raises ValueError when entity_workspace is given and doesn't match
  extraction_backend.workspace_label, instead of silently using the
  mismatched value. Updated/added tests in test_loaders.py.

Verified: ruff check/format clean; ty check clean (gliner2's own unresolved
lazy import is a narrow, commented `ty: ignore[unresolved-import]` -- see
gliner2_backend.py, since it's permanently outside the managed dependency
graph now, not a normal optional extra ty could resolve via --all-extras).
86 unstructured2graph tests (was 62) + 37 lightrag-memgraph tests (was 35)
+ 55 sessions-graph tests all pass, including real-model GLiNER2 e2e tests
run against a manually-installed gliner2 (uv pip install --no-config, to
simulate the documented manual-install path without touching the workspace
lock) and a live dedicated Memgraph.

* unstructured2graph: add LightRAGBackend.afinalize() passthrough

Prompted by review feedback on #332 (the extraction-quality eval), which
was reaching through backend.wrapper.afinalize() to finalize a
LightRAGBackend between repeated runs -- an ExtractionBackend consumer
shouldn't need to know LightRAGBackend wraps a MemgraphLightRAGWrapper at
all. Exposed as a LightRAGBackend-specific method, not added to the
ExtractionBackend protocol: GLiNER2Backend has no equivalent persistent
resource needing teardown, so this isn't a contract every backend must
satisfy.
Deterministic entity/relation-level comparison of the two ExtractionBackend
implementations (#330), separate from context-graph-eval's end-to-end memory
recall loop -- that harness's Golden corpus has no entity/relation gold
labels and reconcile_batch() is hardwired to lightrag_wrapper, neither of
which fits this question. No LLM judge: precision/recall/F1 against gold
spans is a countable metric, not a quality judgment.

- evals/build_gold_corpus.py: generates gold_corpus.jsonl (41 chunks) from
  reviewable Python -- 36 sampled from context-graph/eval's LongMemEval
  corpus (real text, but carries almost no named-entity relations), 5
  authored specifically to exercise works_for/founded/located_in relation
  scoring
- evals/extraction_quality.py: runs both backends through the real
  from_texts() pipeline against a dedicated Memgraph instance, reads back
  what each wrote, scores type-agnostic and type-sensitive entity P/R/F1
  plus relation P/R/F1 (GLiNER2 only -- LightRAG has no relation_types
  concept), tracks latency and (for LightRAG) LLM call/token cost
- evals/README.md: how to run it, what each column means, and a documented
  gold-corpus caveat found by actually running it -- gold marks only
  salient entities, not everything the ontology's broad categories
  (Concept, Event, Artifact...) could arguably cover, so precision is a
  floor, not a measured ceiling

Verified live: GLiNER2 end-to-end against a real dedicated Memgraph instance
(port 7691, isolated from context-graph-eval's own 7689). LightRAG needs
OPENAI_API_KEY, not available in this environment -- run
--backend lightrag yourself before trusting any LightRAG-vs-GLiNER2 delta.
unstructured2graph (and context-graph-eval) had no .env loading anywhere,
unlike integrations/langchain-memgraph and integrations/mcp-memgraph, which
already do a bare load_dotenv() in their own tests -- surfaced concretely
when the extraction-quality eval's lightrag backend silently skipped because
the repo-root .env was never read into the process environment.

main() now does the same bare load_dotenv() call (soft dependency via the
test extra), which finds the repo-root .env through python-dotenv's own
upward search regardless of whether the script runs from unstructured2graph/
or the repo root. Verified with OPENAI_API_KEY unset in the shell and only
present in .env: both backends now run end to end.

First real side-by-side numbers (41-chunk gold corpus, single run each):
lightrag  ent F1 65.3% (P 49.4/R 96.3), 4.83s/chunk median, 83 calls/167k tokens
gliner2   ent F1 52.0% (P 36.0/R 93.9), 0.27s/chunk median, 0 calls, $0
(relation scoring is gliner2-only; ontology conformance 93.2% vs 100%)
Both precision numbers are understated by the same gold-corpus partial-
annotation caveat evals/README.md already documents.
Answers the actual question a token count doesn't: how much did this run
cost. MODEL_PRICING_PER_1M_TOKENS is a small, pinned (point-in-time, not
fetched live) per-model price table; estimated_cost_usd() multiplies it by
the approximate prompt/completion token split already tracked. A model
missing from the table reports "unpriced", never a silent $0 -- the
distinction matters since GLiNER2 genuinely is $0 (n/a, no LLM calls at all)
and an unpriced LightRAG model is a gap in the table, not a finding.

Repeat run against the same 41-chunk gold corpus, same OPENAI_API_KEY-via-
.env path verified last commit: 83 calls, 152k prompt + 16k completion
tokens, ~$0.03 at gpt-4o-mini pricing. Entity F1 moved from 65.3% to 57.7%
between the two runs (both LightRAG, nothing else changed) -- exactly the
LLM-extraction run-to-run variance the README already warns about, now
visible in the wild rather than theoretical.
Follows #330's review-feedback fix: gliner2 is no longer a pyproject.toml
extra (isolated from the workspace's transformers>=5.0.0rc3 floor instead
of weakening it). Updates this eval's own README/module docstring off the
now-removed `pip install unstructured2graph[gliner2]` to the manual
`pip install 'gliner2[local]>=2.0.0'` path.
Standards, hard violations:

- Cypher injection risk. _read_entities/_read_ontology_conformance/
  _read_relations now validate workspace/text_property/relation_types via
  _require_valid_identifier (reused from unstructured2graph.memgraph, same
  helper #330's upsert_typed_relationships review fix added) before
  interpolating them, instead of trusting the caller.

- Incomplete docstrings. Added Attributes/Args/Returns/Raises to GoldEntity,
  GoldRelation, GoldRecord, PRF1 (+ its properties), BackendReport,
  run_backend(), run(), _run_repeated(), print_report(), BackendSpec.

Standards, judgment calls:

- afinalize() reach-through. Eliminated by #330's new
  LightRAGBackend.afinalize() passthrough (merged in) -- run() now calls
  backend.afinalize() instead of backend.wrapper.afinalize().

- Duplicated repeat-loop orchestration. LightRAG's and GLiNER2's near-
  identical "build, run_backend, maybe finalize, print" loops collapsed into
  one shared _run_repeated(), parameterized by a BackendBuilder callable
  (LightRAG rebuilds + resets counters fresh per repeat; GLiNER2 reuses one
  loaded model via _reuse_gliner2_builder) and an optional finalize hook.

- Configuration as loose strings. name/text_property/relation_types/
  llm_model bundled into a new frozen BackendSpec (LIGHTRAG_SPEC/
  GLINER2_SPEC) instead of traveling run_backend()'s call chain as four
  separate parameters.

Spec:

- [High] Zero F1 reported as n/a. PRF1.f1's `if not p or not r` treated a
  real, computed 0.0 the same as "undefined" (None), since 0.0 is falsy.
  Fixed to `if p is None or r is None`, with an explicit 0.0 for the
  precision=recall=0 case (would otherwise divide by zero). Verified: zero-
  overlap now scores f1=0.0; genuine "nothing gold and nothing predicted"
  still scores None; perfect match still scores 1.0.

- [High] LightRAG cost undercounted. _counting_llm_wrapper now tokenizes
  every message in history_messages (each billed as real provider input,
  same as prompt/system_prompt), not just the two it previously counted.
  Verified: a call with two history messages now counts ~5x the prompt
  tokens of an otherwise-identical call with none.

- [High] Backends could target different Memgraph instances. New
  _pin_memgraph_env() sets MEMGRAPH_URL *and* MEMGRAPH_URI (+ empty
  MEMGRAPH_USER/MEMGRAPH_USERNAME/MEMGRAPH_PASSWORD defaults) unconditionally
  before any backend is built, instead of relying on
  lightrag_memgraph.core's _bridge_lightrag_env_names(), which only mirrors
  MEMGRAPH_URL onto MEMGRAPH_URI when MEMGRAPH_URI isn't already set --
  meaning a stale ambient MEMGRAPH_URI would previously have silently won
  over --memgraph-url for LightRAG's storage backend while this eval's own
  client used the correct one. Verified: a pre-set stale MEMGRAPH_URI is now
  overridden.

- [Medium] Gold span scoring discards spans. Not changed to offset-based
  matching (would reintroduce the exact cross-backend brittleness normalized-
  text matching was chosen to avoid -- LightRAG's LLM-normalized text
  doesn't char-align with source spans the way GLiNER2's does). Instead
  documented explicitly and precisely: GoldEntity/GoldRelation's docstrings
  now state start/end exist only for build_gold_corpus.py's own verbatim-text
  assertion and are never read back for scoring, and that matching is by
  unique normalized text per chunk (a set, collapsing repeated identical
  mentions), not span position or per-occurrence count. print_report() prints
  the same caveat.

Verified: ruff check/format clean, ty check clean, 71 unstructured2graph
unit tests pass (was 70 -- LightRAGBackend gained one afinalize() test in
#330). Ran the full rewritten script end to end against both backends (real
GLiNER2 model + real LightRAG/OpenAI) on a live dedicated Memgraph -- both
produce full reports with the new cost/token-breakdown and gold-span-caveat
lines. Independently verified each of the four Spec fixes in isolation
(PRF1.f1's three cases, Cypher-identifier rejection, history_messages token
delta, MEMGRAPH_URI override) rather than relying on the end-to-end run alone
to prove them.
@antejavor
antejavor force-pushed the unstructured2graph-gliner2-vs-lightrag-eval branch from 95de9a4 to d0d00db Compare September 16, 2026 07:16
@antejavor
antejavor merged commit 605b86c into main Sep 16, 2026
18 checks passed
antejavor added a commit that referenced this pull request Sep 16, 2026
main's GLiNER2-backend work (#332) renamed every loaders.py function's
lightrag_wrapper: MemgraphLightRAGWrapper param to
extraction_backend: ExtractionBackend, and dropped the now-unused
MemgraphLightRAGWrapper import along with it. enqueue_texts and
process_enqueued_and_finalize didn't exist on main yet, so the merge
left them calling the old MemgraphLightRAGWrapper-typed
_resolve_entity_workspace signature that no longer exists -- an
undefined name at lint time, and a broken .workspace_label lookup on a
raw wrapper at runtime.

Both functions drive LightRAG's document queue directly (not part of
the ExtractionBackend Protocol, which only exposes aingest_chunk), so
they stay typed on MemgraphLightRAGWrapper rather than joining the
Protocol-based rename. process_enqueued_and_finalize now resolves its
own entity workspace inline instead of through the incompatible shared
helper.

Verified: ruff check/format clean, full-workspace ty check clean,
unstructured2graph 98 passed/9 skipped, sessions-graph 69 passed/1
skipped, context-graph-eval 175 passed/5 skipped (all skips are the
expected OPENAI_API_KEY/gliner2/sample-data gates).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant