Skip to content

unstructured2graph: pluggable extraction backends + GLiNER2 - #330

Merged
antejavor merged 4 commits into
mainfrom
unstructured2graph-gliner2-backend
Sep 16, 2026
Merged

antejavor merged 4 commits into
mainfrom
unstructured2graph-gliner2-backend

Conversation

@antejavor

@antejavor antejavor commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Entity/relation extraction in unstructured2graph is now behind an ExtractionBackend protocol (workspace_label + async aingest_chunk) instead of being hardwired to LightRAG. LightRAGBackend adapts the existing MemgraphLightRAGWrapper; GLiNER2Backend is a new local, LLM-free extractor using GLiNER2 — no API key, no network cost, joint entity+relation extraction in one pass.
  • Ontology gains an optional relation_types field (read directly by GLiNER2Backend as its extraction schema); relations are written as per-label Cypher edge types (e.g. :works_for) rather than one generic edge type.
  • from_texts()/from_unstructured()'s lightrag_wrapper param is renamed to extraction_backend — pre-1.0 package, no compat shim. The one external call site (sessions-graph/core.py) is updated; its own public reconcile_session(lightrag_wrapper=...) API is unchanged.
  • New optional gliner2 extra: pip install unstructured2graph[gliner2].
  • Lowers the workspace's transformers floor from >=5.0.0rc3 to >=4.57.6 (latest patched 4.x release). gliner2[local] hard-pins transformers<5 on every published version, which conflicts with >=5.0.0rc3 (the fix for CVE-2026-1839, an RCE in Trainer's checkpoint loading — no 4.x backport exists). This tradeoff was discussed and explicitly confirmed before making the change; see the inline comment in root pyproject.toml.

Test plan

  • ruff check . / ruff format --check . clean across the repo
  • ty check . clean across the whole workspace (with gliner2 installed via uv sync --all-packages --all-extras)
  • uv lock / uv sync --all-packages --all-extras succeed
  • Unit tests: 62 existing unstructured2graph tests + 13 new (test_extraction_backend.py, test_gliner2_backend.py, fake-model, no real gliner2 needed) all pass
  • sessions-graph test suite (53 tests) passes after the one call-site update
  • Ran the real GLiNER2 model against a live local Memgraph (test_e2e_gliner2.py + a manual smoke check): clean entity typing (Alice Johnson→person, Acme Corp→organization) and a real :works_for typed relation edge

Entity/relation extraction is now behind an ExtractionBackend protocol
(workspace_label + async aingest_chunk) instead of being hardwired to
LightRAG, so from_texts()/from_unstructured() can drive different
extractors through one seam. LightRAGBackend adapts the existing
MemgraphLightRAGWrapper; GLiNER2Backend is a new local, LLM-free
extractor using GLiNER2 (https://github.com/fastino-ai/GLiNER2) for
joint entity+relation extraction with no API key or network cost.

- extraction_backend.py: ExtractionBackend protocol + LightRAGBackend
- gliner2_backend.py: GLiNER2Backend, entity identity scoped to
  (chunk, entity_type, normalized text) since GLiNER2 does no
  cross-chunk coreference, relations written as per-label Cypher edge
  types (e.g. :works_for) drawn from Ontology.relation_types
- ontology.py: Ontology gains an optional relation_types field, read
  directly by GLiNER2Backend as its extraction schema
- memgraph.py: new upsert_typed_relationships() helper
- loaders.py: lightrag_wrapper param renamed to extraction_backend
  (pre-1.0 package, no compat shim; the one external call site in
  sessions-graph/core.py is updated)
- pyproject.toml: new optional `gliner2` extra

Also lowers the workspace's transformers floor from >=5.0.0rc3 to
>=4.57.6 (latest patched 4.x): gliner2[local] hard-pins transformers<5
on every published release, which is incompatible with >=5.0.0rc3 (the
fix for CVE-2026-1839, an RCE in Trainer's checkpoint loading, with no
4.x backport). Confirmed via `uv lock`/`uv sync --all-extras` that no
version satisfies both constraints -- this tradeoff was discussed and
confirmed explicitly before making the change.

Verified against a live Memgraph with the real GLiNER2 model (not just
the fake-model unit tests): clean entity typing and a real typed
relation edge.
Standards, blocking:

- Security floor weakened. Reverted -- transformers>=5.0.0rc3 restored in
  root pyproject.toml. gliner2 is no longer a pyproject.toml extra of
  unstructured2graph at all (gliner2[local] hard-pins transformers<5, which
  is unconditionally incompatible with the floor under uv's workspace-wide
  resolution -- confirmed again via `uv lock`). Isolated instead: manual
  `pip install 'gliner2[local]>=2.0.0'` in your own environment, documented
  in gliner2_backend.py's module docstring, unstructured2graph/README.md,
  the example, and test_e2e_gliner2.py (whose skip in CI is now permanent
  and documented as such, not a coverage gap). uv.lock regenerated --
  gliner2/peft removed, transformers back to 5.17.0.

- Cypher injection risk. upsert_typed_relationships() now validates
  node_label, match_key, every relation type, and every edge-property key
  via a new shared _require_valid_identifier() (memgraph.py), not just
  relation_type. GLiNER2Backend.__init__ now validates `workspace` itself
  (raises ValueError) -- the actual source the review flagged as reaching
  these queries arbitrarily; fixing it there closes every downstream Cypher
  use of workspace (create_nodes_from_list, connect_chunks_to_entities,
  promote_*_to_labels, upsert_typed_relationships) in one place rather than
  requiring the same check duplicated across pre-existing functions this PR
  didn't otherwise touch. New tests for both layers (test_memgraph.py,
  test_e2e.py, test_gliner2_backend.py).

- AGENTS.md stale. Updated the unstructured2graph row to describe the
  pluggable ExtractionBackend architecture.

- Public API documentation incomplete. Added Args/Raises to RelationType
  (+ EntityType, for consistency), GLiNER2Backend.__init__,
  _resolve_entity_workspace, and upsert_typed_relationships.

Standards, judgment calls:

- Typed records. GLiNER2Backend._extract_sync now returns
  list[ExtractedEntity]/list[ExtractedRelation] (frozen dataclasses) instead
  of list[dict[str, Any]].

- Four-object reach for workspace. Added MemgraphLightRAGWrapper.workspace
  (lightrag_memgraph/core.py) so LightRAGBackend.workspace_label is one
  attribute access instead of reaching through get_lightrag().
  chunk_entity_relation_graph. New tests in lightrag-memgraph's own suite.

- GLiNER2 e2e never runs in CI. Now explained rather than left implicit:
  test_e2e_gliner2.py's docstring states this is permanent, not pending,
  given the security-floor isolation above.

Spec:

- [P1] Root quick start crashes. README.md's unstructured2graph quick start
  now uses extraction_backend=LightRAGBackend(lightrag).

- [P2] Advertised joint extraction wasn't joint. _extract_sync now makes one
  model.create_schema().entities(...).relations(...) + model.extract() call
  instead of two separate extract_entities()/extract_relations() calls --
  confirmed against a real gliner2==2.0.0 install that this returns
  {"entities": ..., "relation_extraction": ...} in one shape. Verified live
  this fixes the specific failure mode raised (independent inference passes
  disagreeing on a span for the same entity between the two calls) for
  entities the model recognizes in both its entity and relation output.
  Also verified live that it does NOT make every relation matchable --
  GLiNER2's relation task can name a head/tail its entity task never
  surfaces under the configured schema at all, joint call or not -- and
  corrected _extract_sync's docstring to say so accurately rather than
  overclaim. aingest_chunk() still correctly skips-and-logs that case: no
  extracted entity_type exists to write such an endpoint under.

- [P2] Workspace override silently broke linking. _resolve_entity_workspace
  now raises ValueError when entity_workspace is given and doesn't match
  extraction_backend.workspace_label, instead of silently using the
  mismatched value. Updated/added tests in test_loaders.py.

Verified: ruff check/format clean; ty check clean (gliner2's own unresolved
lazy import is a narrow, commented `ty: ignore[unresolved-import]` -- see
gliner2_backend.py, since it's permanently outside the managed dependency
graph now, not a normal optional extra ty could resolve via --all-extras).
86 unstructured2graph tests (was 62) + 37 lightrag-memgraph tests (was 35)
+ 55 sessions-graph tests all pass, including real-model GLiNER2 e2e tests
run against a manually-installed gliner2 (uv pip install --no-config, to
simulate the documented manual-install path without touching the workspace
lock) and a live dedicated Memgraph.
antejavor added a commit that referenced this pull request Sep 15, 2026
Follows #330's review-feedback fix: gliner2 is no longer a pyproject.toml
extra (isolated from the workspace's transformers>=5.0.0rc3 floor instead
of weakening it). Updates this eval's own README/module docstring off the
now-removed `pip install unstructured2graph[gliner2]` to the manual
`pip install 'gliner2[local]>=2.0.0'` path.
Prompted by review feedback on #332 (the extraction-quality eval), which
was reaching through backend.wrapper.afinalize() to finalize a
LightRAGBackend between repeated runs -- an ExtractionBackend consumer
shouldn't need to know LightRAGBackend wraps a MemgraphLightRAGWrapper at
all. Exposed as a LightRAGBackend-specific method, not added to the
ExtractionBackend protocol: GLiNER2Backend has no equivalent persistent
resource needing teardown, so this isn't a contract every backend must
satisfy.
antejavor added a commit that referenced this pull request Sep 15, 2026
Standards, hard violations:

- Cypher injection risk. _read_entities/_read_ontology_conformance/
  _read_relations now validate workspace/text_property/relation_types via
  _require_valid_identifier (reused from unstructured2graph.memgraph, same
  helper #330's upsert_typed_relationships review fix added) before
  interpolating them, instead of trusting the caller.

- Incomplete docstrings. Added Attributes/Args/Returns/Raises to GoldEntity,
  GoldRelation, GoldRecord, PRF1 (+ its properties), BackendReport,
  run_backend(), run(), _run_repeated(), print_report(), BackendSpec.

Standards, judgment calls:

- afinalize() reach-through. Eliminated by #330's new
  LightRAGBackend.afinalize() passthrough (merged in) -- run() now calls
  backend.afinalize() instead of backend.wrapper.afinalize().

- Duplicated repeat-loop orchestration. LightRAG's and GLiNER2's near-
  identical "build, run_backend, maybe finalize, print" loops collapsed into
  one shared _run_repeated(), parameterized by a BackendBuilder callable
  (LightRAG rebuilds + resets counters fresh per repeat; GLiNER2 reuses one
  loaded model via _reuse_gliner2_builder) and an optional finalize hook.

- Configuration as loose strings. name/text_property/relation_types/
  llm_model bundled into a new frozen BackendSpec (LIGHTRAG_SPEC/
  GLINER2_SPEC) instead of traveling run_backend()'s call chain as four
  separate parameters.

Spec:

- [High] Zero F1 reported as n/a. PRF1.f1's `if not p or not r` treated a
  real, computed 0.0 the same as "undefined" (None), since 0.0 is falsy.
  Fixed to `if p is None or r is None`, with an explicit 0.0 for the
  precision=recall=0 case (would otherwise divide by zero). Verified: zero-
  overlap now scores f1=0.0; genuine "nothing gold and nothing predicted"
  still scores None; perfect match still scores 1.0.

- [High] LightRAG cost undercounted. _counting_llm_wrapper now tokenizes
  every message in history_messages (each billed as real provider input,
  same as prompt/system_prompt), not just the two it previously counted.
  Verified: a call with two history messages now counts ~5x the prompt
  tokens of an otherwise-identical call with none.

- [High] Backends could target different Memgraph instances. New
  _pin_memgraph_env() sets MEMGRAPH_URL *and* MEMGRAPH_URI (+ empty
  MEMGRAPH_USER/MEMGRAPH_USERNAME/MEMGRAPH_PASSWORD defaults) unconditionally
  before any backend is built, instead of relying on
  lightrag_memgraph.core's _bridge_lightrag_env_names(), which only mirrors
  MEMGRAPH_URL onto MEMGRAPH_URI when MEMGRAPH_URI isn't already set --
  meaning a stale ambient MEMGRAPH_URI would previously have silently won
  over --memgraph-url for LightRAG's storage backend while this eval's own
  client used the correct one. Verified: a pre-set stale MEMGRAPH_URI is now
  overridden.

- [Medium] Gold span scoring discards spans. Not changed to offset-based
  matching (would reintroduce the exact cross-backend brittleness normalized-
  text matching was chosen to avoid -- LightRAG's LLM-normalized text
  doesn't char-align with source spans the way GLiNER2's does). Instead
  documented explicitly and precisely: GoldEntity/GoldRelation's docstrings
  now state start/end exist only for build_gold_corpus.py's own verbatim-text
  assertion and are never read back for scoring, and that matching is by
  unique normalized text per chunk (a set, collapsing repeated identical
  mentions), not span position or per-occurrence count. print_report() prints
  the same caveat.

Verified: ruff check/format clean, ty check clean, 71 unstructured2graph
unit tests pass (was 70 -- LightRAGBackend gained one afinalize() test in
#330). Ran the full rewritten script end to end against both backends (real
GLiNER2 model + real LightRAG/OpenAI) on a live dedicated Memgraph -- both
produce full reports with the new cost/token-breakdown and gold-span-caveat
lines. Independently verified each of the four Spec fixes in isolation
(PRF1.f1's three cases, Cypher-identifier rejection, history_messages token
delta, MEMGRAPH_URI override) rather than relying on the end-to-end run alone
to prove them.
@antejavor
antejavor added this pull request to stack #334 September 16, 2026 07:14
@antejavor
antejavor merged commit 32239eb into main Sep 16, 2026
22 checks passed
antejavor added a commit that referenced this pull request Sep 16, 2026
Deterministic entity/relation-level comparison of the two ExtractionBackend
implementations (#330), separate from context-graph-eval's end-to-end memory
recall loop -- that harness's Golden corpus has no entity/relation gold
labels and reconcile_batch() is hardwired to lightrag_wrapper, neither of
which fits this question. No LLM judge: precision/recall/F1 against gold
spans is a countable metric, not a quality judgment.

- evals/build_gold_corpus.py: generates gold_corpus.jsonl (41 chunks) from
  reviewable Python -- 36 sampled from context-graph/eval's LongMemEval
  corpus (real text, but carries almost no named-entity relations), 5
  authored specifically to exercise works_for/founded/located_in relation
  scoring
- evals/extraction_quality.py: runs both backends through the real
  from_texts() pipeline against a dedicated Memgraph instance, reads back
  what each wrote, scores type-agnostic and type-sensitive entity P/R/F1
  plus relation P/R/F1 (GLiNER2 only -- LightRAG has no relation_types
  concept), tracks latency and (for LightRAG) LLM call/token cost
- evals/README.md: how to run it, what each column means, and a documented
  gold-corpus caveat found by actually running it -- gold marks only
  salient entities, not everything the ontology's broad categories
  (Concept, Event, Artifact...) could arguably cover, so precision is a
  floor, not a measured ceiling

Verified live: GLiNER2 end-to-end against a real dedicated Memgraph instance
(port 7691, isolated from context-graph-eval's own 7689). LightRAG needs
OPENAI_API_KEY, not available in this environment -- run
--backend lightrag yourself before trusting any LightRAG-vs-GLiNER2 delta.
antejavor added a commit that referenced this pull request Sep 16, 2026
Follows #330's review-feedback fix: gliner2 is no longer a pyproject.toml
extra (isolated from the workspace's transformers>=5.0.0rc3 floor instead
of weakening it). Updates this eval's own README/module docstring off the
now-removed `pip install unstructured2graph[gliner2]` to the manual
`pip install 'gliner2[local]>=2.0.0'` path.
antejavor added a commit that referenced this pull request Sep 16, 2026
Standards, hard violations:

- Cypher injection risk. _read_entities/_read_ontology_conformance/
  _read_relations now validate workspace/text_property/relation_types via
  _require_valid_identifier (reused from unstructured2graph.memgraph, same
  helper #330's upsert_typed_relationships review fix added) before
  interpolating them, instead of trusting the caller.

- Incomplete docstrings. Added Attributes/Args/Returns/Raises to GoldEntity,
  GoldRelation, GoldRecord, PRF1 (+ its properties), BackendReport,
  run_backend(), run(), _run_repeated(), print_report(), BackendSpec.

Standards, judgment calls:

- afinalize() reach-through. Eliminated by #330's new
  LightRAGBackend.afinalize() passthrough (merged in) -- run() now calls
  backend.afinalize() instead of backend.wrapper.afinalize().

- Duplicated repeat-loop orchestration. LightRAG's and GLiNER2's near-
  identical "build, run_backend, maybe finalize, print" loops collapsed into
  one shared _run_repeated(), parameterized by a BackendBuilder callable
  (LightRAG rebuilds + resets counters fresh per repeat; GLiNER2 reuses one
  loaded model via _reuse_gliner2_builder) and an optional finalize hook.

- Configuration as loose strings. name/text_property/relation_types/
  llm_model bundled into a new frozen BackendSpec (LIGHTRAG_SPEC/
  GLINER2_SPEC) instead of traveling run_backend()'s call chain as four
  separate parameters.

Spec:

- [High] Zero F1 reported as n/a. PRF1.f1's `if not p or not r` treated a
  real, computed 0.0 the same as "undefined" (None), since 0.0 is falsy.
  Fixed to `if p is None or r is None`, with an explicit 0.0 for the
  precision=recall=0 case (would otherwise divide by zero). Verified: zero-
  overlap now scores f1=0.0; genuine "nothing gold and nothing predicted"
  still scores None; perfect match still scores 1.0.

- [High] LightRAG cost undercounted. _counting_llm_wrapper now tokenizes
  every message in history_messages (each billed as real provider input,
  same as prompt/system_prompt), not just the two it previously counted.
  Verified: a call with two history messages now counts ~5x the prompt
  tokens of an otherwise-identical call with none.

- [High] Backends could target different Memgraph instances. New
  _pin_memgraph_env() sets MEMGRAPH_URL *and* MEMGRAPH_URI (+ empty
  MEMGRAPH_USER/MEMGRAPH_USERNAME/MEMGRAPH_PASSWORD defaults) unconditionally
  before any backend is built, instead of relying on
  lightrag_memgraph.core's _bridge_lightrag_env_names(), which only mirrors
  MEMGRAPH_URL onto MEMGRAPH_URI when MEMGRAPH_URI isn't already set --
  meaning a stale ambient MEMGRAPH_URI would previously have silently won
  over --memgraph-url for LightRAG's storage backend while this eval's own
  client used the correct one. Verified: a pre-set stale MEMGRAPH_URI is now
  overridden.

- [Medium] Gold span scoring discards spans. Not changed to offset-based
  matching (would reintroduce the exact cross-backend brittleness normalized-
  text matching was chosen to avoid -- LightRAG's LLM-normalized text
  doesn't char-align with source spans the way GLiNER2's does). Instead
  documented explicitly and precisely: GoldEntity/GoldRelation's docstrings
  now state start/end exist only for build_gold_corpus.py's own verbatim-text
  assertion and are never read back for scoring, and that matching is by
  unique normalized text per chunk (a set, collapsing repeated identical
  mentions), not span position or per-occurrence count. print_report() prints
  the same caveat.

Verified: ruff check/format clean, ty check clean, 71 unstructured2graph
unit tests pass (was 70 -- LightRAGBackend gained one afinalize() test in
#330). Ran the full rewritten script end to end against both backends (real
GLiNER2 model + real LightRAG/OpenAI) on a live dedicated Memgraph -- both
produce full reports with the new cost/token-breakdown and gold-span-caveat
lines. Independently verified each of the four Spec fixes in isolation
(PRF1.f1's three cases, Cypher-identifier rejection, history_messages token
delta, MEMGRAPH_URI override) rather than relying on the end-to-end run alone
to prove them.
antejavor added a commit that referenced this pull request Sep 16, 2026
* unstructured2graph: direct extraction-quality eval, GLiNER2 vs LightRAG

Deterministic entity/relation-level comparison of the two ExtractionBackend
implementations (#330), separate from context-graph-eval's end-to-end memory
recall loop -- that harness's Golden corpus has no entity/relation gold
labels and reconcile_batch() is hardwired to lightrag_wrapper, neither of
which fits this question. No LLM judge: precision/recall/F1 against gold
spans is a countable metric, not a quality judgment.

- evals/build_gold_corpus.py: generates gold_corpus.jsonl (41 chunks) from
  reviewable Python -- 36 sampled from context-graph/eval's LongMemEval
  corpus (real text, but carries almost no named-entity relations), 5
  authored specifically to exercise works_for/founded/located_in relation
  scoring
- evals/extraction_quality.py: runs both backends through the real
  from_texts() pipeline against a dedicated Memgraph instance, reads back
  what each wrote, scores type-agnostic and type-sensitive entity P/R/F1
  plus relation P/R/F1 (GLiNER2 only -- LightRAG has no relation_types
  concept), tracks latency and (for LightRAG) LLM call/token cost
- evals/README.md: how to run it, what each column means, and a documented
  gold-corpus caveat found by actually running it -- gold marks only
  salient entities, not everything the ontology's broad categories
  (Concept, Event, Artifact...) could arguably cover, so precision is a
  floor, not a measured ceiling

Verified live: GLiNER2 end-to-end against a real dedicated Memgraph instance
(port 7691, isolated from context-graph-eval's own 7689). LightRAG needs
OPENAI_API_KEY, not available in this environment -- run
--backend lightrag yourself before trusting any LightRAG-vs-GLiNER2 delta.

* unstructured2graph: auto-load .env for the extraction-quality eval

unstructured2graph (and context-graph-eval) had no .env loading anywhere,
unlike integrations/langchain-memgraph and integrations/mcp-memgraph, which
already do a bare load_dotenv() in their own tests -- surfaced concretely
when the extraction-quality eval's lightrag backend silently skipped because
the repo-root .env was never read into the process environment.

main() now does the same bare load_dotenv() call (soft dependency via the
test extra), which finds the repo-root .env through python-dotenv's own
upward search regardless of whether the script runs from unstructured2graph/
or the repo root. Verified with OPENAI_API_KEY unset in the shell and only
present in .env: both backends now run end to end.

First real side-by-side numbers (41-chunk gold corpus, single run each):
lightrag  ent F1 65.3% (P 49.4/R 96.3), 4.83s/chunk median, 83 calls/167k tokens
gliner2   ent F1 52.0% (P 36.0/R 93.9), 0.27s/chunk median, 0 calls, $0
(relation scoring is gliner2-only; ontology conformance 93.2% vs 100%)
Both precision numbers are understated by the same gold-corpus partial-
annotation caveat evals/README.md already documents.

* unstructured2graph: report LightRAG cost in USD, not just tokens

Answers the actual question a token count doesn't: how much did this run
cost. MODEL_PRICING_PER_1M_TOKENS is a small, pinned (point-in-time, not
fetched live) per-model price table; estimated_cost_usd() multiplies it by
the approximate prompt/completion token split already tracked. A model
missing from the table reports "unpriced", never a silent $0 -- the
distinction matters since GLiNER2 genuinely is $0 (n/a, no LLM calls at all)
and an unpriced LightRAG model is a gap in the table, not a finding.

Repeat run against the same 41-chunk gold corpus, same OPENAI_API_KEY-via-
.env path verified last commit: 83 calls, 152k prompt + 16k completion
tokens, ~$0.03 at gpt-4o-mini pricing. Entity F1 moved from 65.3% to 57.7%
between the two runs (both LightRAG, nothing else changed) -- exactly the
LLM-extraction run-to-run variance the README already warns about, now
visible in the wild rather than theoretical.

* unstructured2graph: update eval docs for gliner2's manual-install fix

Follows #330's review-feedback fix: gliner2 is no longer a pyproject.toml
extra (isolated from the workspace's transformers>=5.0.0rc3 floor instead
of weakening it). Updates this eval's own README/module docstring off the
now-removed `pip install unstructured2graph[gliner2]` to the manual
`pip install 'gliner2[local]>=2.0.0'` path.

* unstructured2graph: address review feedback on #332

Standards, hard violations:

- Cypher injection risk. _read_entities/_read_ontology_conformance/
  _read_relations now validate workspace/text_property/relation_types via
  _require_valid_identifier (reused from unstructured2graph.memgraph, same
  helper #330's upsert_typed_relationships review fix added) before
  interpolating them, instead of trusting the caller.

- Incomplete docstrings. Added Attributes/Args/Returns/Raises to GoldEntity,
  GoldRelation, GoldRecord, PRF1 (+ its properties), BackendReport,
  run_backend(), run(), _run_repeated(), print_report(), BackendSpec.

Standards, judgment calls:

- afinalize() reach-through. Eliminated by #330's new
  LightRAGBackend.afinalize() passthrough (merged in) -- run() now calls
  backend.afinalize() instead of backend.wrapper.afinalize().

- Duplicated repeat-loop orchestration. LightRAG's and GLiNER2's near-
  identical "build, run_backend, maybe finalize, print" loops collapsed into
  one shared _run_repeated(), parameterized by a BackendBuilder callable
  (LightRAG rebuilds + resets counters fresh per repeat; GLiNER2 reuses one
  loaded model via _reuse_gliner2_builder) and an optional finalize hook.

- Configuration as loose strings. name/text_property/relation_types/
  llm_model bundled into a new frozen BackendSpec (LIGHTRAG_SPEC/
  GLINER2_SPEC) instead of traveling run_backend()'s call chain as four
  separate parameters.

Spec:

- [High] Zero F1 reported as n/a. PRF1.f1's `if not p or not r` treated a
  real, computed 0.0 the same as "undefined" (None), since 0.0 is falsy.
  Fixed to `if p is None or r is None`, with an explicit 0.0 for the
  precision=recall=0 case (would otherwise divide by zero). Verified: zero-
  overlap now scores f1=0.0; genuine "nothing gold and nothing predicted"
  still scores None; perfect match still scores 1.0.

- [High] LightRAG cost undercounted. _counting_llm_wrapper now tokenizes
  every message in history_messages (each billed as real provider input,
  same as prompt/system_prompt), not just the two it previously counted.
  Verified: a call with two history messages now counts ~5x the prompt
  tokens of an otherwise-identical call with none.

- [High] Backends could target different Memgraph instances. New
  _pin_memgraph_env() sets MEMGRAPH_URL *and* MEMGRAPH_URI (+ empty
  MEMGRAPH_USER/MEMGRAPH_USERNAME/MEMGRAPH_PASSWORD defaults) unconditionally
  before any backend is built, instead of relying on
  lightrag_memgraph.core's _bridge_lightrag_env_names(), which only mirrors
  MEMGRAPH_URL onto MEMGRAPH_URI when MEMGRAPH_URI isn't already set --
  meaning a stale ambient MEMGRAPH_URI would previously have silently won
  over --memgraph-url for LightRAG's storage backend while this eval's own
  client used the correct one. Verified: a pre-set stale MEMGRAPH_URI is now
  overridden.

- [Medium] Gold span scoring discards spans. Not changed to offset-based
  matching (would reintroduce the exact cross-backend brittleness normalized-
  text matching was chosen to avoid -- LightRAG's LLM-normalized text
  doesn't char-align with source spans the way GLiNER2's does). Instead
  documented explicitly and precisely: GoldEntity/GoldRelation's docstrings
  now state start/end exist only for build_gold_corpus.py's own verbatim-text
  assertion and are never read back for scoring, and that matching is by
  unique normalized text per chunk (a set, collapsing repeated identical
  mentions), not span position or per-occurrence count. print_report() prints
  the same caveat.

Verified: ruff check/format clean, ty check clean, 71 unstructured2graph
unit tests pass (was 70 -- LightRAGBackend gained one afinalize() test in
#330). Ran the full rewritten script end to end against both backends (real
GLiNER2 model + real LightRAG/OpenAI) on a live dedicated Memgraph -- both
produce full reports with the new cost/token-breakdown and gold-span-caveat
lines. Independently verified each of the four Spec fixes in isolation
(PRF1.f1's three cases, Cypher-identifier rejection, history_messages token
delta, MEMGRAPH_URI override) rather than relying on the end-to-end run alone
to prove them.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant