Skip to content

unstructured2graph: GLiNER2Backend has no cross-chunk coreference and no batched reconciliation path (LightRAG has both) #339

Description

@antejavor

What

GLiNER2Backend (#330) and LightRAGBackend are both ExtractionBackend implementations, but two aspects of LightRAG's behavior have no GLiNER2 equivalent today. Neither is a bug (#336/#338 already fixed the actual bug, GLiNER2's context-window truncation) -- these are architecture gaps worth closing deliberately.

1. No cross-chunk entity coreference

LightRAG's LLM merges "the same" entity mentioned across different chunks/documents into one canonical node, with a merged description. GLiNER2Backend._entity_id is deliberately scoped to (chunk_hash, entity_type, normalized_text) -- documented in its own docstring as a known limitation: "unlike LightRAG, this backend does no cross-chunk coreference -- the same real-world entity mentioned in two different chunks becomes two separate Memgraph nodes."

This matters more than it used to: post-#331, a session is one document/chunk, so within a session this is less of an issue, but across sessions (the common case -- the same person, team, or project recurring over many sessions) GLiNER2 will keep creating duplicate nodes indefinitely, where LightRAG consolidates them.

Closing this is a real design decision, not a quick fix: what counts as "the same" entity (exact normalized text? embedding similarity? a threshold?), and where the merge/upsert step lives (write time, per session? a periodic post-process?).

2. No batched/parallel reconciliation path

LightRAG has SessionsGraph.reconcile_sessions_batch() (#333): many sessions staged into one shared LightRAG worker-pool pass, context_graph_eval.reconcile_batch's default path. GLiNER2 only has the plain sequential reconcile_session loop (_reconcile_batch_gliner2 in context_graph_eval/reconcile.py, added for the GLiNER2 eval-pipeline work) -- one session at a time, awaited in order.

Unlike LightRAG's concurrency (network-bound API calls, safe to run many at once), GLiNER2 model-instance concurrency is not clearly safe -- checked directly: batch_extract/extract_long call self.eval() and lazily cache self._inference_collator on every invocation, with no documented thread-safety guarantee, so firing multiple asyncio.to_thread calls at the same model instance from separate sessions is a real race risk (benign at best, a corrupted result at worst), not something to just try.

GLiNER2 does have a real equivalent already in the library, unused: model.batch_extract(texts, schema, ...) performs genuine vectorized batching -- one tokenization/forward pass over many texts together, safe because it's a single call, not concurrent calls. Measured (short synthetic texts, fastino/gliner2.5-base-v1): 20 texts sequentially via extract() took 3.05s, the same 20 via one batch_extract() call took 2.36s (1.29x) -- a real, if modest, win on short text; likely larger on the longer real session texts extract_long's own windowing already produces many chunks for.

Why it matters

Both gaps are specifically about parity between the two backends as pluggable options, not about either one being broken:

Suggested fix

  1. Batching (lower-risk, do first): give _reconcile_batch_gliner2 a batched variant using model.batch_extract() across multiple sessions' texts in one call, rather than looping reconcile_session one at a time. Needs GLiNER2Backend to expose a batch-shaped entry point (aingest_chunks, plural, alongside the existing aingest_chunk) so callers outside the eval pipeline benefit too, not just this one caller.
  2. Coreference (real design work, do second): needs its own discussion on identity/merge semantics before implementation -- not sized here.

Found while comparing GLiNER2Backend and LightRAGBackend after landing #338 (map: unstructured2graph pluggable backends, #330).

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions