Skip to content

unstructured2graph: GLiNER2Backend's single-pass extract() silently undercounts entities on long text (and is much slower than extract_long()) #336

Description

@antejavor

What

GLiNER2Backend._extract_sync calls model.extract(text, schema, ...) — GLiNER2's single-pass API, one forward pass over the whole text with no windowing. GLiNER2 is an encoder-only span/boundary classifier (not an LLM — see below), so it has a fixed effective context, and a session's combined text (post-#331's session-batching, one document per session rather than per-turn) routinely exceeds it. Past that point the call gets both much slower and silently degrades, with no error or warning.

GLiNER2 ships a purpose-built alternative for exactly this, model.extract_long(text, schema, chunk_size=..., chunk_overlap=..., ...), which windows the text with overlap and merges results back into document-global coordinates. _extract_sync never uses it.

Measured

Real session text pulled from the dedicated eval instance (bolt://localhost:7689), SessionsGraph._prepare_session's combined_text for one Tier 1 session — 16,280 chars, 12 turns. Model: fastino/gliner2.5-base-v1, package gliner2==2.0.0.

call time entities found
model.extract(text, schema, ...) (current code) 16.01s 18
model.extract_long(text, schema, chunk_size=384, chunk_overlap=64, ...) 4.34s 249

3.7x faster and ~13.8x more entities recovered — not a tradeoff, extract_long strictly dominates on this sample. Verified extract_long's return shape is identical to extract's (entities + relation_extraction keys, same structure) and span offsets are correctly re-mapped to global (not per-chunk-local) coordinates, so it is a drop-in replacement for _extract_sync's existing parsing logic — no changes needed beyond the one call site.

Reproduce:

from gliner2 import AutoExtractor
from unstructured2graph.ontology import DEFAULT_ONTOLOGY

model = AutoExtractor.from_pretrained("fastino/gliner2.5-base-v1")
schema = model.create_schema().entities({t.label: t.description for t in DEFAULT_ONTOLOGY.entity_types})

r_full = model.extract(text, schema, include_spans=True, include_confidence=True)
r_long = model.extract_long(text, schema, chunk_size=384, chunk_overlap=64, include_spans=True, include_confidence=True)

Why it matters

Blocks a full-scale GLiNER2-backed Tier 1 run (map #297). A first attempt at context-graph-eval run --limit 100 --reconcile-backend gliner2 measured ~14s/session average across 124 real sessions before being killed — extrapolated to the full 4,599-session Tier 1 batch, ~17-18 hours. Profiling one session directly attributed 14.43s of a 15.19s reconcile_session call to this single _extract_sync call (Memgraph I/O: 0.77s) — i.e. this bug is the bottleneck, not GLiNER2 itself, not Memgraph, not concurrency.

Also a silent correctness bug in production. SessionsGraph.reconcile_session builds the same kind of combined per-session text for any real (non-eval) caller too, post-#331. Any session long enough to exceed GLiNER2's effective context silently loses most of its entities, with summary.status == "completed" and no indication anything was dropped.

Suggested fix

Swap GLiNER2Backend._extract_sync's self.model.extract(text, schema, ...) call to self.model.extract_long(text, schema, chunk_size=..., chunk_overlap=..., ...). Pick chunk_size/chunk_overlap deliberately rather than copying the 384/64 used above for measurement (worth a quick sweep — too small trades per-chunk overhead for coverage, too large re-approaches this same bug). Existing entity-identity derivation (_entity_id, keyed on chunk hash + entity_type + normalized text, not on span position) needs no change either way.

Found while attempting the first full-scale Tier 1 run with GLiNER2 reconciliation (map #297). Bug is in GLiNER2Backend._extract_sync, added in #330.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions