Skip to content

unstructured2graph: GLiNER2Backend uses extract_long(), not extract() (#336) - #338

Merged
antejavor merged 1 commit into
mainfrom
unstructured2graph-gliner2-extract-long
Sep 16, 2026
Merged

antejavor merged 1 commit into
mainfrom
unstructured2graph-gliner2-extract-long

Conversation

@antejavor

Copy link
Copy Markdown
Contributor

Summary

Fixes #336: GLiNER2Backend._extract_sync called model.extract() -- a single forward pass with no windowing. GLiNER2 is an encoder-only span/boundary classifier with a fixed effective context, and a session's combined text (one document per session since #331) routinely exceeds it. Past that point extract() both slows down and silently drops most entities, with no error.

  • Swapped to model.extract_long(text, schema, chunk_size=..., chunk_overlap=..., ...) -- GLiNER2's own purpose-built API for this, windowing the text with overlap and merging results back into document-global coordinates.
  • chunk_size/chunk_overlap are now GLiNER2Backend constructor parameters (defaulting to GLiNER2's own 384/64), alongside the existing entity_confidence_threshold/relation_confidence_threshold knobs.

Measured

Real session text pulled from the dedicated eval instance -- 16,280 chars, 12 turns. Model: fastino/gliner2.5-base-v1, package gliner2==2.0.0.

call time entities found
model.extract(text, schema, ...) (before) 16.0s 18
model.extract_long(text, schema, chunk_size=384, chunk_overlap=64, ...) (after) 4.3s 249

3.7x faster and ~13.8x more entities recovered -- not a tradeoff. Verified extract_long's return shape is identical to extract's (entities + relation_extraction keys, same structure, span offsets correctly merged to global coordinates), so _extract_sync's existing parsing logic needed no change beyond the one call site.

Test plan

  • ruff check / ruff format --check clean
  • ty check clean (pre-existing ty: ignore[unresolved-import] unused-locally warning, expected: gliner2 isn't installed in CI, only in a local venv with it manually installed -- see the module's own docstring)
  • Unit tests (fake model, no real gliner2/Memgraph needed): 14 passed, including two new regression tests -- chunk_size/chunk_overlap default to GLiNER2's own values and are configurable, and extract_long() (not extract()) is called with the configured values
  • Real-model e2e tests (test_e2e_gliner2.py, real gliner2==2.0.0 + live Memgraph): 2 passed -- confirms no regression on short text
  • Full unstructured2graph suite: 102 passed, 7 skipped (OPENAI_API_KEY-gated LightRAG e2e tests, unrelated to this change)

…#336)

extract() is a single forward pass with no windowing, and GLiNER2 has a
fixed effective context. A session's combined text (one document per
session since #331) routinely exceeds it, and past that point extract()
both slows down and silently drops most entities, with no error.

Measured on a real 16k-char session: extract() took 16.0s and found 18
entities; extract_long() (chunk_size=384, chunk_overlap=64, GLiNER2's own
defaults) took 4.3s and found 249. Verified extract_long()'s output shape
is identical to extract()'s, so this is a drop-in swap in _extract_sync.
chunk_size/chunk_overlap are now GLiNER2Backend constructor parameters.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

unstructured2graph: GLiNER2Backend's single-pass extract() silently undercounts entities on long text (and is much slower than extract_long())

1 participant