You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
GLiNER2Backend._extract_sync uses none of this — it calls model.create_schema() (returns gliner2.inference.schema.Schema), whose .relations() always writes {"head": "", "tail": ""}.
Open, and everything else on this map leans on the answers:
Entry point.AutoExtractor.extract(text, schema: Union[SchemaAPI, Dict]) — does a JointSchema satisfy that union directly, or does it need .build()/to_dict() first, or a different call entirely (gliner2.joint_ie.engine)?
Output shape. What do relations look like coming back? Still the {"entities": ..., "relation_extraction": ...} per-key shape _extract_sync parses today, or something else? Do include_spans / include_confidence still work?
Real enforcement vs post-filtering. Are max_per_head / no_self_loops / head-tail typing genuinely applied during beam-search decoding (as the write-up claims — "invalid combinations are never admitted into the search"), or applied as a filter afterwards? This matters directly: if it is post-filtering, the "valid by construction" argument for extraction-time constraints collapses and the constraint-model decision changes.
Unconstrained relations. What happens when head/tail are empty or omitted? The plan's backward-compat rule is that start_labels=()/end_labels=() means unconstrained, never a validation failure — check this is expressible.
Entity-schema coupling. The current code observed a located_in relation whose head/tail spans the entities pass never surfaced, and aingest_chunk skips-and-logs those. Does JointSchema make relation endpoints guaranteed-present in the entity output (the blog implies relations only reference entities that exist)? If so, the skip-and-log path may become dead code.
How to answer
Primary sources only: read the installed package source under .venv/lib/python3.11/site-packages/gliner2/joint_ie/ (schema.py, compiler.py, engine.py) and api_client.py, plus the upstream repo https://github.com/fastino-ai/GLiNER2 and https://fastino.ai/blog/gliner2-5-span-free-information-extraction. Where source reading is ambiguous, run it — the model checkpoint fastino/gliner2.5-base-v1 is already downloaded.
Environment note:gliner2 is a manual install that conflicts with the workspace's transformers>=5.0.0rc3 floor, and a bare uv run silently re-syncs the venv and clobbers it. Use uv run --no-sync, and if imports break, reinstall with uv run --no-sync python -m pip install 'gliner2[local]>=2.0.0'.
Part of #344
Question
What does
gliner2.joint_ie.schema.JointSchemaactually give us, and what are its limits?Established already (installed
gliner2==2.0.0, latest on PyPI):JointSchemaexists with.entity(name, description, *, threshold, candidate_threshold, max_candidates, allow_nested)and.relation(name, head, tail, description, *, threshold, candidate_threshold, directed, symmetric, inverse, allow_self, max_per_head, max_per_tail)plus.no_self_loops(relation=None).RelationSpectakeshead: tuple[str, ...]andtail: tuple[str, ...]— typed domain/range.Constraint,AcyclicRelation,MaxRelationsPerHead,MaxRelationsPerTail,NoSelfLoops,EntitySpec.GLiNER2Backend._extract_syncuses none of this — it callsmodel.create_schema()(returnsgliner2.inference.schema.Schema), whose.relations()always writes{"head": "", "tail": ""}.Open, and everything else on this map leans on the answers:
AutoExtractor.extract(text, schema: Union[SchemaAPI, Dict])— does aJointSchemasatisfy that union directly, or does it need.build()/to_dict()first, or a different call entirely (gliner2.joint_ie.engine)?extract_long()'s windowing? This is not optional — unstructured2graph: GLiNER2Backend's single-pass extract() silently undercounts entities on long text (and is much slower than extract_long()) #336/unstructured2graph: GLiNER2Backend uses extract_long(), not extract() (#336) #338 madeextract_long()mandatory because session texts blow the context window, and the blog claims spans are remapped to original-document offsets across chunks. Confirm whether constraints hold across window boundaries or only within one window.{"entities": ..., "relation_extraction": ...}per-key shape_extract_syncparses today, or something else? Doinclude_spans/include_confidencestill work?max_per_head/no_self_loops/ head-tail typing genuinely applied during beam-search decoding (as the write-up claims — "invalid combinations are never admitted into the search"), or applied as a filter afterwards? This matters directly: if it is post-filtering, the "valid by construction" argument for extraction-time constraints collapses and the constraint-model decision changes.head/tailare empty or omitted? The plan's backward-compat rule is thatstart_labels=()/end_labels=()means unconstrained, never a validation failure — check this is expressible.located_inrelation whose head/tail spans the entities pass never surfaced, andaingest_chunkskips-and-logs those. DoesJointSchemamake relation endpoints guaranteed-present in the entity output (the blog implies relations only reference entities that exist)? If so, the skip-and-log path may become dead code.How to answer
Primary sources only: read the installed package source under
.venv/lib/python3.11/site-packages/gliner2/joint_ie/(schema.py,compiler.py,engine.py) andapi_client.py, plus the upstream repo https://github.com/fastino-ai/GLiNER2 and https://fastino.ai/blog/gliner2-5-span-free-information-extraction. Where source reading is ambiguous, run it — the model checkpointfastino/gliner2.5-base-v1is already downloaded.Environment note:
gliner2is a manual install that conflicts with the workspace'stransformers>=5.0.0rc3floor, and a bareuv runsilently re-syncs the venv and clobbers it. Useuv run --no-sync, and if imports break, reinstall withuv run --no-sync python -m pip install 'gliner2[local]>=2.0.0'.