Skip to content

What gliner2 2.0.0 actually offers for entity attributes #362

Description

@antejavor

Part of #344

Question

Can the installed gliner2==2.0.0 attach values to an entity — and if so, through which API, in what output shape, and under what constraints?

#350 found that 3 of 6 question types in its sample have no expressible answer as a relation: a personal-best time (25:50), a count (3), a shift window (8am-4pm). #361 has to decide whether attributes enter the model, and it cannot decide that without knowing what the library can actually do. This is a fact-finding ticket, not a decision: same shape as #345, which is also the quality bar — per-claim file:line and real command output, against the source on disk and a live run of fastino/gliner2.5-base-v1, never the upstream blog.

To answer:

  1. Is there an attribute mechanism on the joint path at all? gliner2.joint_ie.schema.JointSchema.entity() takes name, description, threshold, candidate_threshold, max_candidates, allow_nested — no obvious field slot. Does one exist elsewhere on JointSchema, on the compiled schema, or via **aliases? What JointSchema actually enforces #345 investigated relations only.
  2. The legacy (non-joint) API. gliner2's Schema exposes structure/classification tasks alongside entities and relations. Are those the real mechanism for values, what do they return, and can they run in the same pass as joint entity+relation extraction — or does using them mean a second inference call per window, and losing the entity-id coupling What JointSchema actually enforces #345 verified for relations?
  3. Output shape. Where does a value land: a span with offsets, free text, or a label from a closed set? Does it carry its own confidence? Does it come back attached to a specific entity id, or does it need re-matching the way relation endpoints used to (_extract_sync's span-matching apparatus, which What JointSchema actually enforces #345 retired)?
  4. Are values constrained the way endpoints are? Can an attribute declare a type, an enum of allowed values, or cardinality — and is that enforced during decoding (allow_edge-style) or after? Is there a feasible-equivalent signal when a declared attribute cannot be filled? Wire a constrained ontology and read the edges #350 established that prohibitive constraints never make a window infeasible; check whether attributes change that.
  5. Composition with our pipeline. Does it survive per-window driving (we own the windowing per Wire a constrained ontology and read the edges #350/Chunking and the durability of constraints #352, not extract_long_text)? Does declaring attributes multiply candidates the way permissive endpoints do — Wire a constrained ontology and read the edges #350 measured typed vs all-types endpoint expansion — and what does it cost per window in wall clock?
  6. Version reality. Upstream markets "span attributes" for 2.5. Establish what the installed package and the checkpoint we run actually support, and whether anything here needs a version bump we have not taken (note the manual-install constraint and the transformers floor in gliner2_backend.py's module docstring).

Deliverable: a research note at unstructured2graph/docs/research/gliner2-attributes.md on a research/gliner2-attributes branch — branch only, no PR, same convention as the #345 note — linked back here, with anything inconclusive flagged as such rather than smoothed over.

Environment: the only venv with gliner2 installed is .claude/worktrees/backend-comparison/.venv.

Blocks #361.

Activity

  1. self-assigned this
    on Sep 22, 2026
  2. antejavor commented on Sep 22, 2026

    @antejavor
    ContributorAuthor

    Resolution

    Full write-up, 754 lines with per-claim file:line and command output: unstructured2graph/docs/research/gliner2-attributes.md on branch research/gliner2-attributes (commit 0020973). Branch only, no PR, same convention as the #345 note.

    1. No attribute mechanism exists on the joint path — none, anywhere

    EntitySpec carries six fields, none a value (joint_ie/schema.py:22-32), and entity() is keyword-only with no **aliases — unlike relation(), so the absence reads as deliberate rather than an oversight. JointSchema has no field/structure/classification/entity_attributes; compile_schema hard-codes "json_structures": [] and "classifications": [] (joint_ie/compiler.py:26-28); JointResult/JointEntity have no slot to put a value in.

    Verified by forcing it: a structure appended directly into compiled.model_schema runs without error and the result still contains only ['entities', 'relations'] — silently dropped.

    A trap worth writing down: JointSchema.from_dict ignores any unknown top-level key without complaint. A config file carrying an attributes: block would load cleanly, validate, and do nothing.

    2. The legacy API is the mechanism — three of them — and using it costs typed relations

    • entity_attributes() + AttributeGroup — closed label sets bound to entity spans
    • structure().field() — the only mechanism that extracts an actual value
    • classification() — document-level, entity-free

    All three compose with entities and relations in a single extract() call (verified: result keys ['facts', 'entities', 'topic', 'relation_extraction']). But that call is the legacy runtime #345 retired: relations there compile to {"head": "", "tail": ""} and endpoints come back as span dicts.

    So the trade is stark: values in the same pass, or typed relations — not both. Keeping JointSchema means a second inference call per window (+90% wall clock, measured), and that second call has no entity ids — which resurrects exactly the span-matching apparatus #345 established we could delete.

    3. Three landing places, only one of which is a real value

    mechanism what comes back attached to
    entity_attributes {"label", "confidence"} — no offsets, no text grounding co-located with the entity row, but legacy output has no ids
    structure().field() a real span: {"text": "25:50", "confidence": 0.9999, "start": 82, "end": 87} nothing — document-scoped
    structure(mode="natural", anchor=...) per-anchor instances the only per-entity construct

    Record mode's anchor spans match joint entity ids exactly by (start, end) — verified (414,419) -> e10, (540,545) -> e14 — so coupling the two paths is mechanically recoverable even though the library does not hand it to you.

    4. Constrained by construction; nothing enforced during decoding; no feasibility signal

    Closed sets (AttributeGroup.labels, field choices=) are enforced by only ever scoring declared labels — there is no allow_edge analogue because there is no search. Cardinality is dtype ("str" -> spans[0], "list" -> all); cardinality=/exclusive= are inert outside record mode. validators= (regex) is a provable post-hoc filter: the same span returns at byte-identical confidence 0.9882827401161194 with and without a matching validator. There is no feasible analogue — the legacy result is a plain dict, so #350's infeasibility question has no counterpart here.

    Two sharp edges:

    • A single-label group can never abstain. It is a softmax argmax: three unrelated entities given a planet group all came back Mars.
    • choices= hallucinates. It returned "alpha" at confidence 0.99 from a list whose members appear nowhere in the text.

    The only abstention route is multi_label=True plus a threshold. An unfillable field returns null; an all-unfillable structure is silently dropped entirely.

    5. Composes with per-window driving, is cheap — and conflicts across windows

    Costs on a 572-char window, CPU, median of 5: joint baseline 157ms; +1 attribute group (3 labels) +12ms; +6-field structure +12ms; +18-field structure +88ms; joint-then-legacy as two calls 299ms (+90%). Candidate growth is linear, not quadratic — attribute labels become extra entity queries (5 -> 8), unlike the endpoint expansion #350 measured.

    The real problem is structural, and it lands on #352: a structure is one instance per document, so every window competes to fill every field. Same text, same schema — purchase_date came back as both March 14, 2025 and last Saturday; pair_count as both 3 pairs and Three pairs; personal_best flipped from 26:34 to 25:50 purely by changing window size 220 -> 400. Nothing in the output says which is right.

    6. No version bump exists, let alone is needed

    gliner2==2.0.0 is installed and is the latest on PyPI. "Span attributes" is Schema.entity_attributes — its own docstring says the labels are "decoded as span attributes" (inference/schema.py:331), and AttributeGroup is a public export (__init__.py:42). The "2.5" upstream markets is the checkpoint generation (fastino/gliner2.5-base-v1), not a library version, and we already run it — its BoundaryHeadSettings has enable_records=True, enable_relations=True, enable_abstention=True. Requires-Dist: transformers<5,>=4.38 is still 2.0.0's pin, so adopting attributes needs no new dependency, extra, or version, and nothing touches the CVE-2026-1839 floor.

    Beyond the brief: #350's gap was its vocabulary, not the API

    Declaring Duration, Quantity, TimeWindow, Money, Date as entity types and relating into them expresses all five value facts on the joint path — one pass, entity ids and #345's constraint enforcement intact, feasible=True:

    REL owns_count     Person:'user'  -> Quantity:'3 pairs'        conf=0.985
    REL paid           Person:'user'  -> Money:'$129.99'           conf=0.964
    REL personal_best  Person:'user'  -> Duration:'26:34'          conf=0.971
    REL works_shift    Person:'Admon' -> TimeWindow:'Sunday shift' conf=0.879
    

    So "3 of 6 question types have no expressible answer" (#350) is a property of that prototype's vocabulary, not of the library. Carried to #361 as a third option alongside attributes and reified assertions.

    Three caveats, all visible in that same output: 2 of 4 values are wrong (personal_best bound 26:34, the previous PB — 25:50 was never extracted as a Duration at all; works_shift bound 'Sunday shift' while '8am-4pm' sat two tokens away, extracted at conf 0.998); relations triplicated (the same 3.3x mention duplication #350 measured); 280ms against a 124ms baseline. The structure path got 5/5 on the same text where this got 2/4 — a sample of one, and not a quality result.

    Relatedly, and worth keeping: in the same single legacy pass, the relation mechanism answered personal_best = 26:34 while the structure mechanism answered 25:50. Same model, same text, two mechanisms, two different answers.

    Could not determine

    • Accuracy of any of this. Everything is one 572-char window. The 5/5-vs-2/4 split must not be read as a quality result; that needs the eval over the session corpus.
    • Whether record mode's field-to-anchor binding is controllable from the schema. It bound a value 460 chars away to the wrong anchor — gave a runner's personal best to Admon the physio, twice, and left shift null. occurrence_policy ("all"|"first"|"error_on_ambiguous"|"latent_all") exists and was not swept; error_on_ambiguous in particular may surface ambiguity rather than guess. This is the most consequential unknown here, because record mode is the only per-entity value construct — folded into Values, not links: do entity attributes belong in the model #361 as work its session does.
    • Whether attributes survive extract_long (source says yes; not run, since we own windowing per Wire a constrained ontology and read the edges #350/Chunking and the durability of constraints #352).
    • The span (non-boundary) architecture, which carries a second, separate attribute implementation with an extra overlap-dedupe pass.
    • Whether a value-shaped-entity-type vocabulary could ever be derived (The derivation contract for LlmRecommendationStrategy #353) rather than hand-written.
    • Calibration: structure-field confidences cluster at 0.99+, attribute confidences spread 0.44-0.99. Different heads, not calibrated against each other — a single threshold across both would be unsound.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions