Research

Today we release GLiNER2.5, the most significant upgrade to the GLiNER architecture to date. Across a suite of 16 benchmarks spanning diverse classification and extraction tasks, GLiNER2.5 achieves higher overall average F1 scores in comparison to GLiNER2, including a significant 24.75-point gain on XNLI. We release three variants on Hugging Face under Apache 2.0: gliner2.5-base-v1 (0.2B), gliner2.5-multi-v1 (0.3B), and gliner2.5-small-v1 (74M).
Two changes drive this release. The architecture replaces span enumeration with boundary prediction, scoring entity start and end positions directly, so inference scales linearly with document length. And the training data was generated with the Fastino Data Agent rather than assembled from existing datasets, covering task combinations, document formats, and languages that no public dataset does.
Together, these bring five new capabilities to GLiNER users: long-context extraction and classification over full documents, unlimited span length for entities of any size, joint entity and relation extraction for coherent knowledge graphs, constrained classification for outputs that are valid by construction, and richer extraction context with span attributes.
In this post, we introduce these five capabilities and how they work in practice, give an overview of the updated architecture, and show how the model benchmarks against GLiNER2.
Long-Context Extraction for Full Documents
GLiNER2.5 handles long documents in two ways, one in the model and one in the library.
First, the model itself takes longer inputs. Removing explicit span representations cut memory use enough to train on sequences up to 4,096 words, so most contracts, reports, and transcripts fit in a single forward pass without chunking.
Second, the library now offers native chunking. Every extraction task has a long-document variant with batch versions: entities, classification, JSON schemas, relations, and generic extraction. The document is split into overlapping word chunks, extraction runs on each, and every span is remapped to character offsets in the original document. Duplicates across overlaps are merged under deterministic policies you control, and any returned span can be checked directly against the source text.
GLiNER2 required you to truncate, or write your own chunking and handle offset remapping and cross-chunk duplicates yourself. Now extracting obligations from a full contract or tracking entities across an hour-long transcript is one call, and the schema API is the same at any length.

Unlimited Span Length for Whole-Entity Extraction
GLiNER2.5 removes the maximum span width that constrained GLiNER2. An entity can now begin at the first token and end at the last, so spans of any length are extractable: full postal addresses, clause-length legal references, complete product titles, multi-line table cells.
GLiNER2 enumerated candidate spans up to a fixed width, typically around twelve words, and anything longer was structurally invisible. The model did not score it poorly; it never scored it at all. Users worked around this by raising the width limit, which increased compute with every step, or by extracting fragments and stitching them together afterwards.

The boundary prediction architecture removes the limit rather than raising it. Because the model scores where an entity starts and where it ends instead of scoring enumerated spans, there is no width axis in the computation at all. A forty-word indemnification clause costs the same to locate as a two-word name, and no configuration is required.
This changes what is practical to extract. Legal and contract work can treat whole clauses as entities. Document processing can capture full addresses and references as single spans rather than reassembled parts. And because span length no longer trades against compute, schemas can include long entity types without a performance penalty.
Joint Information Extraction for Coherent Knowledge Graphs
GLiNER2.5 extracts entities and relations together as a single, globally decoded graph, which means that every relation in the output connects entities that actually exist in the result, and the structure as a whole obeys the rules of your schema.
When entities and relations are scored independently, the combined output is frequently inconsistent. Relations may reference entities that fell below the extraction threshold, a person may be assigned two employers where the schema intends one, and hierarchical relations may contain cycles instead of forming a valid tree. Resolving these inconsistencies has traditionally been left to the user, requiring post-hoc filtering rules.
With GLiNER2.5, the user declares entity types, typed relations, and structural rules in a joint schema and the returned output is guaranteed to conform. This is what makes the output suitable for knowledge graphs specifically. A knowledge graph is only as reliable as its least consistent edge, and when triples are extracted independently, every ingestion pipeline needs logic to reject dangling references, enforce cardinality, and break cycles. Joint decoding moves that work into the extraction itself, so what reaches the graph is already well-formed.
The model achieves this by scoring all candidate entities and relations in a single forward pass. A beam search then assembles the highest-scoring combination, checking each candidate against the declared rules as the solution is built rather than filtering the output afterwards. Because invalid combinations are never admitted into the search, the returned graph satisfies the schema by construction.

This is a substantial departure from GLiNER2, where relation extraction returned independently thresholded triples and consistency was the user's responsibility. With GLiNER2.5, the model's output can be treated as a well-formed graph from the moment it is returned, which makes it possible to populate knowledge graphs, build entity-linking pipelines, and feed downstream systems directly, without a validation layer in between.
Constrained Classification
With GLiNER2, classification joined extraction as a first-class capability, spanning multi-label, multi-class, hierarchical, and multi-task setups. GLiNER2.5 builds on this foundation by enforcing user-declared constraints during decoding, which guarantees valid outputs and allows predictions on one task to inform another.
Traditional unconstrained classification can result in predictions that contradict each other across tasks. Consider GLiGuard, our guardrail model, which classifies both the overall safety of a prompt and the type of harm present when it is unsafe. Without constraints, these two tasks are decoded independently, so a prompt can be labeled safe while simultaneously being flagged for prompt injection, a contradiction that downstream code then has to detect and resolve. With constrained classification, a single declared rule, that a harm type may only be assigned when the prompt is labeled unsafe, makes this contradiction impossible: the decoder never admits the invalid combination in the first place. And if no valid assignment exists at all, the classifier raises an error rather than returning an invalid classification.

Richer Extraction Context with Span Attributes
Span attributes allow the model to apply descriptive labels to the spans it extracts, such as a mention of a product and whether that mention is positive or negative in sentiment. GLiNER2.5 decodes these attributes in the same forward pass as the entities themselves, so entities come back qualified rather than flat and lacking context.
When creating the extraction schema, the user defines attribute groups, each a small set of possible values, and attaches them to entity types, so that a group can apply to all entities or only to specific ones. A group can assign a single value per span or several, with a confidence threshold the user controls. Because the attributes are part of the schema, the model scores them alongside the entities in one pass, and each returned span carries its attribute values directly.

GLiNER2 could classify and extract in the same pass, but its classifications applied to the input as a whole rather than to each span, so entities came back flat. With GLiNER2.5, the attributes are decoded span-by-span within the original forward pass, so a single call returns the spans and their qualifications together.
The GLiNER2.5 Boundary Architecture
GLiNER2.5 is built on a new architecture that changes how the model locates entities in text. Earlier GLiNER models enumerated candidate spans: every possible start position paired with every allowed width, each scored against the schema's entity types. This design tied computation to a width axis and imposed a hard ceiling on how long an extracted entity could be.
The new architecture removes the span enumeration entirely. The shared encoder still processes the text and the schema's queries together in one pass, but instead of scoring spans, the model predicts where entities begin and end: for each query, it produces start and end scores over the text's token boundaries and inside scores over the tokens themselves. An entity is located by its boundaries rather than matched against an enumerated list of candidates, and no full position-by-position score matrix is ever built.
From these scores, a sparse proposal stage selects the most promising start and end boundaries per query and pairs them, with no restriction on how far apart a start and end may sit, so a span can open at the first token and close at the last. A reranking head then scores each proposed candidate using the boundary evidence and the span's content, and relation candidates are drawn from this same pool rather than through a separate extraction path. Computation stays linear in sequence length for a fixed schema and candidate budget, and the maximum entity width that constrained earlier GLiNER models is gone.
How we built it
Synthetic data generation
For GLiNER2.5 we generated our training data with the Fastino Data Agent instead of assembling existing datasets.
The data covers the tasks GLiNER2.5 supports: named entity recognition, relation extraction, classification, and structured JSON records. Each appears alone and in combination inside a single schema, so the model learns to handle several tasks in one pass.
Generation runs over a matrix of domains and tasks. For each pair we give the agent a short prompt plus a task spec with the things we want to control, like the number of entity and relation types, document length, languages, and output format. The rest is left to the agent. It picks the document layouts, writes the schemas, varies the surface forms, and applies its own diversity criteria across samples.
Benchmarks
We evaluate zero-shot on 16 public datasets covering classification and extraction: sentiment, topic, intent, and NLI on one side, and general, domain-specific, and multilingual NER on the other. All results are macro F1 against GLiNER2 at matched model sizes, with per-dataset scores included so regressions are visible alongside gains.
Our benchmarks demonstrate that developers do not have to choose between new structural extraction capabilities and the baseline extraction quality of GLiNER2. Despite the addition of joint relational extraction and constrained classification heads, the underlying boundary prediction architecture maintains or improves on previous performance.
Below, we compare the GLiNER2.5 Multilingual (0.3B) and GLiNER2.5 Base (0.2B) models directly against their GLiNER2 predecessors.
Overall Performance
We report overall performance as the macro-average F1 score across a diverse suite of 16 classification and extraction tasks, which we list in the appendix, to evaluate the general utility of the models.
GLiNER2.5 Multilingual achieves an overall average F1 score of 56.17, outperforming GLiNER2 Multilingual (56.09).
GLiNER2.5 Base achieves an overall average F1 score of 54.87, compared to 53.34 for GLiNER2 Base.

Classification
Natural Language Inference (NLI)
NLI benchmarks measure textual entailment and logical reasoning, requiring the model to determine if a pair of sentences entail, contradict, or are neutral toward each other. This is a critical proxy for how well a model understands negation and contextual constraints.
On the XNLI dataset, GLiNER2.5 Multilingual achieves 62.30 F1, a leap of nearly 25 points over GLiNER 2 Multilingual (37.55 F1).
GLiNER2.5 Base's F1 score also increases to 54.49, up from GLiNER2 Base's score of 49.01.
Extraction
General NER
This benchmark evaluates zero-shot entity extraction accuracy across broad, out-of-domain contexts using the Few-NERD dataset.
GLiNER2.5 Base scores 55.14 F1, showing improvement over GLiNER2 Base (47.22 F1).
GLiNER2.5 Multilingual scores 52.37 F1, maintaining stable extraction quality relative to GLiNER2 Multilingual (51.49 F1).
Multilingual NER
This benchmark measures zero-shot entity extraction performance on non-English text using the Romanian (RONEC) dataset. We share this benchmark as compelling evidence for the model's ability to generalize to languages it was not explicitly trained on, as Romanian was not a target language during training.
GLiNER2.5 Multilingual scores 40.13 F1, outperforming GLiNER2 Multilingual (38.86 F1).
GLiNER2.5 Base scores 37.01 F1, representing a substantial gain over GLiNER2 Base (31.55 F1).
Use cases
GLiNER2.5's capabilities are most visible in what they let you build. Each of the use cases below previously required code around the model, chunking logic, validation layers, second passes, or handing the task to an LLM; each now runs as a single call against a single small model.
Model and agent routing. Route tasks to sub-agents, tools, or model tiers with the routing decision and task type decoded jointly under your compatibility rules, rather than as independent predictions that downstream code must reconcile.
Agent guardrails. Screen agent actions with a classifier whose safety verdict and harm type are decoded under a declared rule, so a contradictory verdict is no longer possible.
Knowledge graph construction for agent memory. Give agents a queryable graph of people, projects, and commitments, built from the documents, email, and chat history they read, where joint decoding keeps every edge typed and the structure consistent as the graph grows.
PII detection and redaction. Find personal data across an entire contract or transcript rather than its first window, with each span carrying a global character offset for redaction at the source.
Contract review. Extract parties, obligations, and termination clauses from all three hundred pages, not the portion that fits in the model window, and verify each span against the source text.
Clinical extraction. Pull symptoms and medications with negation status and dosage form attached to each span in the same forward pass, instead of re-classifying every extracted span in a second one.
To see GLiNER2.5 in action, check out tutorials in our GitHub repo.
Conclusion
GLiNER2.5 rests on a new boundary prediction architecture that scores where entities begin and end rather than enumerating candidate spans, and this foundation is what makes its five new capabilities possible: long-context extraction over entire documents, unlimited span length extraction, joint entity and relation extraction, constrained classification, and extraction of span attributes. Together they open use cases that previously required either a pipeline of models or a large language model: populating knowledge graphs directly from raw text, extracting structured records from full contracts and reports, and running guardrail or triage classifiers whose outputs are valid by construction. Across our benchmark suite, GLiNER2.5 matches or exceeds GLiNER2's overall macro-average F1 scores, reaching up to 56.17 F1, demonstrating that the new architecture does not trade away the extraction quality the GLiNER family is known for.
Get started
GLiNER2.5 is available in three variations: gliner2.5-base-v1 (0.2B), gliner2.5-multi-v1 (0.3B), gliner2.5-small-v1 (74M). As always, we release the model weights on the Hugging Face Hub under the Apache 2.0 license.
Check out tutorials in our GitHub repo.
Appendix
Full benchmark results
Task averages
Group | 2.5 Multi | 2.5 Base | GLiNER2 Multi | GLiNER2 Base |
|---|---|---|---|---|
Overall | 56.17 | 54.87 | 56.09 | 53.34 |
Classification | 72.44 | 69.86 | 70.32 | 68.89 |
Extraction | 46.40 | 45.88 | 47.56 | 44.01 |
Classification domains
Domain | 2.5 Multi | 2.5 Base | GLiNER2 Multi | GLiNER2 Base |
|---|---|---|---|---|
Sentiment | 80.01 | 77.58 | 82.94 | 76.73 |
Topic / news | 70.99 | 69.71 | 72.93 | 70.54 |
NLI | 62.30 | 54.49 | 37.55 | 49.01 |
Intent | 61.32 | 62.20 | 62.59 | 63.62 |
Extraction domains
Domain | 2.5 Multi | 2.5 Base | GLiNER2 Multi | GLiNER2 Base |
|---|---|---|---|---|
CrossNER | 54.85 | 58.30 | 57.84 | 59.06 |
General NER | 52.37 | 55.14 | 51.49 | 47.22 |
Multilingual NER | 40.13 | 37.01 | 38.86 | 31.55 |
Per-dataset scores
Dataset | Domain | 2.5 Multi | 2.5 Base | GLiNER2 Multi | GLiNER2 Base |
|---|---|---|---|---|---|
ag_news | Topic / news | 70.99 | 69.71 | 72.93 | 70.54 |
clinc_oos | Intent | 61.32 | 62.20 | 62.59 | 63.62 |
imdb | Sentiment | 85.96 | 88.10 | 89.42 | 89.70 |
multilingual_sentiment | Sentiment | 79.42 | 63.14 | 81.30 | 57.57 |
rotten_tomatoes | Sentiment | 74.67 | 81.51 | 78.11 | 82.91 |
xnli | NLI | 62.30 | 54.49 | 37.55 | 49.01 |
crossner_ai | CrossNER | 45.60 | 50.69 | 50.31 | 52.12 |
crossner_literature | CrossNER | 51.52 | 54.56 | 55.06 | 56.91 |
crossner_music | CrossNER | 65.80 | 68.96 | 63.06 | 64.27 |
crossner_politics | CrossNER | 55.26 | 56.41 | 62.47 | 66.52 |
crossner_science | CrossNER | 56.08 | 60.85 | 58.31 | 55.47 |
few_nerd | General NER | 52.37 | 55.14 | 51.49 | 47.22 |
german_ler | Legal | 21.16 | 11.16 | 22.36 | 6.88 |
hipe2020 | Historical OCR | 45.46 | 39.56 | 41.22 | 29.45 |
mobie | Disaster | 30.65 | 24.44 | 32.47 | 29.66 |
ronec | Multilingual NER | 40.13 | 37.01 | 38.86 | 31.55 |
Fastino Inc. (“Fastino”) develops specialized AI models and provides APIs designed to support structured data extraction, classification, reasoning, and production AI workflows. Fastino is a technology company and does not provide legal, financial, compliance, or advisory services.
Any outputs, predictions, classifications, or decisions generated through Fastino models are based on the configuration, data, and implementation provided by the customer. Fastino does not control, verify, or guarantee the accuracy, completeness, or suitability of model outputs for any specific purpose. By using this website or Fastino’s models and services, you acknowledge that all content and outputs are provided for informational and operational purposes only and agree to our Terms of Use and Privacy Policy.
2026 Fastino Inc.
All rights reserved