Published Research

Published Research

MAY 20, 2026 • RESEARCH

Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers

Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers

Nikhil Nayak, Julia White, Urchade Zaratiana, Kelton Zhang, Henrijs Princis, Dhruv Atreja, Henry Fawcett, Matthew Thomas, George Hurn-Maloney, and Ash Lewis

Preconditioned optimizers like AdamW, Sophia, and Shampoo hide two distinct sources of bias in how they're estimated from minibatches. The paper introduces a lightweight, single-batch correction that fixes both, measurably improving language model training across optimizer families.

Preconditioned optimizers like AdamW, Sophia, and Shampoo hide two distinct sources of bias in how they’re estimated from minibatches. The paper introduces a lightweight, single-batch correction that fixes both, measurably improving language model training across optimizer families.

Read more

Read more

MAY 11, 2026 • MODEL

GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction

GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction

Urchade Zaratiana, Ash Lewis, and George Hurn-Maloney

Fastino Labs builds GLiNER2-PII, a compact 0.3B-parameter model that detects personally identifiable information across 42 entity types in messy, real-world text. It achieves the highest span-level F1 on the SPY benchmark among five compared systems, including OpenAI's Privacy Filter.

Fastino Labs builds GLiNER2-PII, a compact 0.3B-parameter model that detects personally identifiable information across 42 entity types in messy, real-world text. It achieves the highest span-level F1 on the SPY benchmark among five compared systems, including OpenAI’s Privacy Filter.

Read more

Read more

MAY 8, 2026 • MODEL

Accepted to COLM 2026

GLiGuard: Schema-Conditioned Classification for LLM Content Moderation

GLiGuard: Schema-Conditioned Classification for LLM Content Moderation

Urchade Zaratiana, Mary Newhauser, George Hurn-Maloney, and Ash Lewis

GLiGuard, a 0.3B-parameter bidirectional encoder, moderates LLM prompts and responses in a single forward pass, evaluating safety, harm categories, and jailbreak strategies at once instead of generating classifications token by token. It matches decoder-based guard models 23 to 90 times its size, with up to 16x higher throughput.

GLiGuard, a 0.3B-parameter bidirectional encoder, moderates LLM prompts and responses in a single forward pass, evaluating safety, harm categories, and jailbreak strategies at once instead of generating classifications token by token. It matches decoder-based guard models 23 to 90 times its size, with up to 16x higher throughput.

Read more

Read more

APRIL 10, 2026 • RESEARCH

Accepted to EMNLP 2026

Pioneer Agent: Continual Improvement of Small Language Models in Production

Pioneer Agent: Continual Improvement of Small Language Models in Production

Dhruv Atreja, Julia White, Nikhil Nayak, Kelton Zhang, Henrijs Princis, George Hurn-Maloney, Ash Lewis, and Urchade Zaratiana

Pioneer Agent fine-tunes small language models after deployment, automatically diagnosing production failures and retraining to fix them without degrading existing behavior. It substantially improves performance and discovers effective training techniques on its own based on feedback from inference logs.

Pioneer Agent fine-tunes small language models after deployment, automatically diagnosing production failures and retraining to fix them without degrading existing behavior. It substantially improves performance and discovers effective training techniques on its own based on feedback from inference logs.

Read more

Read more

OCTOBER 22, 2025 • RESEARCH

Beyond Reactivity: Measuring Proactive Problem solving in LLM Agents

Beyond Reactivity: Measuring Proactive Problem solving in LLM Agents

Gil Pasternak, Dheeraj Rajagopal, Julia White, Dhruv Atreja, Matthew Thomas, George Hurn-Maloney, and Ash Lewis

Fastino Labs creates PROBE, a benchmark that measures whether LLM agents can proactively spot and resolve problems without being told to, decomposing the task into searching a personalized datastore, identifying the real bottleneck, and executing the correct fix across 1,000 realistic scenarios.

Fastino Labs creates PROBE, a benchmark that measures whether LLM agents can proactively spot and resolve problems without being told to, decomposing the task into searching a personalized datastore, identifying the real bottleneck, and executing the correct fix across 1,000 realistic scenarios.

Read more

Read more

JULY 24, 2025 • MODEL

GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface

GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface

Urchade Zaratiana, Gil Pasternak, Oliver Boyd, George Hurn-Maloney, and Ash Lewis

GLiNER2 is a unified open-source 205M-parameter model that handles named entity recognition, text classification, and hierarchical structured extraction through a single schema-driven interface, matching GPT-4o's accuracy on key benchmarks while running 2.6x faster on CPU.

GLiNER2 is a unified open-source 205M-parameter model that handles named entity recognition, text classification, and hierarchical structured extraction through a single schema-driven interface, matching GPT-4o’s accuracy on key benchmarks while running 2.6x faster on CPU.

Read more

Read more

JUNE 21, 2024 • MODEL

GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer

GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer

Urchade Zaratiana, Nadi Tomeh, Pierre Holat, and Thierry Charnois

GLiNER is a compact bidirectional transformer model that identifies any entity type in a zero-shot setting, extracting entities in parallel rather than generating tokens sequentially like LLMs. It outperforms frontier LLMs on zero-shot NER benchmarks while remaining small and fast enough for resource-constrained environments.

GLiNER is a compact bidirectional transformer model that identifies any entity type in a zero-shot setting, extracting entities in parallel rather than generating tokens sequentially like LLMs. It outperforms frontier LLMs on zero-shot NER benchmarks while remaining small and fast enough for resource-constrained environments.

Read more

Read more

Fine-tune open weight models with the Fastino Fine-Tuning Agent

Deploy production-grade models that you own in hours, not weeks.

Get access

Fastino Inc. (“Fastino”) develops specialized AI models and provides APIs designed to support structured data extraction, classification, reasoning, and production AI workflows. Fastino is a technology company and does not provide legal, financial, compliance, or advisory services.

Any outputs, predictions, classifications, or decisions generated through Fastino models are based on the configuration, data, and implementation provided by the customer. Fastino does not control, verify, or guarantee the accuracy, completeness, or suitability of model outputs for any specific purpose. By using this website or Fastino’s models and services, you acknowledge that all content and outputs are provided for informational and operational purposes only and agree to our Terms of Use and Privacy Policy.

2026 Fastino Inc.

All rights reserved