Back

A Practical Guide to Harness Engineering

A Practical Guide to Harness Engineering

A Practical Guide to Harness Engineering

Research

While raw model weights establish an agent’s latent capability, the harness determines how much of that capability is realized in practice. In isolation, a language model is a stateless next-token predictor: it has no eyes, no hands, and no memory across turns. The harness is everything surrounding the model: the software that assembles context, exposes tools, enforces operational invariants, orchestrates control flow, and preserves execution state across multi-turn tasks.


Request

The incoming message from the user.

Environment

The external workspace where the agent executes actions and reads state.

Result

The final deliverable returned to the caller upon turn completion.

Figure 1: Building a harness requires engineering concrete, modular software.

Across identical, frozen weights, the harness alone can swing performance by tens of percentage points. On the Berkeley Function Calling Leaderboard, changing tool presentation alone shifts Claude Opus 4.5’s V4 overall score by 44 percentage points: from 33.5% when functions are described in prompt text to 77.5% when delivered via native schemas. Single-turn call accuracy is roughly equal; the gap sits almost entirely in the multi-turn, web-search, and memory categories. On SWE-bench Verified, replacing a model-driven loop (SWE-agent) with a fixed localize-repair-validate pipeline (Agentless-1.5) lifts the same GPT-4o snapshot from 23.2% to 38.8%. Part of that 15.6-point gain is extra inference: Agentless samples about 40 candidate patches per issue and selects one with regression and reproduction tests. These spreads suggest that downstream capability is not a function of model weights in isolation, but the joint outcome of a model and its harness. With weights frozen, the harness itself becomes an explicit optimization target: a discrete search space over tool interfaces, memory management policies, and control topologies. Figure 2 holds the model fixed and swaps only the harness: across twelve off-the-shelf harnesses, Nemotron 3 Super resolves between 26 and 60 of the same 100 SWE-bench Verified tasks.

Figure 2: One model, twelve off-the-shelf harnesses: tasks resolved against median tokens per task. Orange points are on the Pareto frontier. Nemotron 3 Super on the same 100 SWE-bench Verified tasks, one attempt per task with a 3,000-second budget, all harnesses run side by side on one inference pool. Cursor CLI and Devin cannot target a self-hosted model and are not included.

What follows is an evidence-backed guide to engineering and optimizing this layer. We begin with the foundational design principles that govern reliable harnesses, examine concrete implementation decisions for each harness component (open any component in Figure 1), and conclude with methods for systematically benchmarking and searching the harness space over time.

The Guiding Principles

Reliable execution is rarely achieved through detailed prompt instructions or open-ended autonomy alone. Because language models are probabilistic components operating across multi-step environments, reliability often depends on how the surrounding software is built: how action spaces are scoped, how interfaces are formatted, how invariants are enforced, and how state is preserved. Below we provide four core principles to establishing a reliable harness.

Minimize the decision surface.

Expose high-level abstractions to the model while absorbing operational complexity into the harness. Sequential decision-making in autonomous agents suffers from compounding errors: granting the model more granular commands or unconstrained axes of freedom sharply increases trajectory failure rates (Ross and Bagnell, 2010; Yang et al., 2024; Kwa et al., 2025). While model weights set the base per-step error rate, the harness governs overall reliability by controlling the decision horizon and the cost of error recovery (Ross et al., 2011; Sutton et al., 1999). Without immediate recovery, an initial error compounds across the remaining horizon: it steers the agent into unfamiliar environment states outside its training distribution (Rajaraman et al., 2020), while polluting the prompt with failed attempts that degrade subsequent reasoning (Sinha et al., 2025).

In practice, give the model fewer high-level actions to choose between and execute deterministic subroutines in code rather than delegating them to the model. Instead of exposing raw system primitives, provide focused semantic actions. SWE-agent demonstrates this with the Agent-Computer Interface, outperforming a shell-only agent (18.0% vs. 11.0% on SWE-bench Lite with GPT-4 Turbo) by providing dedicated, line-bounded file-viewing and editing commands. Similarly, deterministic operations such as applying patches, running linters, or formatting code should execute directly in harness software rather than through an LLM.

Don’t fight the prior.

Structure the model’s inputs and outputs to align with the established conventions from its pretraining over less familiar representations. When generating outputs, language models constantly weigh the task instructions in context against the default patterns they absorbed during pretraining (McCoy et al., 2024). Even when a task is fully deterministic, an unfamiliar format carries a low prior and must compete against patterns the model observed millions of times. Model accuracy tracks the log-frequency of a pattern in pretraining data; when a harness introduces a custom domain-specific language or idiosyncratic delimiter syntax, the failure mode is typically one of expression rather than reasoning (Razeghi et al., 2022; Kandpal et al., 2023). The model grasps the required logic but emits malformed syntax, unbalanced tags, or hallucinated arguments (FormatSpread; Sclar et al., 2024).

In practice, expose tool signatures, argument schemas, environment signals, and intermediate representations in conventions the model encountered repeatedly during pretraining. SWE-bench harnesses operate through standard POSIX terminal conventions and exit codes; Aider settled on search/replace blocks and unified diffs for edits, and OpenHands on a string-replacement file editor, rather than bespoke edit syntaxes; and the BFCL V4 spread above is the price of describing functions in prose where a native provider schema would do. Avoid custom domain-specific languages and opaque database identifiers wherever semantic labels exist.

Hardcode constraints and verifications.

Make invalid actions and counterproductive behavior structurally impossible at the interface boundary rather than discouraged by the prompt. Monitors placed at an execution boundary guarantee safety properties by construction, regardless of which policy generates candidate actions (Alshiekh et al., 2018). In contrast, prompt instructions only adjust probabilities; they provide no formal guarantees and remain vulnerable to context shifts, prompt injections, and sampling noise (Wolf et al., 2024). Enforcing rules in code provides guarantees that standard tests can verify. While deterministic checks easily catch immediate violations like path traversal or out-of-range arguments, software cannot judge an agent’s broader intent (Schneider, 2000). Interfaces must therefore be restricted so dangerous actions are impossible by design.

In practice, enforce parameter validation, file allowlists, rate limits, and loop budgets in deterministic application code. For example, Anthropic advises letting code hold the structure and guardrails rather than relying on prompt-level compliance, while 12-Factor Agents similarly argues for keeping loop budgets, parameter validation, rate limits, and permissions in deterministic code. Cursor extends this to the edit loop, running language-server diagnostics on candidate edits and feeding errors back to the agent before changes are handed to the user (see its shadow workspace design).

Externalize variable states.

Persist evolving task information to durable storage instead of letting it accumulate in the context window. Relying on conversation history as a database degrades both efficiency and accuracy. In transformer attention, softmax normalization over an expanding token sequence dilutes the attention allocated to any single item (Veličković et al., 2025), and long contexts show position-dependent retrieval failures such as lost-in-the-middle (Liu et al., 2024). Furthermore, retaining complete histories of failed attempts pollutes the context; models readily condition on their own prior mistakes, repeating failed strategies even after being prompted to correct them (Sinha et al., 2025). Writing intermediate artifacts to disk and reading them back functions as an external scratchpad, enabling fixed-depth networks to execute multi-step serial reasoning across separate calls (Li et al., 2024; Merrill and Sabharwal, 2024).

In practice, treat the prompt window as a volatile cache and the filesystem as the source of truth. Persist roadmaps, task checklists, and intermediate results in structured markdown or JSON files (Claude Code). Isolate execution in disposable sandboxes and git worktrees where modifications remain inspectable and revertible through standard version control (SWE-agent, OpenHands, Devin), and inspect files via paginated slicing rather than dumping full contents into context.

Systematic Optimization

Viewed as a unified system, the harness becomes an explicit optimization landscape: a structured search space across prompts, memory policies, tool schemas, and control topologies. Systematically improving this architecture requires addressing three distinct challenges: defining where mutable leverage lies across the search space, enforcing code-level isolation to prevent self-grading loops, and applying disciplined experimental guardrails to distinguish architectural gains from compute scaling.

Search space

Weng (2026) traces a progression in what gets optimized, from instructions through structured context and workflows to harness code and optimizer code. We find it useful to read this as a ladder of expressivity and reversibility. At the lowest rung sit system instructions and exemplars, which are fast and safe to mutate but offer a low structural ceiling. Higher up sit memory compaction policies, workflow topology, and harness code (tools and execution middleware), where mutations produce inspectable software diffs with explicit syntax boundaries.

On Terminal-Bench 2, AHE’s ablations attribute its gains to tools, middleware, and long-term memory rather than the system prompt, which on its own slightly hurt performance (−2.3 points) (AHE; Lin, J. et al., 2026). Separately, the quality of harness updates (skills, prompts, and memories) appears largely flat across the scale of the model writing them: updates written by a 9B model yield gains comparable to those from Opus 4.6 (Lin, M. et al., 2026), which suggests cheaper models may suffice as optimizers.

Isolation

Automated harness optimization introduces an acute risk of self-grading failure. An optimizer evaluating its own harness modifications is a generator attempting to verify itself. If the optimizer has write access to the evaluation suite, its own reasoning budget, or the unit tests, the path of least resistance to a higher benchmark score is reward hacking: weakening assertions, deleting edge cases, or expanding compute allocations.

To preserve evaluation integrity, the harness architecture must enforce a strict read-only boundary in code. The evaluation suite, ground-truth oracles, reasoning budgets, and optimizer code must remain permanently immutable to candidate edits. Any mutation that touches files outside the designated execution boundary must trigger immediate rejection.

Guardrails

Automated and manual harness iterations alike require experimental discipline to avoid mistaking variance or compute scaling for architectural progress. Four guardrails keep harness development grounded. First, establish diagnostic prerequisites: never deploy automated search until human engineers have inspected execution histograms and can attribute recurring failures to specific components. Second, enforce budget-matched null tests: because repeated sampling expands solution coverage (Brown et al., 2024), any proposed harness topology must outperform a budget-matched best-of-N plain baseline using the exact same inference spend before being declared superior (Kapoor et al., 2024). Third, maintain multi-dimensional accounting: track per-task token volume, dollar cost, and wall-clock latency alongside accuracy, as benchmark gains that triple inference spend can represent net regressions. Finally, enforce stopping rules to prevent proxy saturation: optimizing aggressively against any fixed benchmark oracle eventually triggers Goodhart’s Law, degrading real-world performance even as benchmark scores climb (Gao et al., 2023).

Summary

Agent capability is the joint product of model weights and harness design. While foundational weights define the latent potential of an agent, the harness determines whether that potential translates into reliable execution.

Individual prompt tricks inevitably expire as capabilities are internalized into future model weights. However, the core design principles of harness engineering, including abstracting complexity behind simple interfaces, aligning representations with pretraining distributions, enforcing safety invariants in code, and externalizing state to durable storage, remain enduring invariants of robust autonomous systems. Before committing to complex scaffolds or unconstrained search loops, measure against simple baselines, isolate your failure modes through tiered telemetry, and let deterministic software hold the structure.

Subscribe to our newsletter

Get the latest updates on model releases, product news, and research.

Fastino Inc. (“Fastino”) develops specialized AI models and provides APIs designed to support structured data extraction, classification, reasoning, and production AI workflows. Fastino is a technology company and does not provide legal, financial, compliance, or advisory services.

Any outputs, predictions, classifications, or decisions generated through Fastino models are based on the configuration, data, and implementation provided by the customer. Fastino does not control, verify, or guarantee the accuracy, completeness, or suitability of model outputs for any specific purpose. By using this website or Fastino’s models and services, you acknowledge that all content and outputs are provided for informational and operational purposes only and agree to our Terms of Use and Privacy Policy.

2026 Fastino Inc.

All rights reserved

Subscribe to our newsletter

Get the latest updates on model releases, product news, and research.

Fastino Inc. (“Fastino”) develops specialized AI models and provides APIs designed to support structured data extraction, classification, reasoning, and production AI workflows. Fastino is a technology company and does not provide legal, financial, compliance, or advisory services.

Any outputs, predictions, classifications, or decisions generated through Fastino models are based on the configuration, data, and implementation provided by the customer. Fastino does not control, verify, or guarantee the accuracy, completeness, or suitability of model outputs for any specific purpose. By using this website or Fastino’s models and services, you acknowledge that all content and outputs are provided for informational and operational purposes only and agree to our Terms of Use and Privacy Policy.

2026 Fastino Inc.

All rights reserved

Subscribe to our newsletter

Get the latest updates on model releases, product news, and research.

Fastino Inc. (“Fastino”) develops specialized AI models and provides APIs designed to support structured data extraction, classification, reasoning, and production AI workflows. Fastino is a technology company and does not provide legal, financial, compliance, or advisory services.

Any outputs, predictions, classifications, or decisions generated through Fastino models are based on the configuration, data, and implementation provided by the customer. Fastino does not control, verify, or guarantee the accuracy, completeness, or suitability of model outputs for any specific purpose. By using this website or Fastino’s models and services, you acknowledge that all content and outputs are provided for informational and operational purposes only and agree to our Terms of Use and Privacy Policy.

2026 Fastino Inc.

All rights reserved