Between a model and its final output sits an important layer of infrastructure: its tools, context management, execution loop, and output checks. This surrounding infrastructure is called harness, and its design can have a substantial impact on agent performance. For example, OpenAI’s Astra improved benchmark performance by almost 40 percentage points, simply by changing the harness.
Adaption Labs’ recent work, “A Better Harness Can Unlock Smaller Models,” further demonstrates the value of harness design and, importantly, shows how effective harnesses can be systematically built. Specifically, the work illustrates:
- A task-specific harness can outperform frontier general-purpose harnesses such as Prime Agent and OpenCode.
- Effective harnesses can be designed systematically by analyzing model failures, identifying recurring patterns, and translating those patterns into targeted improvements.
A good harness can therefore unlock substantial capability from smaller models without changing the models themselves. But designing one still requires considerable manual analysis and iteration.
This raises a natural question: can we automate harness design end-to-end? Adaptive Harness is our attempt to do exactly that.
Adaptive Harness
To clarify, a harness does not give a model new capabilities; instead, it helps the model make better use of the capabilities it already has. This suggests a simple approach to automatic harness design: identify where the model consistently fails, then design targeted tools and interventions to compensate for those weaknesses. By removing these bottlenecks, more of the model’s existing capabilities can translate into successful task performance.
As a prototype, Adaptive Harness takes a model and an existing harness and iteratively improves the harness against a given task benchmark. We start from a baseline harness to reflect real-world deployment: production agents rarely consist of a model API in isolation, and typically have an existing system of tools, context, and execution logic around the model. We also assume access to a task benchmark with clear train / dev / test splits.
Step 1: Train data. The model attempts a fresh batch of tasks from the train split using the current harness, and its full execution trajectories are recorded. These trajectories capture not only whether the model succeeds or fails, but also how it arrives at each outcome.
Step 2: Optimizer LLM. An optimizer LLM, typically a frontier model, is given the trajectories of failed tasks and asked to identify recurring failure modes and modify the harness to address them. To prevent benchmark leakage, the optimizer has no shell or web access and can only modify files within the harness directory. It also retains a history of previous edits and their outcomes, allowing it to avoid repeatedly trying changes that have already failed.
Step 3: Evaluator. The new harness is evaluated on a held-out dev split to determine whether the changes improve performance and generalize beyond the training tasks. The system accepts the new harness only if it improves performance while remaining within a predefined efficiency budget. The process then returns to Step 1 and continues iteratively.
Step 4: Evaluation. Once the optimization loop ends, the best harness (current_harness) is evaluated on the held-out test split. The original harness is evaluated using the same model and settings, so the comparison isolates the effect of harness optimization.
Step 5: Output. Adaptive Harness returns the optimized harness, including the prompts and code that define the surrounding agent system.
Results
Adaptive Harness substantially improves model performance while using far fewer tokens than a frontier general-purpose harness.
Performance. On AppWorld, a benchmark for evaluating agents on multi-step tasks, Adaptive Harness improves Qwen3.5-9B from 19.6% to 49.4%, a 2.5× increase over the baseline harness. For Qwen3.5-27B, performance increases from 35.1% to 72.0%, roughly 2× the baseline. Adaptive Harness is also competitive with the frontier OpenCode harness: it substantially outperforms OpenCode on the 9B model (49.4% vs. 28.0%) and approaches its performance on the 27B model (72.0% vs. 76.8%).
Efficiency. Adaptive Harness achieves these improvements while using far fewer tokens. For the 9B model, Adaptive Harness uses 27.9M input tokens, 63% fewer than OpenCode’s 75.0M. For the 27B model, it uses 25.5M input tokens, 43% fewer than OpenCode’s 45.0M. In both cases, Adaptive Harness stays close to the token usage of the baseline harness.
Limitations
Limited benchmark and model coverage. Our experiments are limited to AppWorld and two models from the Qwen3.5 family. While Adaptive Harness itself is not specific to either, we have not yet shown that the improvements generalize to other model families or domains, such as legal and financial tasks.
Small and noisy dev set. Our dev set contains only 20 tasks, meaning that a few random changes in task outcomes can significantly affect the score. Using a larger dev set, or requiring a clearer improvement before accepting an edit, could make the optimization more reliable, but would also increase the cost of each round.
Unknown trade-offs across tasks. Adaptive Harness deliberately specializes a harness for a specific task distribution. We have not yet tested whether improving performance on that task comes at the cost of worse performance elsewhere. A task-specific harness may therefore outperform a general-purpose harness on its target task while becoming less effective on unrelated tasks. Testing this trade-off would help determine how narrowly an optimized harness should be deployed.
Limited optimization scope. In this prototype, Adaptive Harness only modifies the system prompt and agent-loop code. Other parts of the harness, such as tool interfaces, memory, and context-compaction mechanisms, remain fixed. This means the system is not yet searching over the full harness design space, and further gains (or degradation) may be possible by allowing these components to be optimized as well.
Path to adoption
Adaptive Harness is still a research prototype. To make it useful in practice, we also need to consider what deployment would look like for real users. Imagine an enterprise that wants to run models on-premise, but has limited GPU capacity and cannot deploy frontier-scale models. Instead, it deploys a smaller model and uses Adaptive Harness to get more performance from it.
Let’s assume Adaptive Harness generalizes beyond AppWorld and the limitations above can be addressed. Four practical challenges still remain:
User input as a benchmark
Adaptive Harness needs a benchmark, but most users will not already have one. AppWorld gives us tasks, graders, and train / dev / test splits out of the box. A real team may instead have only a task description, a few examples, and historical traces from its workflow.
A practical solution is to build the benchmark from this existing data. Real examples can be expanded with synthetic variants, while the system generates a small preview set of tasks and graders for the user to review before creating the full benchmark. Graders should rely on deterministic outcome checks—such as database state, API outputs, unit tests, or business rules—whenever possible. Existing systems such as Adaption Labs’ Invent a Dataset could support this stage.
Affordable optimization
Adaptive Harness is only practical if the cost of optimization remains reasonable. Each round requires running the model across train and dev tasks and using a frontier model to analyze failures and edit the harness.
A useful approach is to treat evaluation as a funnel: test edits on a small dev subset first, reject weak changes early, and run full evaluation only on promising ones. Repeated context should be cached, inference should run near optimal concurrency, and previous trajectories should be reused where possible. The goal is not the cheapest run, but the cheapest evaluation that still produces trustworthy decisions.
Harness initialization
The final result depends partly on where Adaptive Harness starts. We used AppWorld’s ReAct-style baseline because it already supported the benchmark environment, but real users may have a stronger production agent, a task-specific harness, or no harness at all.
A sensible hierarchy is: start from the user’s existing production harness when available; otherwise use a task-adjacent harness, then a strong general-purpose harness such as OpenCode, then a harness transferred from a similar previous Adaptive Harness run. Generating a harness from scratch should be the fallback. This reduces the amount of optimization spent rediscovering basic design choices.
Harness routing
A harness optimized for one task may not remain strong on unrelated tasks. Adaptive Harness deliberately specializes around a target task distribution, so replacing a general-purpose harness globally could improve one workflow while hurting another.
A better deployment strategy is to route each request to a harness optimized for that type of task, with a general-purpose harness as fallback. Each specialized harness should also be tested on a broader regression set before deployment. This turns Adaptive Harness from a system that searches for one universal harness into one that maintains and routes between multiple specialized harnesses.
Together, these changes would turn Adaptive Harness from a benchmark optimization loop into a practical workflow: describe the task, build a reliable benchmark, choose the strongest available starting harness, optimize within a controlled budget, and route requests to the harness best suited for each task.
Acknowledgements
Thanks to Adaption Labs’ post A Better Harness Can Unlock Smaller Models for the insights that enabled this project. The project also builds on AppWorld (Trivedi et al., 2024) and its baseline agent, Qwen’s open models, and OpenCode as a point of comparison.
The full implementation is available on GitHub.
Appendix
Setup
| Setting | Value |
|---|---|
| Models | Qwen3.5-9B (Modal L40S) and Qwen3.5-27B (Modal H100), served with SGLang, thinking off |
| Sampling | Temperature 0, replies capped at 1,500 tokens, at most 50 steps per task, 32k context |
| Starting harness | AppWorld’s simplified ReAct code agent and its prompt |
| Optimizer | Claude Opus 4.8 in Claude Code (headless); edits limited to the harness folder; no shell or web access |
| Splits | Train: a fresh 15 tasks per round. Dev: a fixed 20 tasks. Test: AppWorld’s test_normal, 168 tasks, run once |
| Keep rule | Dev pass@1 must rise (ties broken by unit tests passed); input tokens may rise at most 20% |
| Rounds | 3 per model |
| Statistics | Exact McNemar test on paired task outcomes; 95% bootstrap intervals over tasks and over scenarios |
What the loop tried
Qwen3.5-9B
| Round | Edit proposed from training failures | Dev pass@1 | Verdict |
|---|---|---|---|
| 0 | AppWorld ReAct baseline | 15% | Start |
| 1 | Don’t return an answer on action tasks | 45% | Kept |
| 2 | A step-by-step recipe for reading every page of API results | 45% | Rejected (fewer unit tests passed) |
| 3 | Read API fields by their exact names; sanity-check filters | 20% | Rejected |
Qwen3.5-27B
| Round | Edit proposed from training failures | Dev pass@1 | Verdict |
|---|---|---|---|
| 0 | AppWorld ReAct baseline | 20% | Start |
| 1 | Answer only explicit questions; never a confirmation or summary | 55% | Kept |
| 2 | Loop guard: warn when the model repeats identical code, stop after 3 repeats | 70% | Kept (noise: the guard never fired on dev) |
| 3 | Always page through search and list results (APIs return 5 items by default) | 80% | Kept (within noise) |
Related work
Automatic harness optimization became an active area in 2026. DeepMind’s AutoHarness (Lou et al., 2026), an unrelated system, has Gemini write code harnesses that stop game-playing agents from making illegal moves, which lets the smaller Gemini Flash beat Gemini Pro. AutoSaddler learns harness updates from agent execution traces, HarnessOpt-Bench evaluates LLMs as harness optimizers when evaluation is expensive and noisy, and Harness-R1 uses reinforcement learning to train a dedicated harness editor. This project is closest to Adaption Labs’ hand-tuning work. Its focus is small open models, a public benchmark with fixed splits and a test set used once, paired significance tests, and a breakdown of what the gains actually are.