ExpHarness

ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness

Tao Feng1 Chongrui Ye1 Fangxu Yu2 Tianyang Luo1 Jingjun Xu3 Xueqiang Xu1 Haozhen Zhang4 Weizhi Zhang5 Zijie Lei6 Zhigang Hua6 Yan Xie6 Shuang Yang6 Jiaxuan You1
1University of Illinois Urbana-Champaign   2University of Maryland   3Harvard University
4Nanyang Technological University   5University of Illinois Chicago   6Meta Monetization AI
+12.1%
Small executor
+4.5%
Large executor
ExpSuite-Static
Relative gain in weighted average score over the strongest baseline (MemRL) on 10 benchmarks.
+21.4%
Small executor
+12.7%
Large executor
ExpSuite-Agentic
Relative gain in average score over the strongest baseline (S3) on ALFWorld and AppWorld.
−21.6%
Interaction Steps
vs. the most step-efficient baseline (large executor).
0
Executor Updates
The executor stays frozen; only the harness learns.

Small / large executors. ExpSuite-Static: Llama-3.2-3B-Instruct / Llama-3.1-8B-Instruct. ExpSuite-Agentic: Qwen3-32B / Gemini-3.1-Flash-Lite. Gains are relative improvements (Sours − Sbaseline) / Sbaseline of the averaged score.

Overview

Let the Harness Learn, Keep the Executor Frozen

LLM agents increasingly operate within a harness—the scaffolding that determines what enters the executor's context—yet the experience they accumulate across tasks rarely flows back into this harness. Fine-tuning the executor on collected experience ties what was learned to one model, and has to be repeated whenever a stronger or more suitable executor appears.

ExpHarness is a learnable experience harness that improves frozen and replaceable LLM executors without modifying their parameters. It distills trajectories into reusable skills and failure lessons within a self-evolving experience graph, and trains a lightweight retrieval copilot that decides, per task, how broadly to explore the graph and how strongly to favor historically useful experiences over merely similar ones. The copilot is optimized with reinforcement learning from a utility-grounded reward, and the same reward updates the graph during training.

Self-Evolving Experience Graph

Successful trajectories become skills, failed ones become lessons. Each entry stores its text, embedding, a utility estimate and a retrieval count, and is linked to its semantically similar neighbors.

Trainable Retrieval Copilot

A small LM (Qwen2.5-3B-Instruct) reads the task and outputs two controls: R sets how far graph diffusion spreads from the seeds, W trades semantic similarity for historical utility. It never answers the task itself.

Utility-Grounded Co-Evolution

The executor is run with and without the retrieved experiences. The score gain plus the absolute score trains the copilot with PPO and updates the utility of every retrieved node.

Method

Copilot-Controlled, Utility-Guided Retrieval

ExpHarness framework: a retrieval copilot predicts R and W for each task; retrieval on the experience graph runs semantic seeding, personalized-PageRank diffusion and utility-aware ranking; the frozen executor is evaluated with and without memory; the reward updates both the copilot and the graph.
Overview of ExpHarness. For each task, the retrieval copilot predicts R (diffusion depth) and W (similarity–utility trade-off). Retrieval on the experience graph runs (a) semantic seeding, (b) personalized-PageRank diffusion and (c) utility-aware ranking. The frozen executor is evaluated with and without the retrieved experiences; the resulting reward updates the copilot via PPO and the graph via node utility updates, new experiences, and pruning. Only the harness evolves; the executor stays frozen.
Step (a)
Semantic Seeding

Embed the task with Contriever and take the top-m most similar experiences as seeds.

Step (b)
Structure-Aware Expansion

Personalized PageRank from the seeds. A larger R means a higher restart probability, so retrieval stays local; a smaller R lets it spread across the graph.

Step (c)
Utility-Aware Ranking

Rank candidates by (1−λ)·similarity + λ·UCB(utility, count), with λ = W/100, and pass the top-K to the executor.

r = ( swith − swithout ) + η · swith Utility-grounded reward: the executor's gain from the retrieved experiences plus its absolute score. It trains the copilot (PPO) and updates uv ← (1−β)uv + βr for every retrieved node.

Benchmark

ExpSuite: 10 Static Benchmarks + 2 Agentic Environments

ExpSuite tests whether experience reuse helps both single-turn reasoning and long-horizon interaction. ExpSuite-Static uses Llama-3.2-3B-Instruct and Llama-3.1-8B-Instruct as executors; ExpSuite-Agentic uses Qwen3-32B and Gemini-3.1-Flash-Lite. All executors stay frozen and no manually curated few-shot examples are used.

Question Answering 5
ARC-C · CommonsenseQA · GPQA-Diamond · MMLU · OBQA
Metric: accuracy
Math Reasoning 3
GSM8K · GSM-Symbolic · MATH
Metric: exact match
Code Generation 2
HumanEval+ · MBPP+
Metric: Pass@1
ALFWorld 2
Seen · Unseen splits
Metric: success rate, #steps
AppWorld 2
Test-Normal · Test-Challenge
Metric: pass rate, #steps

Main Results

Highest Weighted Average in Every Evaluated Setting

We compare against a no-memory executor, retrieval-centric experience learning methods (ReasoningBank, ExpeL, LightMem, Mem0, AWM, MemRL), LLM-centric methods (IRCoT, Search-o1, S3) and, for the agentic setting, prompt-based agents (ReAct, Reflexion). All experience-based baselines receive the same historical trajectories whenever applicable.

ExpSuite-Static

MethodQuestion Answering ReasoningCoding Weighted
Avg.
ARC-CCSQAGPQA-DMMLUOBQA GSM8KGSM-SymMATHHumanEval+MBPP+
No Memory51.5654.4418.3342.8954.2269.5660.4437.5643.5938.7551.88
Retrieval-Centric Experience Learning Baselines
ReasoningBank71.7861.5621.6754.0072.6770.4456.4445.5651.2857.5060.83
ExpeL44.6757.7816.6746.4443.1177.1180.4423.3320.5122.5051.49
LightMem71.5664.0026.6758.0072.8944.8942.0049.1148.7242.5056.47
Mem068.0059.5628.3356.4465.7852.0039.7828.2241.0360.0052.42
AWM72.4460.4415.0053.3369.3355.3338.8945.7853.8560.0055.81
MemRL67.3360.0023.3340.2267.1183.7868.0055.3346.1557.5062.06
LLM-Centric Experience Learning Baselines
IRCoT62.6759.7815.0038.0063.3377.7872.6729.5651.2850.0056.65
Search-o166.6760.2228.3350.0062.4478.4473.3338.4430.7728.7559.63
S364.8959.1123.3344.0064.0082.0074.0053.1148.7255.0061.94
ExpHarness (Ours)74.0063.1126.6760.0074.0084.8982.0057.1156.4162.5069.57

Accuracy (QA), exact match (reasoning) and Pass@1 (coding), in %. The weighted average weights each benchmark by its number of test instances. Bold: best; underline: second best.

ExpSuite-Agentic

MethodALF-SeenALF-Unseen AppWorld Test-NAppWorld Test-C Average
SR#Steps↓SR#Steps↓ PR#Steps↓PR#Steps↓ Score#Steps↓
No Memory19.341.235.842.530.130.821.631.525.134.7
Prompt-based Agentic Baselines
ReAct31.436.747.038.630.032.221.831.928.933.8
Reflexion23.640.640.338.731.631.321.932.726.934.6
Retrieval-Centric Experience Learning Baselines
ReasoningBank37.935.326.142.629.425.622.224.326.829.2
ExpeL54.327.529.141.426.735.128.225.632.330.2
LightMem49.229.625.039.228.534.223.737.529.035.8
Mem030.037.318.742.820.833.415.735.219.536.4
AWM37.935.332.840.128.028.822.833.027.833.7
MemRL25.740.320.249.645.218.628.924.730.229.9
LLM-Centric Experience Learning Baselines
IRCoT33.637.927.640.652.811.436.117.837.623.4
Search-o132.138.011.247.345.416.526.514.228.723.7
S358.623.542.522.555.49.735.015.044.016.5
ExpHarness (Ours)70.020.075.418.057.18.839.313.653.414.4

Success rate (SR) on ALFWorld and pass rate (PR) on AppWorld, in %, with the average number of environment interaction steps (lower is better). The average is weighted across all test tasks.

ExpHarness improves the weighted average over the strongest baseline by 12.1% / 4.5% on ExpSuite-Static (vs. MemRL) and by 21.4% / 12.7% on ExpSuite-Agentic (vs. S3) with the small / large executor, while using 12.7% / 21.6% fewer interaction steps than the most step-efficient baselines. Gains are larger for the smaller executors and for the interactive tasks.

Transfer

The Learned Harness Transfers Across Executors

Because all learning resides in the harness, a harness trained with one executor can be reused when the executor is replaced. We transfer the experience graph only, the copilot only, or both, and compare with ExpHarness trained directly on the target executor. For the reasoning setting we transfer from Llama-3.1-8B-Instruct to DeepSeek-R1-Distill-Llama-8B (static) and from Gemini-3.1-Flash-Lite to Claude-Sonnet-4 (agentic).

Radar chart: small-to-large executor transfer
(a) Small → large
Radar chart: large-to-small executor transfer
(b) Large → small
Radar chart: non-reasoning-to-reasoning executor transfer
(c) Non-reasoning → reasoning
Harness transfer across executor shifts. Graph+Copilot transfer (the full harness) is closest to target-specific training in the small-to-large and non-reasoning-to-reasoning settings and is the best transfer variant in most domains for large-to-small transfer, where a gap to target-specific training remains. Panels use different axis ranges.

Ablation

Every Retrieval Design Choice Contributes

We ablate four experience-retrieval design choices across QA, Reasoning, Coding, ALFWorld and AppWorld: w/o Similarity Filtering (no near-duplicate filtering at insertion), Flat Experience (a flat pool retrieved purely by similarity), w/o Graph Diffusion (retrieve only from the semantic seeds) and w/o Utility Ranking (rank by similarity only). Every variant falls below the full method.

Bar chart of ExpHarness versus four ablation variants across QA, Reasoning, Coding, ALFWorld and AppWorld
Performance under four ablation settings. Flat Experience shows the largest drop (it changes both candidate generation and ranking); removing graph diffusion or utility ranking also lowers scores, while removing similarity filtering causes a smaller decline.

Code & Data

Get Started

The release contains the experience graph server, the copilot PPO trainer (built on verl), executors for ALFWorld, AppWorld and the static benchmarks, processed data splits and the cold-start experience graphs. Executors are plain API calls (Gemini, NVIDIA NIM, any OpenAI-compatible endpoint such as vLLM, or Claude).

# install
git clone https://github.com/ulab-uiuc/ExpHarness && cd ExpHarness
pip install torch==2.4.0 vllm==0.6.3 && pip install -r requirements.txt

# train the retrieval copilot (the executor stays frozen)
bash run/alfworld/train.sh
bash run/appworld/start_servers.sh && bash run/appworld/train.sh
bash run/static/train.sh all            # all | reasoning | coding

# evaluate a checkpoint together with its evolved graph
bash run/alfworld/eval.sh checkpoints/expharness_alfworld_42 300

BibTeX

@article{feng2026expharness,
  title  = {ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness},
  author = {Feng, Tao and Ye, Chongrui and Yu, Fangxu and Luo, Tianyang and Xu, Jingjun and Xu, Xueqiang and Zhang, Haozhen and Zhang, Weizhi and Lei, Zijie and Hua, Zhigang and Xie, Yan and Yang, Shuang and You, Jiaxuan},
  year   = {2026}
}