Dynamic Harness Search: Building Multi-Agent Systems Per-Query via Prediction
SHIFT searches per-query agent harnesses and beats 17 baselines by 7.2 accuracy points.
SHIFT builds a multi-agent harness per query by learning a policy over harness-building actions and a value function that predicts accuracy versus execution cost, then using Monte Carlo tree search. Across 9,193 tasks in six benchmarks, with a Gemini 3.5 Flash executor, it reaches about 80% mean accuracy and exceeds the strongest of 17 prompting, prompt-optimization, and workflow-search baselines by 7.2 percentage points. A cheaper mode still beats every baseline while using 32% fewer execution tokens than the strongest baseline. Choosing structure, instructions, and tools jointly beats choosing only instructions or only tools by up to 9.1 points.
- SHIFT predicts harness utility so search avoids executing every alternative.
- Across 9,193 tasks it reaches about 80% mean accuracy with Gemini 3.5 Flash.
- It beats the strongest of 17 baselines by 7.2 percentage points.
- A cheaper mode uses 32% fewer execution tokens than that baseline.
- Joint structure, instruction, and tool choice gains up to 9.1 points.
Full article217 words · extracted from huggingface.co · click to collapse
Agent harnesses specify the roles, instructions, tools, and communication structure used to solve a task, and the right harness depends on the query. Because the value of each design choice is observable only through execution, tailoring a harness to each query has required either executing alternatives at inference time or costly manual design. We introduce SHIFT, which moves execution out of the per-query search loop. A local LLM architect learns a policy over harness-building actions from search, and a value function that predicts, from measured executions, a utility balancing accuracy against execution cost. For each query, Monte Carlo tree search uses these predictions to construct a harness. Across 9,193 tasks in six benchmarks, from math to document and general-assistant tasks, with a Gemini 3.5 Flash executor, SHIFT attains the highest mean accuracy, about 80%, outperforming 17 baselines that span prompting, prompt optimization, and workflow search, and exceeding the strongest baseline by 7.2 percentage points. A cheaper mode of SHIFT also attains a higher mean accuracy than every baseline while using 32% fewer execution tokens than the strongest baseline. We further show that choosing structure, instructions, and tools jointly beats choosing only instructions or only tools by up to 9.1 percentage points, and that learned value selection identifies more accurate harnesses with lower execution cost from candidate pools.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.04137