Harness Engineering for Software Engineering via Modular Executable Dev-Primitives
HERMES uses modular Dev-Primitives to beat baseline software-engineering agent harnesses by 12.4% on four benchmarks.
The paper introduces Dev-Primitives, pairing repository artifacts with resident LLMs so components can reason, communicate, and self-modify. HERMES activates those primitives at repository scale using dependency-aware activation and a bug-diagnosis mechanism that maps execution evidence to components that need revision. On four software-engineering benchmarks, HERMES beats matched baseline harnesses by 12.4% on average. With strong activation and diagnosis models, Qwen3-8B Dev-Primitives stay within 4.5% of a homogeneous GPT-5.6 Sol setup while cutting inference cost 26.2% on Terminal-Bench 4.0.
- Dev-Primitives pair each repo artifact with a resident LLM interface.
- HERMES activates primitives by dependency and maps execution evidence to buggy components.
- Average gain is 12.4% over matched baseline harnesses on four benchmarks.
- Qwen3-8B primitives stay within 4.5% of GPT-5.6 Sol and cut Terminal-Bench inference cost 26.2%.
Full article233 words · extracted from huggingface.co · click to collapse
Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce Dev-Primitives (Development Primitives), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose HERMES, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4\% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5\% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2\% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.07832