Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization
InFlowOp optimizes multi-agent workflows label-free, gaining up to 11.97% over single-agent baselines.
InFlowOp optimizes multi-agent LLM workflows with a label-free cost that balances each agent's competence against its runtime. Before execution it chooses task granularity and assignments from that cost; during execution it applies the cheapest correction using the same cost. The authors introduce Braid, a benchmark whose tasks require coordination beyond a single agent. Across domains and backbones, InFlowOp beats single-agent baselines by up to 11.97% and gains 9.64% from in-flow optimization.
- InFlowOp prices decomposition and agent assignment with one label-free cost.
- It sets granularity before execution and applies the cheapest fault fix in flow.
- Braid requires multi-agent coordination beyond a single agent's capability.
- Gains reach +11.97% over single-agent baselines and +9.64% from in-flow optimization.
Full article228 words · extracted from huggingface.co · click to collapse
Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle up front. Thus, whether each subtask succeeds remains unknown until the workflow runs. Yet, improving a workflow is costly. Locating a fault usually requires a reference answer, a graded outcome, or a trained assessor, and the fix is applied to the whole workflow through re-execution, re-search, or retraining. We propose InFlowOp, which prices every decision in one label-free cost that weighs how well an agent's competence meets what a subtask demands against how much that agent takes to run. Before execution, InFlowOp bidirectionally determines the granularity of task decomposition and agent assignment following from the cost rather than from a fixed template. During execution, InFlowOp corrects a fault with the cheapest move via the same cost that serves the workflow both as it is built and as it runs. Facing the workflow-level evaluation challenge, we introduce Braid, a benchmark whose tasks require multi-agent coordination beyond single-agent capability. Across various domains and backbones, InFlowOp outperforms single agent baselines by up to +11.97%, achieving +9.64% with in-flow optimization. Our project page: https://xhguo7.github.io/InFlowOp/.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.01017