The Delegation Blind Spot: Auditing Product Decisions from Agent Choices
Audit finds agent execution accuracy does not identify which product improvement a user would value.
The paper presents a decision-specific audit that maps a declared observation channel and product-value contrast to compatible intervals and witness populations. A frozen experiment made 4,800 requests to two pinned model snapshots; all 36 conservative primary intervals remained unresolved despite different execution accuracy. An exploratory 2,400-call follow-up resolved three of nine comparisons per model, while a deterministic extractor resolved seven of nine without model calls. Further 14,400 multinomial simulations distinguish structural ambiguity from weak identification, and the authors propose a source-labeled decision receipt. The study included no human participants or real customer outcomes.
- 4,800 calls left all 36 primary decision intervals unresolved.
- A deterministic extractor resolved seven of nine comparisons without model calls.
- 14,400 simulations separate ambiguity from weak identification.
- Study used synthetic tasks and no real customer outcomes.
Full article189 words · extracted from arxiv.org · click to collapse
Successful agent execution need not identify which future product improvement its user would value. We present a decision-specific audit that maps a declared observation channel and product-value contrast to compatible intervals and witness populations. Its foundations are established identification and decision theory; the contribution is an executable measurement workflow and a controlled study of its limits. A frozen experiment makes 4,800 requests to two pinned model snapshots on shared synthetic tasks. All 36 conservative primary intervals remain unresolved despite different execution accuracy. An exploratory 2,400-call follow-up records supplied preferences and resolves three of nine comparisons per model. A deterministic extractor resolves seven of nine without model calls or calibration observations, exposing unnecessary uncertainty introduced by model-generated reports. A further 14,400 controlled multinomial simulations distinguish structural ambiguity from weak identification and finite calibration precision. We propose a source-labeled decision receipt and provide an offline viewer for inspecting the audit. These results motivate preserving decision-relevant structured input and diagnosing why a decision is unresolved before collecting more telemetry. The study contains no human participants or real customer outcomes. Full proofs, raw model provenance, controlled experiments, and reproducible analyses accompany the report.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.26642