ZeroHour
arXiv cs.CRpublished ()ingested Guosen Wu

ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

infoAI safety & securityimportance 45
AI summary · glm-5.3-flash

ASLEval benchmark shows local privacy proxies miss 46.9% of LLM agent session exposure recovered by measuring all visible exits.

Researchers introduce privacy exposure displacement, the mismatch between local evaluation proxies and target-grounded exposure across full LLM agent sessions, and ASLEval, an authorization-aware framework that pre-registers hidden target sets and measures all declared visible exits. Across enterprise-style environments and independently implemented runtimes, expected-outlet-only views missed 46.9% of exposure recovered by the visible-exit union, and attacker self-reports combined omissions with high false discovery. Schema-aligned internal evidence usually preceded visible exposure at the request/probe level. The authors argue benchmarks should declare the complete visible boundary and report privacy alongside task utility.

  • Expected-outlet-only views miss 46.9% of exposure recovered by the visible-exit union in agent sessions.
  • Attacker self-reports combine omissions with high false discovery rates.
  • Schema-aligned internal evidence usually precedes visible exposure at the request/probe level.
  • ASLEval pre-registers hidden target sets and measures all declared visible exits across enterprise-style environments.
ProductsASLEval
Full article171 words · extracted from arxiv.org · click to collapse

Privacy evaluations of tool-using LLM agents often inspect a designated action, final response, or attacker report. These local proxies can miss unauthorized exposure elsewhere in a multi-step session and lack common ground truth across outlets, reports, and tool paths. We introduce privacy exposure displacement, the mismatch between a local evaluation proxy and target-grounded session exposure, and ASLEval, an authorization-aware framework that pre-registers a hidden target set, measures all declared visible exits, and reserves internal traces for diagnosis. Across multiple enterprise-style environments and independently implemented runtimes, we observe three recurring patterns. An expected-outlet-only view misses 46.9% of exposure recovered by the visible-exit union; attacker self-reports combine omissions with high false discovery; and schema-aligned internal evidence usually precedes visible exposure at the request/probe level. Reducing model-visible returns changes this path but can eliminate normal-task success. Independent human review supports the adjudication pipeline while identifying harder console and candidate cases. These findings motivate benchmarks that declare the complete visible boundary, ground claims in pre-specified targets and authorization, and report privacy together with task utility.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.18864