ZeroHour
arXiv cs.CRpublished ()ingested Jiamin Zheng1

PIA-Bench: Towards Automated Privacy Impact Assessment with Large Language Models

infoResearchimportance 48
AI summary · glm-5.3

Researchers release PIA-Bench, the first open benchmark evaluating how accurately LLMs can automate privacy impact assessments using 73 curated federal PIAs.

PIA-Bench is the first open benchmark for evaluating large language models on real-world privacy impact assessments (PIAs). The authors audited 499 expert-authored PIAs published by US federal agencies and curated 73 structured PIAs comprising 451 privacy risk items and 831 mitigation items. Off-the-shelf LLMs were found to produce meaningful assessments while identifying clear avenues for improvement. The paper calls for domain-specific LLM agent workflows, accountable LLM infrastructure, and new quality standards for PIAs.

  • First open benchmark for evaluating LLMs on real-world privacy impact assessments
  • Curated 73 PIAs from 499 US federal agency documents: 451 risks, 831 mitigations
  • Off-the-shelf LLMs produce meaningful but imperfect privacy risk assessments
  • Authors urge domain-specific agent workflows and accountable LLM infrastructure
Full article170 words · extracted from arxiv.org · click to collapse

Privacy impact assessment (PIA) is a critical instrument for institutions to proactively identify privacy risks and develop mitigation strategies before system deployment. While mandated across regulatory and institutional contexts, executing PIA requires extensive privacy and technical expertise, posing a particular challenge for teams without access to such resources. Prior work shows the potential of leveraging large language models (LLMs) to assist practitioners' privacy decisions, but little is known about how accurately and reliably LLMs can automate PIA. To this end, we develop PIA-Bench, the first open benchmark for evaluating LLMs on real-world PIAs. We first audited 499 expert-authored PIAs published by US federal agencies and curated 73 structured PIAs, comprising a total of 451 privacy risk and 831 mitigation items, to evaluate LLMs' ability to assess privacy risks and propose mitigations of complex systems. Our results show that off-the-shelf LLMs produce meaningful assessments and identify avenues for future improvement. Finally, we call for improving domain-specific workflows for LLM agents, developing accountable LLM infrastructure, and designing new quality standards for PIAs.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.12571