APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
APort Vault shows a passport check stopped unauthorized payment recipients across 225,964 agent evaluations.
APort Vault replays 4,371 human-written attacks from a public capture-the-flag against a live payment agent. The study covers 14 models from eight labs, five policy levels, and two replay tracks, with and without an Open Agent Passport pre-action check, totaling 225,964 evaluations. At levels 2–4, transfers to recipients the passport did not permit were 140 of 76,842 with the model alone and zero of 69,297 behind the check. The authors release the evaluations, passports, and scoring code.