Practical Secrets Extraction against Black-box LLMs
Researchers recover memorized credentials from black-box commercial LLMs, including OpenAI and Claude Code.
Researchers present a black-box framework that extracts memorized secrets from commercial API language models using only model outputs. Semantics-preserving prompts and response cross-validation distill secret-related behavior into a local proxy, which then guides candidate filtering. On controlled API-key benchmarks the method improves recovery effectiveness and real-key rates while reducing latency. A responsible evaluation recovered masked provider-specific credentials from three deployed systems spanning OpenAI and Claude Code.
- Output-only access is enough; weights and token probabilities are not required.
- A local proxy and cross-validation improve recovery on controlled API-key benchmarks.
- Masked credentials were recovered from three deployed OpenAI and Claude Code systems.
Full article171 words · extracted from arxiv.org · click to collapse
Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extraction audits, however, largely assume access to model weights or token probabilities. In this work, we present a black-box secret extraction framework for commercial, API-based LLMs under output-only access. It comprises (i) \emph{Cross-Validated Secret Knowledge Distillation}, which uses semantics-preserving prompt variants, response cross-validation, and provider-specific format filtering to distill secret-relevant behavior into a local white-box proxy; and (ii) \emph{Proxy-Guided Secret Extraction and Candidate Filtering}, which combines truncated top-$p$ sampling with local token entropy, $N$-gram frequency profiling, and provider-specific structural priors. On controlled API-key benchmarks, our framework improves recovery effectiveness and real-key rates over representative baselines while reducing extraction latency. A responsible real-world evaluation further recovers masked provider-specific credentials from three independently deployed black-box LLM systems spanning OpenAI and Claude Code, showing that memorized secrets can be exposed under output-only access.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.36941