Configuration, Not Conscience: A Large-Scale Empirical Study of LLM System Prompts
A study of 407 LLM system prompts finds they function as operational configs, not value statements.
Researchers analyzed a merged corpus of 407 leaked, reconstructed, or officially published system prompts from 62 vendors across four community collections, identifying 29 near-duplicate clusters covering 66 files. A simple block-level classifier assigns roughly 58% of classified words to tool and protocol content and about 5% to safety policy, while the strictest rule lines guard tool use and file safety over harmful content by an 11:1 margin. Literal reuse concentrates in a few cross-vendor pairs, and version chains turn over thousands of words per release. The authors treat these prompts as operational configuration and frame reuse and prompt rot as supply-chain concerns, cautioning that most documents are adversarial in origin so the magnitudes are directional.
- Corpus covers 407 prompts from 62 vendors across four collections.
- About 58% of classified words cover tools and protocols, versus 5% safety policy.
- Strictest rules favor tool and file safety over harmful content 11 to 1.
- Cross-vendor copying and prompt rot are treated as supply-chain issues.
Full article175 words · extracted from arxiv.org · click to collapse
Leaked system prompts are often treated as windows into the hidden values of commercial language models, yet their composition is rarely studied at scale. We analyze a merged corpus of 407 leaked, reconstructed, or officially published system prompts from 62 vendors across four community collections, identifying 29 near-duplicate clusters covering 66 files. Operational content rather than ethical statements dominates the corpus; a deliberately simple block-level classifier assigns roughly 58\% of classified words to tool/protocol and roughly 5\% to safety policy, while the strictest rule-lines guard tool use and file safety over harmful content by an 11:1 margin. Literal text transfer concentrates in a small set of cross-vendor pairs. Prompts also carry measurable maintenance debt, with version chains turning over thousands of words per release. The evidence supports treating leaked prompts as operational specifications, closer to configuration files than value statements, and treats reuse and prompt rot as engineering and supply-chain concerns. Because most documents are adversarial in origin and the detectors are deliberately simple, all magnitudes are directional; we audit the main classifier's error modes.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.31575