配置,而非良知:LLM 系统提示的大规模实证研究
Configuration, Not Conscience: A Large-Scale Empirical Study of LLM System Prompts
- University of Piraeus(比雷埃夫斯大学)
- Athena Research Centre(雅典研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究大规模分析407个泄露系统提示,发现其以操作性内容为主,接近配置文件而非价值声明,并揭示跨供应商重用与维护债务问题。
AI中文摘要:
泄露的系统提示常被视为窥探商业语言模型隐藏价值观的窗口,然而其构成却鲜少在大规模范围内得到研究。我们分析了来自四个社区集合中62家供应商的407个泄露、重建或官方发布的系统提示的合并语料库,识别出覆盖66个文件的29个近重复簇。操作性内容而非道德声明主导了该语料库;一个刻意简化的块级分类器将约58%的分类词汇归为工具/协议类,约5%归为安全策略类,而最严格的规则行对工具使用和文件安全的保护超过对有害内容的保护,比例为11:1。字面文本转移集中在少量跨供应商配对中。提示还带有可测量的维护债务,版本链每次发布都会更替数千词。证据支持将泄露的提示视为操作性规范,更接近配置文件而非价值声明,并将重用和提示腐烂视为工程和供应链问题。由于大多数文档来源具有对抗性且检测器刻意简化,所有量级均为方向性的;我们审计了主要分类器的错误模式。
英文摘要:
Leaked system prompts are often treated as windows into the hidden values of commercial language models, yet their composition is rarely studied at scale. We analyze a merged corpus of 407 leaked, reconstructed, or officially published system prompts from 62 vendors across four community collections, identifying 29 near-duplicate clusters covering 66 files. Operational content rather than ethical statements dominates the corpus; a deliberately simple block-level classifier assigns roughly 58\% of classified words to tool/protocol and roughly 5\% to safety policy, while the strictest rule-lines guard tool use and file safety over harmful content by an 11:1 margin. Literal text transfer concentrates in a small set of cross-vendor pairs. Prompts also carry measurable maintenance debt, with version chains turning over thousands of words per release. The evidence supports treating leaked prompts as operational specifications, closer to configuration files than value statements, and treats reuse and prompt rot as engineering and supply-chain concerns. Because most documents are adversarial in origin and the detectors are deliberately simple, all magnitudes are directional; we audit the main classifier's error modes.