野外提示:代码中事务性提示的大型分析集合
Prompts in the Wild: A Large Analyzed Collection of Transactional Prompts in Code
- Bar-Ilan University(巴伊兰大学)
- Allen Institute for Artificial Intelligence(艾伦人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文从GitHub收集57500个代码事务性提示样本,构建结构化本体将其结构化,分析得使用模式呈齐普夫分布,经错误分析验证标注可靠性,发布数据集及浏览界面。
AI中文摘要:
当代生成式大型语言模型(LLM)的行为直接由提示塑造,提示是描述期望输出和模型行为的非结构化文本。本文认为,提示是值得自身研究的语言对象。为此,我们从GitHub收集了57500个独特的提示样本,特别关注事务性提示:集成到软件中的可复现自然语言指令。为实现对提示的实证定量研究,我们引入了结构化本体,捕捉提示的属性及其形式和语义组件。基于该本体,我们将非结构化原始文本形式的提示转化为结构丰富的语言对象。对这些结构化数据的分析显示,不同语言、领域、任务和模态的使用模式存在显著多样性,呈现典型的齐普夫分布:部分提示明显占主导,其余更多样的提示则出现在长尾中。为验证基于本体的提示标注的可靠性,我们对所有领域进行了全面的错误分析,提供了标注质量的详细评估。我们发布了该数据集及一个浏览探索界面(此https URL)。
英文摘要:
The behavior of contemporary generative Large Language Models (LLMs) is directly shaped by prompts, unstructured texts that describe the desired output and model behavior. In this paper we argue that prompts are linguistic objects that merit investigation in their own right. To this end, we collect 57.5K unique samples of prompts from GitHub. Specifically, we focus on transactional prompts: reproducible natural language instructions that are integrated into software. To enable the empirical, quantitative study of prompts, we introduce a structured ontology, capturing the properties of prompts as well as their formal and semantic components. Based on this ontology, we transform prompts from unstructured raw texts into richly structured linguistic objects. Analysis of these structured data reveals significant diversity of usage patterns across languages, domains, tasks, and modalities, in a typical Zipf-like distribution where some clearly prevail and others, more diverse, appear in the long tail. To validate the reliability of the ontology-based annotation of the prompts, we perform a comprehensive error analysis across all fields, providing a detailed assessment of annotation quality. We release the dataset together with a browsing and exploration interface (https://github.com/OnlpLab/transactionalPromptsCollection ).