发表机构
University of California, Berkeley(加利福尼亚大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究跨越七个对话数据集,发现神经网络分类器仅凭用户消息即可高准确率识别数据来源,揭示数据集存在独特特征,且这些特征影响用户模型训练与评估,并可为数据选择提供指导。
AI 中文摘要
人类与大语言模型(LLM)的交互数据集塑造了我们对人工智能使用的理解,并为下游研究(包括用户模型的训练与评估)提供了基础。近年来,越来越多的数据集试图捕捉人类与LLM交互的代表性图景。但这些数据集所提供的图景之间有多大差异,这些差异对基于它们的研究又意味着什么?我们跨越七个对话数据集(涵盖野外聊天日志和人类偏好数据)研究了这些问题。我们首先重新审视了Torralba和Efros的数据集分类实验,发现神经网络分类器仅从用户消息就能以远高于随机水平的准确率识别对话的来源,这表明存在独特的数据集特征。这种可分离性在将数据集按人类设计的分类法维度进行匹配后依然存在,暗示了这些分类法未能捕捉的微妙差异。随后,我们考察了对用户建模的影响:数据集特征如何传播到在这些数据集上训练的用户模型的输出中;数据集选择如何影响用户模型质量的评估以及随后与这些用户模型配对的LLM助手的评估;以及数据集分类器如何指导用于训练用户模型的数据选择。虽然每个数据集都旨在捕捉“真实世界”交互的一个片段,但我们的发现揭示了这些片段分歧的程度,以及这些差异对建立在这些基础上的研究的影响。
英文摘要
Human--LLM interaction datasets shape our understanding of AI use and provide a foundation for downstream research, including training and evaluation of user models. In recent years, a growing number of datasets have sought to capture a representative picture of human--LLM interactions. But how different are the pictures these datasets provide, and what do those differences mean for research built on them? We study these questions across seven conversation datasets, spanning in-the-wild chat logs and human preference data. We begin by revisiting the dataset classification experiment of Torralba & Efros and find that neural network classifiers identify the source of a conversation from user messages alone well above chance, indicating distinctive dataset signatures. This separability persists after matching datasets on the dimensions of human-designed taxonomies, implying subtle differences that these taxonomies do not capture. We then examine the implications for user modeling: how dataset signatures propagate to the outputs of user models trained on these datasets; how dataset choice influences evaluations of user model quality and subsequent evaluations of LLM assistants paired with these user models; and how dataset classifiers can guide data selection for training user models. While each dataset is meant to capture a slice of 'real-world' interactions, our findings reveal the extent to which these slices diverge, and the consequences of those differences for research built on these foundations.