arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

压力下的行为:六十个语言模型在用户施压时的表现

Conduct Under Pressure: What Sixty Language Models Do When a User Pushes

Tapan Parikh

arXiv 2609.25447首次发表:更新:

发表机构

Cornell Tech(康奈尔科技学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过向60个语言模型发送多轮施压场景,发现模型是否坚持立场取决于其能力,而坚持方式则与供应商相关,并验证了LLM编码者在行为标注中的可靠性。

AI 中文摘要

我们研究了当用户在令人不适的情境中施加压力时,大型语言模型(LLM)会做什么:用户坚持、恳求、奉承或悲痛,而模型会放弃一个正确的事实、撰写本应拒绝的文件,或为一个将使用户花钱的计划叫好。我们向来自13家供应商的60个模型发送了冻结的多轮场景,每个模型无论回复如何都使用相同的场景,并使用通过开放式编码构建然后冻结的编码本对每个对话记录进行标注:一个轨迹(模型坚持立场或屈服)和一个方式(它如何坚持或屈服)。两个发现截然不同。模型是否坚持与其生成能力相关,即其新旧程度:屈服率与公共能力指数在Spearman相关系数为-0.64时相关,供应商效应很小。模型如何坚持则与供应商相关:17个方式编码中有6个按供应商排序,置换检验p≤0.001,并针对整个编码本进行了校正。我们报告了在通过可靠性检验的编码上的四个供应商画像。我们还询问了标注的哪些部分需要人工参与。来自三家供应商的六个LLM编码者比三个人类编码者更一致地应用编码本(Krippendorff's alpha为0.66对0.46),在编码本示例从未涉及的对话记录上,与编码本作者在轨迹上的kappa为0.84至0.91,并与经裁决的人类参考在0.83上匹配。盲机器读数恢复了编码本的类别,但无法判断其中哪些类别第二位读者会以相同方式应用。我们得出结论,对于非专业人士可以判断的行为,人类的贡献是编写和界定编码并拥有一个小型参考集,而不是大量生成标签。

英文摘要

We study what LLMs do when a user applies pressure in an uncomfortable situation: a user insists, begs, flatters or grieves, and the model gives up a correct fact, writes a document it should refuse, or cheers a plan that will cost the user money. We send frozen multi-turn scenes, identical for every model regardless of the reply, to 60 models from 13 vendors, and label each transcript with a codebook built by open coding and then frozen: a trajectory (the model held its position or folded) and a manner (how it held or folded). Two findings separate. Whether a model holds tracks its generation, meaning how recent it is: fold rate correlates with a public capability index at Spearman -0.64, with little vendor effect. How it holds tracks the vendor: six of the 17 manner codes sort by vendor at permutation p <= 0.001, corrected across the codebook. We report four vendor profiles on the codes that cleared reliability. We also ask which parts of the labeling need a person. Six LLM coders from three vendors apply the codebook more consistently than three human coders do (Krippendorff's alpha 0.66 against 0.46), agree with the codebook's author on trajectory at kappa 0.84 to 0.91 on transcripts the codebook's examples never touched, and match an adjudicated human reference at 0.83. Blind machine readings recover the codebook's categories but cannot tell which of them a second reader would apply the same way. We conclude that for behavior a non-specialist can judge, the human contribution is authoring and bounding the codes and owning a small reference, not producing labels at volume.

Comments16 pages, 1 figure, 4 tables. v2 adds a preregistered replication on six further scenes. Code, data and labels: https://github.com/tap2k/modelun/studies/conduct

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑