arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型的偏好智能体协同设计:参与可能导致过度信任

Co-design of LLM-based preference agents: participation may drive overtrust

Michael J. Fell

arXiv 2607.21757首次发表:更新:

发表机构

University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨大语言模型模拟人类偏好引发的问题,通过12名参与者在家庭能源领域协同设计个人偏好智能体的定性研究,发现参与和过程透明度或成“过度信任引擎”,并将此发展为参与式偏好智能体设计核心机制。

AI 中文摘要

大语言模型在研究和实际应用中越来越多地用于模拟人类偏好,引发了对验证、误传和排斥的担忧。与所代表的人协同设计智能体是解决这些问题的一种有前景的方式,但参与也可能掩盖其看似解决的问题。本文通过一项主要的定性研究来探讨这种矛盾,12名参与者通过背景调查、协同设计访谈和验证调查,在家庭能源领域协同设计个人偏好智能体。参与者积极参与,大多认为他们的智能体很好地代表了自己。然而,独立验证显示人机一致性参差不齐,智能体的回答比人类样本明显更趋同、果断和抽象。本文认为参与和过程透明度可能成为一个“过度信任引擎”,促进信任的同时掩盖与潜在结构后果的系统性不一致。本文将此发展为参与式偏好智能体设计的核心机制,将个体一致性视为一个动态过程而非固定状态。

英文摘要

Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designed personal preference agents in the domain of household energy, via a background survey, co-design interview, and validation survey. Participants engaged readily and mostly came to see their agents as representing them well. Independent validation, however, revealed mixed human-agent alignment, with agent responses markedly more homogeneous, decisive, and abstract than the human sample. I argue that participation and process transparency can act as an "overtrust engine" that promotes trust while concealing systematic misalignment with potential structural consequences at scale. I develop this as a core mechanism in participatory preference agent design, treating individual alignment not as a fixed state but as an enacted process.

Comments43 pages (18 main text plus supplementary material), 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑