arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

回归正轨:一种用于用户模拟中角色保真度创作与测量的社会语言学方法

Back in Style: A Sociolinguistic Approach to Authoring and Measuring Persona Fidelity in User Simulation

Lex Konnelly, Elena Khasanova, Riqiang Wang, Matthias Lee, Harsh Saini, Parsa Kavehzadeh

arXiv 2610.10988首次发表:更新:

发表机构

Dialpad Inc.(戴尔德公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种社会语言学方法,将用户角色以具体风格比率创作,结合作者身份验证风格计量法与词典内容分析,提升了多数客服智能体模拟的风格保真度,为用户模拟的错误归因分析提供了有效路径。

AI 中文摘要

随着智能体系统在商业领域日益普及,用户模拟器逐渐成为评估这些系统的测量工具。然而,模拟用户与真实人类用户的保真度通常较低,且通常需要通过成本高昂、主观性强的大语言模型(LLM)评判者进行评估。在这项试点研究中,我们探究是否可以通过社会语言学方法确定性地测量保真度,将用户角色视为一种从可观察语言风格中浮现的社会类型,而非模型必须从标签或描述中推断出行为的类型。我们将角色创作成具体的风格比率,这使我们能够将两种成熟的无模型工具——作者身份验证风格计量法和基于词典的内容分析——作为保真度诊断工具。我们针对五个面向任务的客服智能体,对该社会语言学方案与单一描述性基线进行了A/B测试。结果显示,对于大多数测试模型,该社会语言学方案提升了风格贴合度和风格计量可区分性,但需注意角色风格保真度并不等同于角色的“自然性”。我们认为,角色设计的社会语言学方法是实现更多样化、更具代表性用户角色的有前景路径,且这些指标在错误归因分析中最具价值,可定位保真度失效的位置。这是推动用户模拟更接近忠实呈现多样且可变语言输出的干预措施的第一步。

英文摘要

As agentic systems gain commercial popularity, user simulators increasingly serve as measurement instrument for their evaluation. However, the fidelity of simulated users in comparison to real human users is generally low, and typically assessed by costly, subjective LLM judges. In this pilot study, we ask whether fidelity can instead be measured deterministically by treating a user persona sociolinguistically: as a social type that emerges from observable linguistic style, rather than one predicted by labels or descriptions a model must extrapolate into behaviour. We author personas as concrete stylistic rates, which lets us transfer two established, model-free instruments -- authorship-verification stylometry and lexicon-based content analysis -- as fidelity diagnostics. We A/B-test the sociolinguistic schema against a flat descriptive baseline across five task-oriented customer-service agents. Results show that the sociolinguistic schema improves both stylistic adherence and stylometric distinguishability for most of the tested models, with a caveat that persona style fidelity does not necessarily equal persona "naturalness". We argue that a sociolinguistic approach to persona design is a promising path towards more diverse and representative user personas, and that these metrics are most valuable in an error-attribution analysis, localizing where fidelity breaks down. This is a first step towards interventions that move user simulations closer to faithful renderings of diverse and variable linguistic outputs.

CommentsAccepted to the UserSim @ NeurIPS 2026 workshop (non-archival)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑