arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KnowSim:用学习的用户模拟器评估大语言模型助手的信息校准

KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn

Yoonjoo Lee, Hyoungwook Jin, Tae Soo Kim, Shaoyang Zhang, Philippe Laban, Q. Vera Liao

arXiv 2608.17150首次发表:更新:

发表机构

University of Michigan; KAIST; Microsoft Research(密歇根大学; 韩国科学技术院; 微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出KNOWSIM框架,通过带显式知识状态的用户模拟器评估LLMs的信息校准,验证后其性能优于基线,可揭示用户知识水平与LLMs的适配关系。

AI 中文摘要

为了在知识密集型任务中与用户有效协作,大语言模型(LLMs)必须进行信息校准:使内容匹配用户不断变化的理解程度和认知能力。然而,用于评估和训练LLMs的用户模拟器未明确建模用户知识,因此既无法产生不同知识水平下的真实交互,也无法反映该知识随时间演变的交互过程。为弥合这一差距,我们引入KNOWSIM,这是一个围绕用户模拟器构建的评估框架,该模拟器维护明确的知识状态,以具有先决关系的信息单元图表示,并在基于学习理论的更新规则下演变。KNOWSIM直接从知识状态轨迹计算三个指标(知识增益、交付校准、认知过载),反映信息校准的关键机制方面。我们针对两个领域中按知识水平分层的705个人类-AI会话验证KNOWSIM:其排名与人类判断显著一致(73%-74%符号一致性),优于三个基线模拟器。将其应用于9个LLMs时,KNOWSIM揭示最佳模型随用户知识水平变化,显示出标准评估无法发现的能力-处理交互作用。

英文摘要

To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.

Comments30 pages, 6 figures, 16 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑