arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

衡量长时程人类活动模拟的行为保真度

Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations

Yi Fei Cheng, Fan Yang, Iremsu Bas, Koichiro Niinuma, Narishige Abe, David Lindlbauer

arXiv 2609.01257首次发表:更新:

发表机构

Human-Computer Interaction Institute, School of Computer Science, Carnegie Mellon University; Fujitsu Ltd.; Fujitsu Research of America(卡内基梅隆大学计算机科学学院人机交互研究所; 富士通有限公司; 富士通美国研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出评估长时程人类活动模拟行为保真度的框架,通过43小时办公活动数据集对比不同条件机制,发现统计先验最贴近真实行为但存在缺陷,推动了更全面的评估方式。

AI 中文摘要

基于大语言模型(LLM)的人类模拟器正越来越多地被用于政策制定、评估与训练,它们必须忠实地再现真实行为模式。尽管已有研究调查了调查回应与对话中的行为保真度,但长时程现实世界活动仍在很大程度上未被探索。我们提出了一个用于评估长时程活动模拟行为保真度的框架,该框架覆盖时间粒度与分析层级。作为案例研究,我们收集了一个时长43小时的野外办公活动多摄像头数据集,并比较了由轨迹推导的条件机制:角色描述符、少样本示例,以及统计转移与时段先验。我们发现行为保真度在不同指标间并不均匀:统计先验使活动与序列分布最接近真实行为,但过度碎片化日常流程并抑制个体内部变异性。这些发现推动了需覆盖多指标、时间粒度与分析层级的更全面评估。

英文摘要

As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined behavioral fidelity in survey responses and dialogue, longer-horizon real-world activity remains largely unexplored. We introduce a framework for evaluating behavioral fidelity in long-horizon activity simulations across temporal granularities and levels of analysis. As a case study, we collect a 43-hour multi-camera dataset of in-the-wild office activity and compare trace-derived conditioning mechanisms: persona descriptors, few-shot exemplars, and statistical transition and time-of-day priors. We find that behavioral fidelity is not uniform across metrics: statistical priors bring activity and sequence distributions closest to real behavior, yet over-fragment routines and suppress within-person variability. These findings motivate a more holistic evaluation that spans multiple metrics, temporal granularities, and levels of analysis.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑