arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SenWorld:用于生成富含上下文评估数据的数字孪生模拟

SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data

Zenghui Zhou, Xiaoyang Li, Xiaoxuan Qiao, Zhilang Wei, Tianming Lei

arXiv 2607.19949首次发表:更新:

AI 中文总结

研究针对智能手机助手评估缺富含上下文且隐私安全数据的问题,提出SenWorld数字孪生模拟方法,通过真实数据构建场景生成评估数据,经实验验证其在数据分布等方面与真实用户基准匹配,能暴露助手故障且隐私安全。

AI 中文摘要

智能手机个人助手基于纵向个人数据进行推理,但评估它们需要已知正确答案的富含上下文的评估数据,而真实设备轨迹因隐私敏感而无法共享。为应对这一挑战,我们提出了SenWorld,这是一种基于物理的、确定性的、事件驱动的数字孪生模拟,通过构建固定的地面真值来生成此类数据。在SenWorld中,角色在由真实地图、天气、假期和网络数据构建的世界中度过一整天;每个可观察信号都保存在全系统快照中;每个评估案例都通过指向现有记录的指针进行标记,而不是通过事后注释或大语言模型(LLM)判断。我们在北京用16个角色对该方法进行了评估。生成的数据在类别分布(詹森-香农散度(JSD)为0.070)和通信记录的日常节奏(JSD低于0.1)方面与保留的真实用户基准紧密匹配,尽管生成的记录比真实记录短。在没有脚本化交互的情况下,角色形成了完全往复的对话子图和不同的行为模式。将生成的数据投影到717个评估案例中,发现生产中的智能手机助手有78个故障,主要集中在通话和短信记录上,而联系人、日程安排和闹钟从未出现故障。快照指针确认每个故障都是助手端的检索错误,无需LLM判断。总体而言,SenWorld提供了一条隐私安全、可重现且经过分布检查的评估数据路径,其标签通过构建固定。

英文摘要

Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address this challenge, we present SenWorld, a physically grounded, deterministic, event-sourced digital-twin simulation that generates such data with ground truth fixed by construction. In SenWorld, personas live through a full day in a world built from real map, weather, holiday, and network data; every observable signal is archived in full-system snapshots; and each evaluation case is labeled by a pointer to an existing record rather than by post-hoc annotation or a large language model (LLM) judge. We evaluate this method with 16 personas in Beijing. The generated data closely matches the held-out real-user benchmark in category distribution (Jensen--Shannon divergence (JSD) 0.070) and in the daily rhythm of communication records (JSD below 0.1), though generated records remain shorter than real ones. Without scripted interaction, personas form a fully reciprocated dialogue subgraph and differentiated behavioral repertoires. Projected into 717 evaluation cases, the generated data exposes 78 failures in a production smartphone assistant, concentrating on call and Short Message Service (SMS) records while contacts, schedules, and alarms never fail. The snapshot pointer confirms each failure as an assistant-side retrieval error, with no LLM judge involved. Overall, SenWorld offers a privacy-safe, reproducible, and distribution-checked path to evaluation data whose labels are fixed by construction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑