合成TLX:使用智能体模拟预测人类工作负荷
Synthetic TLX: Forecasting Human Workload Using Agent Simulation
浏览论文内容
中文总结 AI 辅助
本文提出合成TLX范式,利用智能体模拟在任务执行前预测NASA TLX工作负荷评分,通过三项实验验证其与人类评分的对齐性,并展示其应用潜力。
中文摘要 AI 辅助
评估技术中介任务中的人类工作负荷有助于防止因技术设计不佳而导致的任务失败。传统上,工作负荷是在人类完成任务后,使用NASA任务负荷指数(TLX)进行回顾性评估。如果我们能在人类尝试任务之前,使用智能体模拟来预测工作负荷,会怎样?我们引入了合成TLX(Synthetic TLX),这是一种用于主动工作负荷估计的新范式,能够预测给定任务的NASA TLX评分,从而开启新颖的交互机会和评估方法。为了理解其可行性,我们进行了三项实验,比较人类和智能体生成的评分,以评估它们在哪些方面一致和分歧。我们发现,当智能体被赋予人类角色提示并进行主动任务模拟时,其估计与人类评分尤为一致。然而,智能体和人类在其敏感的工作负荷来源上存在分歧。基于我们的发现,我们展示了三个应用来展示合成TLX的潜力,并讨论了未来工作负荷感知的人机交互。
英文摘要
Assessing human workload for technology-mediated tasks helps prevent task failure caused by poor technology design. Traditionally, workload is assessed retrospectively using the NASA Task Load Index (TLX) after humans complete a task. What if we could forecast workload before a human attempts a task using agent simulation? We introduce Synthetic TLX, a new paradigm for proactive workload estimation that predicts NASA TLX scores for a given task, unlocking novel interaction opportunities and evaluation methods. To understand its viability, we conducted three experiments comparing human and agent-generated scores to evaluate where they align and diverge. We found agent estimates align with human scores particularly when prompted with a human persona and active task simulation. However, agents and humans diverge in the sources of workload they are sensitive to. Based on our findings, we present three applications to showcase Synthetic TLX's potential and discuss the future of workload-aware human-AI interaction.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。