arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01656cs.HC

你无法优化你无法衡量的:多任务评估作为AI介导的平视交互的缺失基础

You Cannot Optimize What You Cannot Measure: Multitasking Evaluation as the Missing Foundation of AI-Mediated Heads-Up Interaction

Nuwan Janaka, Runze Cai, Yang Chen, Chenyu Zhao, Shengdong Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

针对AI介导的平视AR界面评估多基于快照、存在干扰难量化等局限,提出转向多任务轨迹评估等三项改进,为该领域提供缺失的评估基础。

中文摘要 AI 辅助

AI介导的平视增强现实(AR)用动态适配的界面取代固定界面,这类界面会基于不断变化、无法完全提前预知的上下文,决定呈现的信息内容、形式和时机。尽管它仍是一种界面,但其随时间的行为在设计时仅部分被明确。我们认为,这种转变需要相应的评估方式改变:从快照转向轨迹。固定界面在快照中评估——单一上下文、单一会话、一组任务绩效指标;而流动界面必须在轨迹中评估——从界面实际会遇到的分布中采样的一系列带转换的上下文序列,且需跟踪足够长时间以使用户信任形成、演变甚至恶化。借鉴光学透视头戴式显示器(OST-HMDs)支持的平视AR多任务文献,我们发现当前评估实践大多仍基于快照:多数研究采用固定条件、单一会话设计;并发任务间的干扰很少被直接量化;常用的工作负荷指标无法区分归因于单个任务的认知负荷。为解决这些局限,我们主张三个转变:从孤立指标转向性能操作特征(POC)干扰边界,从固定条件转向上下文轨迹评估,从单一会话快照转向纵向信任测量。

英文摘要

AI-mediated heads-up augmented reality (AR) replaces fixed interfaces with dynamically adapting ones that decide what information to present, in what form, and when, based on a continually changing context that cannot be fully anticipated beforehand. Although it remains an interface, its behavior over time is only partially specified at design time. We argue that this shift requires a corresponding change in evaluation: from snapshots to trajectories. A fixed interface is evaluated in a snapshot --- one context, one session, one set of task-performance metrics. A fluid interface must be evaluated over a trajectory --- a sequence of contexts with transitions, sampled from the distribution the interface will actually encounter, and tracked long enough for user trust to form, evolve, and potentially deteriorate. Drawing on the literature for heads-up AR multitasking enabled by optical see-through head-mounted displays (OST-HMDs), we find that current evaluation practice remains largely snapshot-based. Most studies use fixed-condition, single-session designs; interference between concurrent tasks is rarely quantified directly; and commonly used workload measures cannot disentangle cognitive load attributable to individual tasks. To address these limitations, we argue for three shifts: from isolated metrics to Performance Operating Characteristic (POC) interference frontiers, from fixed conditions to evaluation over context trajectories, and from single-session snapshots to longitudinal trust measurement.

补充信息

↑