arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从计算机使用轨迹中诱导任务模型

Inducing Task Models from Computer-Use Traces

Yucheng Jiang, Zora Zhiruo Wang, Ruishi Chen, Diyi Yang

arXiv 2608.20319首次发表:更新:

发表机构

Stanford University; Carnegie Mellon University(斯坦福大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出TMI方法,可从计算机使用轨迹中分离并发任务并诱导结构化任务模型,在恢复交织任务、重构执行步骤及提升任务准确率上均优于现有基线。

AI 中文摘要

自然主义的计算机使用轨迹(即被动记录的截图及鼠标、键盘操作)是推导日常工作完成方式的可审计、可重用符号模型的宝贵资源。这类模型至关重要,因为计算机使用智能体进入实际工作场景后,需要学习任务的实际执行方式,而企业也需要对这类知识进行审计和重用。然而,诱导此类任务模型颇具挑战性,因为活动仅以低级别事件的形式被观测到,且实际工作是多线程的,存在交织的目标。现有方法假设存在给定的任务或单一工作流,仅生成步骤级摘要而非结构化任务模型。我们提出任务模型诱导(Task Model Induction,TMI),该方法可:(i)在无约束轨迹中发现潜在任务,分离交织的并发活动;(ii)为每个潜在任务诱导任务模型,将递归目标分解的分层目标模型与组织执行的控制流过程模型配对。在受控的人类和智能体轨迹上,TMI在恢复交织任务时与真实分组的一致性达0.974,重构74.9%的观测执行步骤,远优于最强的工作流诱导基线。在外部实验中,从TMI任务模型衍生的技能较最强基线提升了30.0%的保留任务准确率。

英文摘要

Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbolic, auditable, and reusable models of how everyday work is done. Such models matter as computer-use agents enter real work, where agents need to learn how tasks are actually performed, and organizations need to audit and reuse that knowledge. However, inducing such task models is challenging, as activity is observed only as low-level events and real-world work is multi-threaded with interleaved goals. Existing methods assume a given task or a single workflow, and produce step-level summaries rather than structured task models. We introduce Task Model Induction (TMI), which (i) discovers the latent tasks in an unconstrained trace, disentangling concurrent activity, and (ii) for each latent task, induces a task model pairing a hierarchical objective model of recursive goal decomposition with a procedure model of the control flow that organized the execution. Intrinsically, on controlled human and agent trajectories, TMI recovers interleaved tasks with 0.974 agreement against ground-truth groupings and reconstructs 74.9% of the observed execution steps, far more than the strongest workflow induction baseline. Extrinsically, skills derived from TMI's task models improve held-out task accuracy by 30.0% over the strongest baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑