arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越任务完成:衡量终端用户界面中的交互成本

Beyond Task Completion: Measuring Interaction Cost in Terminal User Interfaces

Ruida Hu, Yuanhao Wang, Chao Peng, Yakun Zhang, Cuiyun Gao

arXiv 2610.05047首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen; University of Edinburgh(哈尔滨工业大学(深圳); 爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Agent-as-a-User范式,通过智能体模拟用户,分离解释与操作成本,衡量终端用户界面的交互成本,实现可重复、基于执行的任务可用性评估。

AI 中文摘要

大型语言模型(LLM)越来越多地通过终端用户界面(TUI)使用,然而仅凭任务完成情况并不能反映界面理解和操作的难度。现有的人类评估以及LLM生成的评分或报告,无法提供基于验证任务执行的可重复的交互努力度量。我们提出“Agent-as-a-User”评估范式,将LLM智能体置于用户角色。Agent-KLM将智能体端的解释成本和操作成本分离,并将其与人类交互在结构上关联。TUINaut通过记录交互轨迹并使用与执行者无关的预言机验证结果来操作化该范式。利用TUINaut-Bench,我们研究了15个真实TUI中的84个任务,涉及六名人类参与者和五种观察-评估配置,并评估了由三个领先技术栈生成的45个TUI。相似的总体成功率掩盖了人类和智能体完成任务上的差异,而解释成本和操作成本遵循相关的任务排名,尤其是在操作方面。这种模式在不同配置中持续存在。生成的TUI可以实现正确的功能,但仍需可用性改进。这些结果将Agent-as-a-User定位为一种可重复、基于执行的任务可用性度量范式。

英文摘要

Large language models (LLMs) are increasingly used through terminal user interfaces (TUIs), yet task completion alone does not capture how difficult an interface is to understand and operate. Existing human assessments and LLM-generated ratings or reports do not provide repeatable measurements of interaction effort grounded in verified task execution. We propose Agent-as-a-User, an evaluation paradigm that places an LLM agent in the user role. Agent-KLM separates agent-side interpretation and operation costs and relates them structurally to human interaction. TUINaut operationalizes the paradigm by recording interaction trajectories and verifying outcomes with actor-independent oracles. Using TUINaut-Bench, we study 84 tasks across 15 real-world TUIs with six human participants and five observation-evaluator configurations, and evaluate 45 TUIs generated by three leading stacks. Similar aggregate success rates mask differences in which tasks humans and agents complete, whereas interpretation and operation costs follow correlated task rankings, especially for operation. This pattern persists across configurations. Generated TUIs can implement correct functionality while still requiring usability improvements. These results position Agent-as-a-User as a repeatable, execution-grounded paradigm for measuring task-based usability.

CommentsWeb Page: https://tuinaut.pages.dev/ Source Code: https://github.com/kinesiatricssxilm14/Agent-as-a-User

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑