发表机构
Stony Brook University; Old Dominion University(石溪大学; 奥多明尼昂大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过三周日记研究评估了OLLA原型及GPT-5等模型对盲用户桌面应用交互的有效性,发现GPT-5成功率达52.5%,同时揭示了CUAs的不足及用户的额外需求。
AI 中文摘要
计算机使用智能体(Computer-use agents,CUAs)正作为一种具身人机交互范式兴起,它结合语言推理与多模态界面接地技术来操作图形用户界面(GUIs)。然而,其在真实桌面工作流中对使用屏幕阅读器的盲用户的有效性仍不明确。我们开展了一项为期三周的日记研究,8名盲用户使用OLLA(一款支持屏幕阅读器访问的CUA原型),在12个应用中收集了1258条命令,附带截图、UI树、模型响应及操作轨迹。部署期间我们评估了GPT-5,并用另外4个模型重新执行相同命令。GPT-5的成功率最高,达52.5%。轨迹分析显示存在接地、规划、约束跟踪及终止失败,访谈则揭示了超越自动化的需求。
英文摘要
Computer-use agents are emerging as a paradigm for agentic human-AI interaction, combining language reasoning with multi-modal interface grounding to operate GUIs. Yet their effectiveness for blind screen-reader users in real-world desktop workflows remains unclear. We present a three-week diary study with 8 blind users using OLLA, a screen-reader-accessible CUA prototype, collecting 1,258 commands across 12 applications with screenshots, UI trees, model responses, and action traces. We evaluate GPT-5 during deployment and re-execute the same commands with four additional models. GPT-5 achieved the highest success rate at 52.5%. Trace analysis reveals grounding, planning, constraint-tracking, and termination failures, while interviews reveal beyond-automation needs.
CommentsAccepted to the EMNLP 2026 Main Conference