发表机构
The University of Hong Kong; Microsoft Research(香港大学; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出GenUI-Harness多智能体框架,结合工具智能体与GUI编码智能体,通过Dynamic UX和Reward Auditor解决UI生成训练的奖励问题,推出UI-TAU Bench基准,在任务完成率和对话轮次上均优于现有模型。
AI 中文摘要
当前多数人机交互仍基于文本,自然语言会给复杂任务带来认知过载、歧义、信息混乱及输入缓慢等问题;短暂的生成式UI可呈现结构化信息并引导用户完成任务。我们提出GenUI-Harness,这是一种多智能体harness,将用于信息检索和任务执行的工具智能体,与用于识别歧义并生成结构化界面前端代码的GUI编码智能体配对。用强化学习训练该编码智能体存在挑战:可验证的交互式UI生成奖励需要成本高昂的执行,而大语言模型作为评判者(LLM-as-a-Judge)的奖励易受奖励黑客行为影响。我们通过Dynamic UX解决第一个挑战,这是一个用于在单个沙箱中进行动态交互和奖励收集的轻量级包;通过Reward Auditor解决第二个挑战,这是一种元奖励机制,用于监控奖励分布并将诊断模式提炼为共享规则和评分规范。我们推出UI-TAU Bench,这是一个用于通过生成式UI代码进行主动人机交互的基准,基于从公共数据源构建的10个真实领域数据库,并基于Tau-Bench工具使用设置,包含Lite(300个任务)和Full(1000个任务)两个拆分版本。GenUI-Harness在Lite版本上相比smolagents实现了4.48个百分点的平均Pass@3提升;用GenUI-Harness训练4B主干模型后,其Pass@3从9.33%提升至58.00%,超过Claude Opus 5等更大的前沿模型(46.67%);GenUI-Harness在歧义与非歧义查询上均保持稳健,在对比沟通渠道的评审调查中,生成式UI将平均对话轮次从3.4降至1.2。这些结果表明,数据感知生成式界面可支持有效任务完成,并减少经评估的数据库支持工作流中的对话轮次。
英文摘要
Most human-agent interaction today remains text-based. Natural language can impose cognitive overload, ambiguity, information chaos, and slow input for complex tasks; ephemeral generative UIs can present structured information and guide users toward task completion. We propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces. Training the coder with reinforcement learning is challenging: verifiable rewards for interactive UI generation require costly execution, while LLM-as-a-Judge rewards are prone to reward hacking. We address the first challenge with Dynamic UX, a lightweight package for dynamic interaction and reward collection in a single sandbox, and the second with Reward Auditor, a meta-reward mechanism that monitors reward distributions and distills diagnostic patterns into a shared rubric and scoring specification. We introduce UI-TAU Bench, a benchmark for active human-agent interaction through generated UI code, built on 10 real-world domain databases constructed from public data sources and based on Tau-Bench tool-use settings, with Lite (300 tasks) and Full (1,000 tasks) splits. GenUI-Harness achieves an average Pass@3 gain of 4.48 percentage points over smolagents on Lite. Training with GenUI-Harness improves a 4B backbone from 9.33% to 58.00% Pass@3, outperforming larger frontier models such as Claude Opus 5 (46.67%). GenUI-Harness also remains robust on ambiguous and non-ambiguous queries. In a reviewer survey comparing communication channels, generated UIs reduce average dialogue rounds from 3.4 to 1.2. These results show that data-aware generative interfaces can support effective task completion and reduce dialogue rounds in evaluated database-backed workflows.
Comments35 pages, 7 figures. Code: https://github.com/bird-bench/GenUI-Agent