发表机构
University of Minnesota; School of Statistics(明尼苏达大学; 统计学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人机协作中人类监督与人工智能效率的矛盾,提出在人机协作中放置监督阶段的问题,开发非均匀性原则,即在工作流程中以非递减间隔放置监督阶段,并在撰写文献综述和构建网站的工作流程中实证验证。
AI 中文摘要
随着生成式人工智能越来越多地应用于自动化多步骤和高风险工作流程,人类判断和参与对于确保人工智能生成输出的质量仍然至关重要。在实践中,虽然希望人类专家定期对人工智能进行监督,通常是通过审查中间输出、提供反馈、进行纠正和指导后续步骤,但这种监督受到人类可支配时间和资源的限制。这就产生了人类监督需求与人工智能在较少干预下提供更多输出的效率之间的矛盾。一个重要但未充分探索的问题是如何在人机协作中最优地让人类参与。这项工作最初源于我们的实证观察,即在长人工智能工作流程中,人类监督通常会提高用户满意度,同时减少不必要的返工和令牌消耗。从那里,我们提出了在人机协作中何处放置监督阶段的问题。在合理假设下,我们开发了非均匀性原则,该原则指出最优调度应在工作流程中以非递减间隔放置监督阶段。我们在两个常见的人工智能代理工作流程中对这一原则进行了实证验证:撰写文献综述和构建网站。
英文摘要
As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and involvement remain essential for ensuring the quality of AI-generated outputs. In practice, while it is desirable for human experts to provide oversight on AI regularly, often by reviewing intermediate outputs, giving feedback, making corrections, and steering subsequent steps, such oversight is constrained by the time and resources that humans can afford. This creates a tension between the need for human oversight and AI's efficiency in delivering more output with less intervention. An important but underexplored question, then, is how to optimally engage humans in human-AI coworking. This work was originally motivated by our empirical observation that in long AI workflows, human oversight often improves user satisfaction while reducing unnecessary rework and token consumption. From there, we formulate the problem of where to place oversight stages in human-AI coworking. Under reasonable assumptions, we then develop the nonuniformity principle, which states that the optimal schedule places oversight stages with non-decreasing gaps along the workflow. We empirically validate this principle in two common AI agent workflows: writing literature reviews and constructing websites.