arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工智能代理知道任务何时简单吗?迈向复杂性感知推理与执行

Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution

Junjie Yin, Xinyu Feng

arXiv 2607.13034首次发表:更新:

AI 中文总结

研究探讨LLM代理缺乏任务感知执行范围估计能力,提出E3方法,在MSE-Bench基准测试及真实模型工具中验证,该方法能在保证成功率时大幅降低成本,推动迈向工程基础人工智能,还发布了框架和基准。

AI 中文摘要

大语言模型(LLM)代理越来越多地自动化多步骤工程和信息学工作流程,但它们很少考虑任务实际需要多少工作量。它们常采用最大上下文优先策略,将一行编辑变成小型代码库审核。我们认为缺少的能力是任务感知执行范围估计,即判断任务难度、真正所需信息及最短可靠路径。我们形式化了最小充分执行和代理认知冗余率(ACRR),并提出E3(估计、执行、扩展):代理估计初始操作点,执行最小可行路径,仅在验证失败时扩展范围。在MSE-Bench上,E3与最强基线的成功率相同,同时成本降低85%,令牌减少91%,检查文件减少92%,还比强大的自适应检索基线高出16%。配套的真实模型工具(LLM-Case)在实时gpt-4o代理编辑真实开源库时证实了该效果。我们将此视为对执行冗余的可控探究,而非对任何已部署代理的测量,并将任务感知执行定位为迈向工程基础人工智能(EGAI)的一步,我们还发布了框架和基准。

英文摘要

Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires. They often follow a maximum-context-first strategy--re-reading files and dependencies they have already seen--turning a one-line edit into a small code-base audit. We argue the missing capability is task-aware execution-scope estimation: judging a task's difficulty, the information it truly needs, and the shortest reliable path before committing budget. We formalize minimum-sufficient execution and the Agent Cognitive Redundancy Ratio (ACRR), and propose E3 (Estimate, Execute, Expand): the agent estimates an initial operating point, executes a minimum viable path, and expands scope only when verification fails. On MSE-Bench--a deterministic benchmark of 121 edits in a capability-controlled simulator--E3 matches the strongest baseline's 100% success while cutting cost by 85%, tokens by 91%, and inspected files by 92%, and further beats a strong adaptive retrieval baseline by 16%; the gains survive held-out instruction wording and essentially every cost weighting. A companion real-model harness (LLM-Case) corroborates the effect on a live gpt-4o agent editing a real open-source library, with every candidate patch graded by actually running the project's real pytest suite against a measured oracle: the over-reading is milder but real, and E3 is the leanest and fastest policy at comparable task success--its one shortfall a provider rate-limit, not a wrong edit. We frame this as a controlled probe of execution redundancy, not a measurement of any deployed agent, and position task-aware execution as a step toward engineering-grounded AI (EGAI)--agents whose effort is anchored in the engineering reality of the task. We release the framework and benchmark.

Comments27 pages, 8 figures, 8 tables. Code and benchmark: https://github.com/eejyin/Do-AI-Agents-Know-When-a-Task-Is-Simple-Toward-Complexity-Aware-Reasoning-and-Execution

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑