arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

感知底层约束的AI智能体:将执行上下文作为一等输入

Substrate-Aware AI Agents: Execution Context as a First-Class Input

Manu Agrawal

arXiv 2609.05232首次发表:更新:

AI 中文总结

该研究提出「底层约束盲视」概念,通过数值代码生成实验发现,向Anthropic Claude Opus 5等AI模型披露执行内存与时间约束,可使其生成代码主动适配,大幅提升资源利用率与任务成功率。

AI 中文摘要

自主AI智能体越来越多地在环境中选择动作,这些环境的内存、执行时间、运行时、计算及操作约束决定了何种方案是合适的。我们将智能体规划状态中缺失此类执行上下文的现象称为「底层约束盲视」。我们通过数值代码生成来验证这一普遍假设,其中选定的实现选择和操作后果是可直接观测的。三种前沿模型配置——Anthropic Claude Opus 5、OpenAI GPT-5.6-Sol和Google Gemini 3.7 Flash——针对高维成对欧氏距离任务生成代码,仅基于任务本身,或在128 MB RAM和10.0 s wall-time(挂钟时间)的约束下生成。约束披露在14个可执行的、索引对齐的「仅任务」与「约束披露」对比中,将测得的峰值进程内存降低了13个,且在所有三个组中均降低了平均挂钟时间,使执行速度最高提升3.1倍。在审核的语料库中,披露产生了结构性代码变更,包括有界阻塞、float32保留、上三角遍历以及原地或内存映射缓冲区。在更严格的96 MB约束下,独立抽样的「约束披露」组中,Claude Opus 5的正确且符合预算的结果为4/5,GPT-5.6-Sol为5/5,Gemini 3.7 Flash为3/5;而「仅任务」组的对应结果分别为0/5、1/5和0/5;各组的平均MaxRSS(最大驻留集大小)和挂钟时间比其「仅任务」参考值低49-74%和35-64%。这些结果为感知底层约束的智能体规划建立了受控的概念验证:一个最小的执行约束会在生成的程序中引发主动的结构性适应,将计算从无约束分配中转移,在执行前大幅改善观测到的资源时间分布。

英文摘要

Autonomous AI agents increasingly select actions in environments whose memory, execution-time, runtime, compute, and operational constraints determine what counts as a suitable plan. We call the absence of this execution context from an agent's planning state substrate blindness. We test this general proposition through numerical code generation, where selected implementation choices and operational consequences are directly observable. Three frontier model configurations--Anthropic Claude Opus 5, OpenAI GPT-5.6-Sol, and Google Gemini 3.7 Flash--generate code for a high-dimensional pairwise Euclidean-distance task either from the task alone or with a 128 MB RAM and 10.0 s wall-time contract. Contract disclosure reduced measured peak process memory in 13 of 14 executable index-aligned task-only versus contract-disclosed comparisons and reduced mean wall time in all three cohorts, making execution up to 3.1x faster. Across the audited corpus, disclosure produced structural code changes including bounded blocking, float32 retention, upper-triangle traversal, and in-place or memory-mapped buffers. At a tighter 96 MB contract, independently sampled contract-disclosed cohorts achieved correct-and-within-budget outcomes of 4/5 for Claude Opus 5, 5/5 for GPT-5.6-Sol, and 3/5 for Gemini 3.7 Flash, compared with task-only outcomes of 0/5, 1/5, and 0/5; cohort mean MaxRSS and wall time were 49-74% and 35-64% lower than their task-only references. These results establish a controlled proof of concept for substrate-aware agent planning: a minimal execution contract induces proactive structural adaptation in generated programs, shifting computation away from unconstrained allocations and substantially improving observed resource-time profiles before execution.

Comments8 pages, 3 figures. Reproducibility artifacts and source-linked evaluation code: https://github.com/manu2/Context-Aware-Agent-Experiment

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑