arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-01-08 至 2026-01-08 共收录 4 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 软件智能体 4 篇

2601.03556 2026-01-08 cs.SE 88%

Do Autonomous Agents Contribute Test Code? A Study of Tests in Agentic Pull Requests

自主代理是否贡献测试代码?对代理拉取请求中测试的调查

Sabrina Haque, Sarvesh Ingale, Christoph Csallner

专题命中 软件智能体 :agentic(title,abstract);autonomous agent(title);agent(abstract);分类 cs.SE

AI总结 研究分析了代理拉取请求中测试代码的出现频率及影响因素,揭示了测试代码对PR规模和处理时间的影响,为自主软件开发研究提供实证依据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04171 2026-01-08 cs.LG 83%

Agentic Rubrics as Contextual Verifiers for SWE Agents

代理 rubrics 作为 SWE 代理的上下文验证器

Mohit Raghavendra, Anisha Gunjal, Bing Liu, Yunzhong He

专题命中 软件智能体 :agentic(title,abstract);agent(abstract);分类 cs.LG

AI总结 代理 rubrics 通过上下文感知的检查清单提升 SWE 代理的验证效率与准确性。

Comments 31 pages, 11 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03335 2026-01-08 cs.AI cs.NE 70%

Digital Red Queen: Adversarial Program Evolution in Core War with LLMs

数字红皇后:基于大语言模型的Core War中的对抗程序进化

Akarsh Kumar, Ryan Bahlous-Boldi, Prafull Sharma, Phillip Isola, Sebastian Risi, Yujin Tang, David Ha

机构 * MIT(麻省理工学院) Sakana AI

专题命中 软件智能体 :agent(abstract);multi-agent(abstract);分类 cs.AI

AI总结 本文提出DRQ算法,利用大语言模型在Core War游戏中通过持续适应变化的目标进化出通用且高效的战士,揭示了动态对抗进化在人工智能系统中的潜力。

Comments 14 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03878 2026-01-08 cs.SE 57%

Understanding Specification-Driven Code Generation with LLMs: An Empirical Study Design

通过LLM理解基于规范的代码生成:一项实证研究设计

Giovanni Rosa, David Moreno-Lumbreras, Gregorio Robles, Jesús M. González-Barahona

专题命中 软件智能体 :workflow(abstract);分类 cs.SE

AI总结 本文通过实证研究设计,探讨人类干预在基于规范的LLM代码生成过程中对代码质量和动态的影响。

Comments This paper is a Stage 1 Registered Report. The study protocol and analysis plan were peer reviewed and accepted at SANER 2026 with a Continuity Acceptance (CA) score for Stage 2

详情

展开后加载摘要…

URL PDF HTML 收藏