DeepRNG: Towards Deep Reinforcement Learning-Assisted Generative Testing of Software
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.LG、cs.SE
Comments Workshop on ML for Systems, 35th Conference on Neural Information Processing Systems (NeurIPS 2021)
AI 大模型
智能体、工具调用、规划、工作流、多智能体和自主任务执行。
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.LG、cs.SE
Comments Workshop on ML for Systems, 35th Conference on Neural Information Processing Systems (NeurIPS 2021)
专题命中 软件智能体 :agent(abstract);multi-agent(abstract)
专题命中 软件智能体 :workflow(abstract);分类 cs.AI、cs.CL、cs.LG
Comments Accepted for publication at TACL. This version is a pre-MIT Press publication version
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL、cs.LG
Comments ACL-IJCNLP 2021 (demo paper)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL、cs.LG
Comments Contains main article and supplementaries
Journal ref Neurips 2021
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.LG、cs.SE
Comments 15 pages, 7 figures
专题命中 软件智能体 :agent(abstract);AI agent(abstract)
Comments Code is available at: https://github.com/amazon-research/progressive-coordinate-transforms
专题命中 软件智能体 :agent(abstract);autonomous agent(abstract)
Comments SAIAD CVPR21 Workshop
专题命中 软件智能体 :agent(abstract);multi-agent(abstract)
Comments 15 pages, 9 figures, journal
专题命中 软件智能体 :agent(abstract);AI agent(abstract)
Comments arXiv admin note: text overlap with arXiv:2003.10286
专题命中 软件智能体 :agent(abstract);autonomous agent(abstract)
Journal ref IEEE Transactions on Robotics, 2020
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL、cs.LG
Comments Accepted at WACV 2019. Also at NeurIPS 2017 workshop on Visually-Grounded Interaction and Language (ViGIL)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL、cs.LG
Comments Published in Advances in Neural Information Processing Systems (NIPS) 30, December 2017
Journal ref Rothe, A., Lake, B. M., and Gureckis, T. M. (2017). Question asking as program generation. Advances in Neural Information Processing Systems 30
专题命中 软件智能体 :agent(abstract);autonomous agent(abstract)
专题命中 软件智能体 :agent(abstract);multi-agent(abstract)
Journal ref The World of Computer Science and Information Technology Journal (WSCIT). 2014, Volume 4, Issue 2. pp. 18.25
专题命中 软件智能体 :agent(abstract);multi-agent(abstract)
Comments 6 pages
Journal ref IJCSI International Journal of Computer Science Issues, Vol. 10, Issue 2, No 3, March 2013
专题命中 软件智能体 :agent(abstract);planning(abstract)
Comments 12 Pages, 3 Tables, 3 Figures
Journal ref Proc. European Spreadsheet Risks Int. Grp. (EuSpRIG) 2003 147-159 ISBN 1 86166 199 1
跨模型大语言模型代码审查:应该用Claude审查Codex还是相反?
机构 * University of California, Davis(加州大学戴维斯分校) ; Johns Hopkins University(约翰霍普金斯大学) ; California State University, Long Beach(长滩加州州立大学)
专题命中 软件智能体 :workflow(abstract);分类 cs.AI、cs.SE;agentic(comments)
AI总结 研究开发者同时使用Claude和Codex进行代码审查的成本、时间及配对顺序问题,通过对116个任务的六种条件实验发现,Claude审查Codex草稿效果好,反向则不佳,有用的配对是不对称的,应Claude审查Codex。
Comments This paper had been accepted by Agentic SE @ KDD'26
用强化学习建模客户轨迹以获得实际零售洞察
机构 * McGill University(麦吉尔大学) ; Mila - Quebec AI Institute(魁北克人工智能研究所)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)
AI总结 本文提出了一种基于智能体的建模框架,将客户轨迹预测转化为最大熵强化学习问题,以更准确地反映具有有限理性的客户行为,从而提供更精确的冲动购买率和货架交通密度估计。
Comments Proceeding of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)
Banana100: 通过100次迭代图像复制打破NR-IQA度量标准
机构 * University of California, Santa Barbara(加州大学圣塔芭芭拉分校)
专题命中 软件智能体 :agentic(abstract,comments);分类 cs.AI、cs.LG
AI总结 Banana100通过100次迭代编辑生成28000张退化图像,揭示多轮编辑中图像质量退化问题,发现现有NR-IQA度量标准无法检测严重退化图像,威胁未来模型训练稳定性。
Comments Accepted to CVPR 2026 Workshop on Agentic AI for Visual Media
基于反射的可信代码代理控制
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.SE;agentic(comments)
AI总结 本文提出反射驱动控制方法,通过内部反思循环提升代码生成的安全性和合规性,实现自主、安全且可审计的AI代码代理。
Comments Accepted to AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)
机构 * LinkedIn(领英)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL;agentic(comments)
Comments 11 pages, 8 figures, Workshop on Agentic AI for Enterprise at KDD '25
专题命中 软件智能体 :agent(abstract,comments);分类 cs.AI、cs.CL
Comments Code, data, and over 24k agent trajectories are released at https://github.com/Leezekun/SOPBench
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)
Comments To appear in 18th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2019) as a full paper. arXiv admin note: substantial text overlap with arXiv:1806.08055
超越满意:从安慰性到可操作性解释以提升可理解性
机构 * The University of Tulsa(图兰大学)
专题命中 软件智能体 :agent(abstract,journal_ref);分类 cs.AI;multi-agent(journal_ref)
AI总结 本文探讨了可解释性在提升系统可理解性中的作用,通过实验发现可操作性解释在任务表现上优于安慰性解释,但用户满意度评分相同,强调需结合客观指标与主观评估来衡量解释质量。
Comments 21 pages, 7 figures, 6 tables. EXTRAAMAS 2025 submission. Preprint version
Journal ref In: Calvaresi, D., et al. Explainable, Trustworthy, and Responsible AI and Multi-Agent Systems. EXTRAAMAS 2025. Lecture Notes in Computer Science. Springer, Cham
机构 * Center for Modeling Social Systems(社会科学建模中心) ; NORCE Norwegian Research Center AS(挪威NORCE研究机构) ; Kristiansand, Norway(挪威克里斯蒂安桑)
专题命中 软件智能体 :agent(abstract,comments);分类 cs.AI;multi-agent(comments)
Comments 13 pages, 3 figures, 23rd International Conference on Practical applications of Agents and Multi-Agent Systems (PAAMS 2025)
探索时持异议,提交时持共识:面向软件智能体的路由引导测试时缩放
机构 * Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 软件智能体 :tool-use(abstract);分类 cs.AI、cs.SE
AI总结 该研究针对软件智能体测试时缩放难题,提出Risa方法,利用MoE路由轨迹引导探索与仲裁,在SWE-bench等基准上提升了仓库级软件工程任务的解决率。
Repo2Skill-Evo:仓库技能在静默中过时
机构 * ByteDance(字节跳动) ; Peking University(北京大学) ; Beijing Jiaotong University(北京交通大学)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.SE
AI总结 Repo2Skill-Evo研究发现,在57个真实仓库的105次版本转换中,前沿智能体难以可靠维护仓库技能,其平均@3 macro F1仅29.9%-69.7%,仓库技能会在无明确信号的情况下静默过时。
GameXpert-Bench:编码智能体距离专家级游戏开发还有多远?
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Tsinghua University(清华大学) ; The Hong Kong University of Science and Technology(香港科技大学) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Lightspeed Studios, Tencent(腾讯光速工作室)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL
AI总结 GameXpert-Bench是覆盖游戏开发全生命周期的基准,含三个赛道,测评显示当前编码智能体在生成可玩游戏基础上表现较好,在缺陷发现等方面仍有不足。
具备权威的AI:从应用到硅片
专题命中 软件智能体 :AI agent(abstract);分类 cs.AI、cs.SE
AI总结 该研究提出Salt方法,借助生成式AI与机器验证,在五周内由一名研究人员指挥AI智能体完成RISC-V处理器的流片,无人工审核证明与RTL,实现高效可靠的自主机器工作。
Comments 17 pages, 6 figures