CTD4 -- A Deep Continuous Distributional Actor-Critic Agent with a Kalman Fusion of Multiple Critics
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG
AI 大模型
智能体、工具调用、规划、工作流、多智能体和自主任务执行。
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG
专题命中 Agent评测 :planning(title);分类 cs.AI、cs.CL
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG
Comments https://cybercapabilities.org/
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL
Comments EMNLP 2024 Main
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL
专题命中 Agent评测 :agent(abstract,comments);multi-agent(abstract,comments);分类 cs.AI、cs.LG
Comments Proc. of the Main Track of 22nd International Conference on Practical Applications of Agents and Multi-Agent Systems, 26th-28th June, 2024, https://www.paams.net/. Includes 6 figures, 1 table and 32 references
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.SE
Comments 2 pages, "JFMS 2020 -- Les Journees Francophones de la Modelisation et de la Simulation -- Convergences entre la Theorie de la Modelisation et la Simulation et les Systemes Multi-Agents"
专题命中 Agent评测 :AI agent(title);分类 cs.AI、cs.CL
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG
Comments 32 pages, 6 figures, ICLR 2023
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG
Comments 33 pages
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG
Comments This paper was published In Proceeding of NeurIPS 2022
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG
Comments The Second International Conference on AIML Systems, October 12--15, 2022, Bangalore, India
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL
Comments ACL 2022 Insights Workshop (6 pages)
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL
Comments 7 pages, 3 figures, 3 Tables. Accepted to be presented at COLING 2020 conference: https://coling2020.org/pages/accepted_papers_industry_track
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG
Comments Deep Reinforcement Learning Symposium, NIPS 2017
专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL
Comments 10 pages. In Proceedings of Conference on Human Computation & Crowdsourcing (HCOMP 2016), 2016, Austin, TX, USA
小型语言模型用于高效的代理工具调用:通过针对性微调超越大模型
专题命中 Agent评测 :agentic(title,comments);分类 cs.AI
AI总结 本文通过针对性微调小型语言模型,在工具调用任务中超越大模型,展示了SLMs在成本优化和效率提升方面的潜力。
Comments Accepted at AAAI 2026 Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks
专题命中 Agent评测 :agent(title,comments);分类 cs.LG
Comments The main results in this draft have been subsumed by our later paper "Persuading a Learning Agent"
专题命中 Agent评测 :agent(title);分类 cs.LG;autonomous agent(comments)
Comments -Updated to the most recent and completed version (to be presented at AAMAS 2015) -Updated author list. in Autonomous Agents and Multiagent Systems (AAMAS) 2015, Istanbul, Turkey, May 2015
专题命中 Agent评测 :agent(title);分类 cs.CL;autonomous agent(journal_ref)
Comments 10 pages, uses aaai.sty, lingmacros.sty, psfig.sty
Journal ref Proceedings of the First International Conference on Autonomous Agents, Marina del Rey, California, USA. 1997. pp 96-105
工具可及性对LLM代理安全对齐的因果影响
机构 * Cardiff Metropolitan University(卡地夫 Metropolitan 大学) ; Harvard Medical School(哈佛医学院) ; Clark University(克拉克大学) ; Harvard University(哈佛大学)
专题命中 Agent评测 :agent(abstract);agentic(abstract);分类 cs.AI、cs.LG、cs.SE
AI总结 研究通过对比文本聊天机器人与具备工具访问权限的代理行为,发现工具可及性显著增加安全违规率,表明文本评估不足以评估代理系统。
连贯性更弱,互惠性更弱:对比Moltbook与Reddit的语义及社交组织
专题命中 Agent评测 :agent(abstract);AI agent(abstract);autonomous agent(abstract)
AI总结 该研究对比AI智能体社交网络Moltbook与早期Reddit,发现Reddit在语义连贯性、多样性及交互互惠性上更优,且未被Moltbook复现。
VICBench:一个用于代码漏洞检测的多语言基准测试集
专题命中 Agent评测 :workflow(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.SE
AI总结 本研究构建了多语言代码漏洞检测基准VICBench,包含100个CVE对应的VIC,规模显著大于现有数据集,经评估现有漏洞检测算法性能有限,该基准可用于可靠评估漏洞检测方法。
评估自主编程代理中的计划合规性
机构 * University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校) ; IBM(IBM公司)
专题命中 Agent评测 :agent(abstract,abstract_cn);分类 cs.AI、cs.CL、cs.SE
AI总结 本文系统分析了编程代理在执行任务时对计划的遵循情况,通过16991条轨迹评估不同计划变体的影响,发现计划提醒能提高任务成功率,但不恰当的计划补充可能降低性能。
HarnessOpt-Bench:评估大型语言模型的 harness 优化能力
机构 * Scale AI
专题命中 Agent评测 :agent(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.LG
AI总结 本研究提出HarnessOpt-Bench基准,评估前沿LLM在高成本随机评估下的端到端harness优化能力,发现优化器模型差异大于编码harness、原生harness非始终更优,该能力具可测性与提升空间。
面向AI安全评估的对抗语用学:指令冲突、嵌入命令与策略模糊性基准
机构 * Humber Polytechnic(汉博理工学院) ; University of Toronto(多伦多大学)
专题命中 Agent评测 :agent(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.SE
AI总结 提出对抗语用学基准和标注协议,通过语言学控制的分类法评估模型在指令冲突、嵌入命令等场景下的行为,为安全评估提供实证和方法论工具。
Comments 32-page main paper plus 13-page supplement; 6 figures and 17 tables total; code and data artifact available at the linked repository
Nautilus:从一个提示到即插即用的机器人学习
机构 * TU Darmstadt(图宾根大学) ; KIT(卡尔斯鲁厄理工学院) ; FZI(弗劳恩霍夫研究所) ; Robotics Institute Germany(德国机器人研究所) ; Honda Research Institute Europe(本田欧洲研究院)
专题命中 Agent评测 :agent(abstract);workflow(abstract);agentic(abstract)
AI总结 Nautilus通过单个提示生成可复现、评估、微调和部署的机器人学习工作流程,提供即插即用的代理技能集和统一接口,减少跨家族验证的工程负担。