DIRAC framework evaluation for the $\boldsymbol{Fermi}$-LAT and CTA experiments
专题命中 Agent评测 :agent(abstract);workflow(abstract)
Comments proceedings to CHEP 2013 conference : http://www.chep2013.org/
AI 大模型
智能体、工具调用、规划、工作流、多智能体和自主任务执行。
专题命中 Agent评测 :agent(abstract);workflow(abstract)
Comments proceedings to CHEP 2013 conference : http://www.chep2013.org/
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)
专题命中 Agent评测 :agent(abstract);multi-agent(abstract)
专题命中 Agent评测 :agent(abstract);planning(abstract)
Comments 20 pages
专题命中 Agent评测 :agent(abstract);multi-agent(abstract)
专题命中 Agent评测 :agent(abstract);planning(abstract)
Comments 20 pages, 3 figures, 1 Table
专题命中 Agent评测 :agent(abstract);multi-agent(abstract)
Comments Poster at the 93rd Transportation Research Board annual meeting, Washington, January 2014 - Committee number AHB45 - TRB Committee on Traffic Flow Theory and Characteristics
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)
Comments 21 pages, 15 figures, supplementary videos at https://www.youtube.com/user/BristleBotChannel
Journal ref Proc. R. Soc. A 469, 20120637 (2013)
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)
Comments 1 zip file containing 1 .tex files, and 39 .eps files. The paper (including the appendix) contains 13 Figures
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract)
Comments 15 pages, 8 figures
专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL、cs.SE
Comments 6 pages
Journal ref Proceedings of RANLP 2001 - EuroConference on Recent Advances in Natural Language Processing, September 5-7, 2001, Tzigov-Chark, Bulgaria
盲选策展人:有偏见的评判者如何在自我进化的智能体中悄然阻碍技能淘汰
机构 * AWS Generative AI Innovation Center(亚马逊云科技生成式人工智能创新中心) ; HSBC Holdings Plc., HSBC Technology Center, China(汇丰控股有限公司,汇丰科技中心,中国)
专题命中 Agent评测 :agent(abstract,comments);分类 cs.AI、cs.CL
AI总结 研究自我进化智能体中,有偏见的评判者对技能淘汰的影响。通过损坏奖励分析等方法发现,“误判通过”偏差会导致技能淘汰机制失效,此为行为安全结果。还提出廉价审计可判断评判者是否越过阈值影响智能体。
Comments Published at COLM 2026 Workshop on Agent Behavior
作为替罪羊的护栏:审计工具增强型大语言模型智能体中不忠实的安全拒绝行为
专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG;agentic(comments)
AI总结 研究工具增强型大语言模型智能体安全拒绝行为审计问题,引入轻量级黑盒审计框架,将智能体响应分类。实验发现伪造行为占主导,不忠实安全拒绝行为在基线时少,增强安全语言会显著增加该行为,还提出检测方法及治理影响。
Comments 10 pages, 3 figures. Accepted at the ACM KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI
EvolveTool-Bench:评估LLM生成的工具库作为软件 artifact 的质量
机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) ; Uber Technologies(优步科技公司)
专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.SE;agentic(comments)
AI总结 本文提出EvolveTool-Bench,通过评估LLM生成的工具库在软件工程流程中的质量,揭示任务完成率之外的软件质量风险,强调需将工具库视为首要软件 artifact。
Comments 11 pages, 4 figures; accepted at KDD 2026 Workshop on Agentic AI Evaluation and Trustworthiness
可重播的金融代理:用于工具使用LLM代理的确定性-忠实性保证框架
专题命中 Agent评测 :agentic(abstract,comments);分类 cs.AI、cs.CL
AI总结 本文提出DFAH框架,用于评估金融代理的确定性和忠实性,发现决策确定性与准确性无显著相关性,强调需独立测量两者的多维评估方法。
Comments 27 pages, 5 figures, 9 tables | Code and data: https://github.com/ibm-client-engineering/output-drift-financial-llms | To appear in the 2nd ICLR Workshop on Advances in Financial AI: Towards Agentic and Responsible Systems (ICLR 2026)
离线学习纳什稳定的联盟结构(可能有重叠的联盟)
机构 * Bar Ilan University(巴伊兰大学) ; University of Oxford(牛津大学)
专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)
AI总结 本文提出了一种允许重叠联盟的离线学习模型,通过分析代理级和联盟级效用反馈,设计样本高效的算法以推断纳什稳定的联盟结构。
Comments To Appear in the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2026
机构 * University of Pennsylvania(宾夕法尼亚大学)
专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL;AI agent(comments)
Comments 9 pages, 5 figures. COLM 2025 Workshop on AI Agents
专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)
Comments Accepted at International Conference on Autonomous Agents and Multiagent Systems (AAMAS) Workshop, 2025
专题命中 Agent评测 :AI agent(abstract);分类 cs.AI、cs.LG;agent(comments)
Comments 10 pages, To be published in the International Conference on Human-Agent Interaction (HAI '24) proceedings
专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)
Comments Accepted as a full paper to the 23rd International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2024)
专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)
Comments To appear in the Proceedings of the 18th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2019). (Extended Abstract)
基于强化学习的道德智能体的元规范理论
机构 * University of Luxembourg(卢森堡大学) ; University of Bergen(卑尔根大学)
专题命中 Agent评测 :agent(abstract,comments);分类 cs.AI;multi-agent(comments)
AI总结 本文针对强化学习(RL)道德智能体设计中哲学文献被边缘化的问题,提炼元规范理论的相关理念审视RL架构,以明确RL智能体道德行为的判定标准,为评估相关RL方法奠定基础。
Comments The 14th International Workshop on Engineering Multi-Agent Systems (EMAS 2026) held May 25-26, 2026 Co-located with AAMAS 2026 Paphos, Cyprus
双层自动研究:元自动研究自身
机构 * Independent Researcher(独立研究者)
专题命中 Agent评测 :agent(abstract);分类 cs.AI;AI agent(comments);agentic(comments)
AI总结 提出双层自动研究框架,外层循环通过读取内层循环代码和轨迹、识别瓶颈并注入可执行Python搜索机制来改进内层循环,在GPT预训练基准上实现5倍改进。
Comments 16 pages, 5 figures, 3 tables. v2 expands the framing as mechanism-level agentic self-improvement and updates related work and limitations; core method and experiments unchanged. This paper was primarily drafted by AI agents with human oversight and direction
多机器人协同中的目的地到传送口任务映射优化
机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)
专题命中 Agent评测 :planning(abstract);分类 cs.AI;agent(comments);multi-agent(comments)
AI总结 本文提出基于进化算法和混合整数线性规划的任务映射优化方法,用于提升多机器人分拣系统的吞吐量。
Comments Accepted to IEEE International Symposium on Multi-Robot and Multi-Agent Systems (MRS) 2025
机构 * Department of Psychology, University of Wisconsin-Madison(威斯康星大学麦迪逊分校心理学系) ; Department of Electrical and Computer Engineering, University of Wisconsin-Madison(威斯康星大学麦迪逊分校电气与计算机工程系) ; Department of Psychological Science, University of California, Irvine(加州大学伊文斯分校心理学科学系) ; Department of Psychology, University of Michigan, Ann Arbor(密歇根大学安娜堡分校心理学系)
专题命中 Agent评测 :agent(abstract,comments);分类 cs.LG;multi-agent(comments)
Comments ICML 2025 Workshop on Multi-Agent Systems in the Era of Foundation Models: Opportunities, Challenges and Futures
专题命中 Agent评测 :agent(abstract,comments);分类 cs.LG;multi-agent(comments)
Comments In proceedings of the Reincarnating Reinforcement Learning (RRL) Workshop at ICLR 2023 and the Neuro-Symbolic AI for Agent and Multi-Agent Systems (NeSyMAS) Workshop at AAMAS 2023
专题命中 Agent评测 :agent(abstract,journal_ref);分类 cs.AI;multi-agent(journal_ref)
Journal ref Proceedings of the 18th International Conference on Principles and Practice of Multi-Agent Systems (PRIMA 2015). pp 20-35. Lecture Notes in Computer Science, vol 9387. Springer
从准确性到可审计性:金融AI系统中的确定性综述
专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.LG
AI总结 本文从系统视角综述了金融AI中表格模型、图网络和基于LLM的智能体工作流三种模态的不可重现性问题,通过实验量化了确定性指标并提出了分层评估框架。
演化不确定性下的鲁棒风险:熵值风险(Entropic Value-at-Risk)的Wasserstein对应
机构 * Technical University of Munich(慕尼黑工业大学) ; Masaryk University(马萨里克大学)
专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG
AI总结 该研究针对演化不确定性下的鲁棒风险问题,提出Wasserstein熵值风险,弥补熵值风险无法对冲名义模型认定不可能的灾难的缺陷,经数值验证其变分对偶性,并构造出随信念变化的闭式鲁棒动态规划算子。
Comments Best Paper Award at the 2nd Workshop on Safe AI at UAI (SafeAI@UAI 2026, non-archival), Amsterdam