arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-02-16 至 2026-02-16 共收录 11 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 多智能体 11 篇

2602.12520 2026-02-16 cs.LG cs.MA 89%

Multi-Agent Model-Based Reinforcement Learning with Joint State-Action Learned Embeddings

基于联合状态-动作学习嵌入的多智能体模型驱动强化学习

Zhizun Wang, David Meger

机构 * McGill University(麦吉尔大学)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);planning(abstract);分类 cs.LG

AI总结 本文提出了一种基于联合状态-动作学习嵌入的多智能体模型驱动强化学习框架,通过统一表示学习与想象式回放,提升智能体在动态环境中协调能力与长期规划效率。

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18138 2026-02-16 cs.LG 88%

B3C: A Minimalist Approach to Offline Multi-Agent Reinforcement Learning

B3C: 一种针对离线多智能体强化学习的极简方法

Woojun Kim, Katia Sycara

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.LG;autonomous agent(comments)

AI总结 B3C通过引入批评者剪裁和非线性价值分解,有效解决离线多智能体强化学习中的过估计问题,提升性能。

Comments Accepted at the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13059 2026-02-16 cs.CL 88%

TraceBack: Multi-Agent Decomposition for Fine-Grained Table Attribution

TraceBack: 多智能体分解用于细粒度表格归因

Tejas Anvekar, Junha Park, Rajat Jha, Devanshu Gupta, Poojah Ganesan, Puneeth Mathur, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学) Adobe Research(Adobe研究院)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.CL

AI总结 TraceBack通过多智能体框架实现细粒度表格归因,提供可验证的单元格支持证据,提升问题回答的透明度和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00602 2026-02-16 cs.LG cs.SY eess.SY 88%

Multi-Agent Stage-wise Conservative Linear Bandits

多智能体分阶段保守线性老虎机

Amirhossein Afsharrad, Ahmadreza Moradipari, Sanjay Lall

机构 * Stanford University(斯坦福大学) University of California, Santa Barbara(加州大学圣巴巴拉分校)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.LG

AI总结 本文提出MA-SCLUCB算法,通过分阶段保守约束实现多智能体在安全保证下的高效分布式学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22890 2026-02-16 cs.CV cs.CR 88%

CP-uniGuard: A Unified, Probability-Agnostic, and Adaptive Framework for Malicious Agent Detection and Defense in Multi-Agent Embodied Perception Systems

CP-uniGuard: 多智能体具身体验系统中恶意代理检测与防御的统一、概率无关和自适应框架

Senkang Hu, Yihang Tao, Guowen Xu, Xinyuan Qian, Yiqin Deng, Xianhao Chen, Sam Tak Wu Kwong, Yuguang Fang

机构 * Hong Kong JC STEM Lab of Smart City and Department of Computer Science, City University of Hong Kong(香港JC STEM实验室及城市大学计算机科学系) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) Department of Electrical and Electronic Engineering, The University of Hong Kong(香港大学电子与电气工程系) School of Data Science, Lingnan University(岭南大学数据科学学院)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract)

AI总结 CP-uniGuard通过概率无关的样本共识和自适应阈值,实现多智能体系统中恶意代理的检测与防御。

Comments Accepted by IEEE Transactions on Mobile Computing (TMC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12953 2026-02-16 cs.HC 78%

Human Tool: An MCP-Style Framework for Human-Agent Collaboration

Human Tool: 一种MCP风格的人机协作框架

Yuanrong Tang, Huiling Peng, Bingxi Zhao, Hengyang Ding, Hanchao Song, Tianhong Wang, Chen Zhong, Jiangtao Gong

专题命中 多智能体 :agent(title,abstract)

AI总结 Human Tool通过MCP风格框架实现人机协作平衡,利用结构化工具方案提升任务表现并减少人类负担。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07117 2026-02-16 cs.AI cs.LG 73%

The Conditions of Physical Embodiment Enable Generalization and Care

物理具身的条件使泛化和关怀成为可能

Leonardo Christov-Moore, Arthur Juliani, Alex Kiefer, Joel Lehman, Nicco Reggente, B. Scot Rousse, Adam Safron, Nicolás Hinrichs, Daniel Polani, Antonio Damasio

机构 * Institute for Advanced Consciousness Studies(先进意识研究所) Monash Centre for Consciousness and Contemplative Studies(莫纳什意识与冥想研究中心) University of Oxford(牛津大学) Topos Institute(拓斯研究所) Allen Discovery Center(艾伦发现中心) Okinawa Institute of Science and Technology(冲绳科学技术研究所) Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所) University of Hertfordshire(赫特福德郡大学) Brain and Creativity Institute(大脑与创造力研究所)

专题命中 多智能体 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出物理具身的条件是泛化和关怀的基础,通过稳态驱动和因果建模实现智能体在开放环境中的鲁棒性和可信对齐。

Comments 15 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21164 2026-02-16 cs.MA 67%

Learning Large-Scale Competitive Team Behaviors with Mean-Field Interactions and Online Opponent Modeling

通过均场交互和在线对手建模学习大规模竞争团队行为

Bhavini Jeloka, Yue Guan, Panagiotis Tsiotras

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 MF-MAPPO通过均场交互和在线对手建模,实现了大规模竞争团队行为的学习,有效整合团队内合作与团队间竞争,展示了在大规模场景下的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12583 2026-02-16 cs.GT 67%

Opinion dynamics and mutual influence with LLM agents through dialog simulation

基于对话模拟的LLM代理意见动态与相互影响

Yulong He, Dutao Zhang, Sergey Kovalchuk, Pengyi Li, Artem Sedakov

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 本文提出利用LLM代理进行对话模拟,以解决意见动态研究中现实数据不足的问题,通过模拟锚定效应和多代理互动,提供可扩展的意见形成分析工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12502 2026-02-16 cs.MA 67%

Building Large-Scale Drone Defenses from Small-Team Strategies

构建大规模无人机防御的小组策略

Grant Douglas, Stephen Franklin, Claudia Szabo, Mingyu Guo

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 本文提出通过整合小规模防御策略模块化组件,构建大规模无人机防御策略,实现高效扩展和有效合作行为发现。

Comments 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12830 2026-02-16 cs.GT cs.MA 50%

Decentralized Optimal Equilibrium Learning in Stochastic Games via Single-bit Feedback

在随机博弈中通过单比特反馈实现去中心化最优均衡学习

Seref Taha Kiremitci, Ahmed Said Donmez, Muhammed O. Sayin

专题命中 多智能体 :agent(abstract)

AI总结 本文提出了一种在随机博弈中通过单比特反馈实现去中心化最优均衡学习的方法,通过探索-承诺和在线变体,实现了对社会福利目标的优化,并建立了有限时间遗憾保证。

详情

展开后加载摘要…

URL PDF HTML 收藏