Emergent Translation in Multi-Agent Communication
专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI
Comments Accepted to ICLR 2018
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI
Comments Accepted to ICLR 2018
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI
Comments 10 pages, RoboNLP Workshop from ACL Conference
专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI
Comments 11 pages, 5 appendix pages, 11 figures, 3 tables, under review as a conference paper at ICLR 2017
专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL
Comments 9 pages, manuscript under submission
AutoTool: 面向智能体推理的动态工具选择与集成
机构 * Nanyang Technological University(南洋理工大学)
专题命中 多模态Agent :multimodal(abstract);分类 cs.CL;multi-modal(comments)
AI总结 提出AutoTool框架,通过双阶段优化(SFT+RL轨迹稳定化和KL正则化Plackett-Luce排序)使大语言模型具备动态工具选择能力,在数学、科学、代码和多模态推理等任务上平均提升6.4%-7.7%。
Comments ICML2026; Best Paper Award at ICCV 2025 Workshop on Multi-Modal Reasoning for Agentic Intelligence
专题命中 多模态Agent :multimodal(abstract,comments);分类 cs.AI
Comments CCC Blue Sky Ideas paper, published at the ACM International Conference on Multimodal Interaction (ICMI '22), November 7-11, 2022
专题命中 多模态Agent :multimodal(abstract,comments);分类 cs.CL
Comments 6 pages, ICLR 2021 Embodied Multimodal Learning Workshop
专题命中 多模态Agent :multimodal(abstract,comments);分类 cs.CV
Comments Accepted to 2019 International MICCAI Brainlesion Workshop -- Multimodal Brain Tumor Segmentation Challenge (BraTS) 2019. arXiv admin note: substantial text overlap with arXiv:1810.11654
专题命中 多模态Agent :multimodal(abstract,journal_ref);分类 cs.CV
Journal ref ICMI 2018 - Workshop at 20th ACM International Conference on Multimodal Interaction, Oct 2018, Boulder, Colorado, United States. pp.1-13
MetaCaster:用于轻量级时间序列预测器小样本端到端学习的元调控优化智能体
机构 * University of Houston(休斯顿大学) ; NEC Labs(NEC实验室) ; University of Waterloo(滑铁卢大学) ; University of Connecticut(康涅狄格大学) ; University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Singapore Management University(新加坡管理大学)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 针对资源受限场景下轻量级时间序列预测器小样本学习的困境,提出MetaCaster多智能体框架,可高效训练专用预测器,在18个数据集等实验中兼顾数据、计算效率与预测性能。
Comments Accepted by EMNLP 2026
EMPIRE:将显式操作规划作为可学习中间表征用于自我中心视角下手运动预测
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 该研究针对现有手运动预测方法忽略操作过程、梯度干扰的问题,提出两阶段框架EMPIRE,构建EMPIRE-651K数据集,实现了更优的手运动预测精度。
Comments 14 pages, 10 figures, 18 tables
超越成功与失败:面向GUI智能体的长度感知对比学习
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Baidu Inc.(百度公司) ; University of Alberta(阿尔伯塔大学)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 针对GUI智能体现有对比RLVR方法无法捕获轨迹细粒度质量差异的问题,提出LACL-GUI框架,引入轨迹级质量信号,在基准测试中实现性能提升。
自进化科学智能体发现可泛化的物理推理流体控制
机构 * National University of Singapore(新加坡国立大学)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 提出一种由大语言模型驱动的自进化科学智能体工作流,通过迭代代码生成和物理仿真诊断,自动构建可解释的控制器,并在欠驱动双关节狗鲨游泳器目标到达任务中实现零样本泛化。
MindClaw: 用于精确干预的闭环具身心理状态推理
机构 * Jilin University(吉林大学) ; Microsoft Asia(微软亚洲) ; National Taiwan University(国立台湾大学)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 提出MindClaw框架,通过闭环具身心理状态推理实现精确干预,结合多源输入、信念记忆、认知触发技能和动作生成,在动态环境中优化干预时机。
Comments Extended version of the CVPR 2026 paper *MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents*. This work is in progress
MAVEN-T:用于实时多智能体轨迹预测的强化异构蒸馏
机构 * School of Mathematical Sciences, Shanghai Jiao Tong University(上海交通大学数学科学学院) ; Bio-X Institutes, Key Laboratory for the Genetics of Developmental and Neuropsychiatric Disorders, Shanghai Jiao Tong University(上海交通大学Bio-X研究院、发育与神经精神疾病遗传学重点实验室) ; Shanghai Key Laboratory of Psychotic Disorders, Brain Science and Technology Research Center, Shanghai Jiao Tong University(上海精神疾病重点实验室、脑科学与技术研究中心,上海交通大学)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 提出MAVEN-T框架,通过高容量教师模型和紧凑学生模型的异构蒸馏,结合强化学习优化,实现实时多智能体轨迹预测,在多个数据集上达到高精度与低延迟。
AEGIS:防止MCP中的跨域资源滥用
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 本文提出AEGIS,该组件借助LLM将MCP工具调用归一化为统一表示,结合Open Policy Agent与ContextForge AI Gateway,可防止跨异构MCP工具和模态的资源滥用。
大规模AI手写物理评估评分:评分一致性与奥林匹克团队选拔结果
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 本研究评估基于GPT-5.5的AI对三类物理手写评估的评分表现,其与官方评分相关性高且能准确选拔奥林匹克团队,第二轮评分优化了一致性,AI评分可作为考官控制下的辅助评分工具。
迈向通用具身智能:整合大语言模型、知识库与推理能力以构建下一代AI智能体
机构 * College of Mechanical Engineering, Chongqing University of Technology(重庆理工大学机械工程学院) ; James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院) ; School of Energy and Power, Jiangsu University of Science and Technology(江苏科技大学能源与动力学院) ; Magnesium Research Center, Kumamoto University(熊本大学镁研究中心) ; Department of Information and Communication Engineering, Nagoya University(名古屋大学信息与通信工程系) ; State Key Laboratory of Fluid Power and Mechatronic Systems, Zhejiang University(浙江大学流体动力与机电系统国家重点实验室)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 本文综述以LLM为核心的智能系统演进,提出整合LLMs、KBs、RA与具身性的概念框架,明确高效LLM部署等五大挑战,为开发复杂动态环境下的自适应多模态智能体提供路线图。
Journal ref International Journal of Hydromechatronics 9(2) (2026) 250-316
基于无人机的基础设施检查:面向AEC+FM的文献综述与提出框架
专题命中 多模态Agent :multimodal(abstract);分类 cs.CV
AI总结 本文综述了无人机在基础设施检查中的应用,提出框架整合多模态数据与Transformer架构,以提高检测准确性和可靠性,未来研究方向包括轻量级AI模型和合成数据集。
Comments Accepted for publication in the Proceedings of the International Conference on Computing in Civil Engineering (i3CE 2025)
Teach a Molmo2Fish:面向自然语言引导的交互式鱼类追踪
机构 * Massachusetts Institute of Technology(麻省理工学院)
专题命中 多模态Agent :multimodal(abstract);分类 cs.CV
AI总结 本研究针对声呐鱼类追踪数据集定制了多模态大语言模型工具Molmo2Fish,通过交互式预测修正工作流开展实验,发现其在鱼类追踪和轨迹修正上性能良好,但自然语言引导的融入仍需提升。
Comments 29 pages, 6 figures, to be published in Third Workshop on Computer Vision for Ecology at ECCV 2026
基于视觉语言模型的X射线荧光(XRF)与光学显微镜视场(FOV)定位
机构 * Northwestern University(西北大学) ; University of Chicago(芝加哥大学) ; Oregon Health and Science University(俄勒冈健康与科学大学) ; Argonne National Laboratory(阿贡国家实验室)
专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV
AI总结 本文针对跨模态显微图像的视场定位难题,提出结合视觉语言模型(VLM)的候选生成-验证工作流,在低对应度的相邻切片成像数据中实现了有效定位,为关联XRF与光学显微测量提供了支撑。
DeAR:基于能力锚定与协作思维导航的去中心化智能体推理
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 针对现有智能体推理系统的路由瓶颈与静态角色分配问题,提出DeAR框架,通过三种机制实现去中心化协作,在9类基准测试中性能优于基线方法,提升了知识密集型推理任务的准确性。
MistyPilot:通过多智能体大语言模型技能编排实现社交机器人控制
机构 * State University of New York at Buffalo(纽约州立大学布法罗分校)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 该研究提出多智能体大语言模型框架MistyPilot,可解释自然语言指令并编排Misty社交机器人技能,经评估其在多项任务上准确率高、方差低,用户反馈积极,代码将公开。
Comments Accepted at the ECCV 2026 ACVR Workshop
ETHOS:面向临床多智能体系统的模块化伦理框架
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 ETHOS是可与现有临床多智能体系统集成的模块化伦理框架,通过分层治理提升决策可靠性,将AI伦理原则转化为可部署的安全保障。
Comments Preprint of an article submitted for consideration in Pacific Symposium on Biocomputing \textcopyright\ 2027 World Scientific Publishing Company. \url{https://psb.stanford.edu/}
ForceU-VLA:用于实体超声扫描的力感知视觉-语言-动作模型
机构 * Faculty of Computer Science and Technology, Ocean University of China(中国海洋大学计算机科学与技术学院) ; School of Information Science and Engineering, Shandong University(山东大学信息科学与工程学院) ; Innovation School of Artificial Intelligence, Hefei University of Technology(合肥工业大学人工智能创新学院)
专题命中 多模态Agent :multimodal(abstract);分类 cs.CV
AI总结 本文针对现有实体超声扫描方法的不足,提出力感知视觉-语言-动作模型ForceU-VLA,设计FUSFM与SAMM模块,构建ForceU-VLA-Data数据集,实验证实其可提升超声扫描的接触稳定性与压力调节能力。
视觉语言模型(VLMs)无法规划,但它们能进行形式化吗?
机构 * Drexel University(德雷塞尔大学) ; University of Pennsylvania(宾夕法尼亚大学) ; Johns Hopkins University(约翰霍普金斯大学)
专题命中 多模态Agent :multimodal(abstract);分类 cs.CL
AI总结 本研究提出5种VLM作为形式化器的流水线,评估后发现其表现优于端到端规划生成,较弱VLMs的瓶颈为物体关系视觉grounding,较强模型已克服该限制。
AdvDex:通过关节对齐动作与对抗学习从人类演示中学习灵巧操作
机构 * Zhejiang University(浙江大学) ; Shanghai Innovation Institute(上海创新研究院) ; Fudan University(复旦大学) ; Shanghai Jiao Tong University(上海交通大学) ; Paxini Tech(帕西尼科技)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 AdvDex是一种统一视觉-语言-动作框架,通过OmniShare数据集、JAAS动作空间与领域对抗学习,实现从人类和机器人演示中学习灵巧操作,提升跨实体泛化与技能迁移能力。
PlayWorld:基于智能体玩家的长程目标世界模型基准测试
机构 * The Chinese University of Hong Kong(香港中文大学) ; The University of Hong Kong(香港大学) ; Zhejiang University(浙江大学) ; Kuaishou Technology(快手科技)
专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV
AI总结 该研究针对现有世界模型跨模型公平比较的难题,推出含171个场景的PlayWorld基准,通过多模态智能体玩家从多维度评估9种先进世界模型,发现其长程交互式目标表现仍不可靠。
Comments project page: https://kxding.github.io/project/PlayWorld/
LLM-Advisor: 一种用于多地形高效路径规划的LLM基准
机构 * Graduate School of Information Science and Technology, Hokkaido University(弘前大学信息科学与技术研究生院) ; Department of Information and Communication Engineering, The University of Tokyo(东京大学信息与通信工程系)
专题命中 多模态Agent :multimodal(abstract);分类 cs.AI
AI总结 LLM-Advisor利用大型语言模型优化多地形路径规划,提升路径成本效率,适用于现实场景。
Comments This paper has been accepted by IEEE Transactions on Automation Science and Engineering
Lines and Ladders:面向大规模零售价格分类的上下文感知多智能体框架
机构 * Walmart Global Tech(沃尔玛全球科技)
专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI
AI总结 针对大规模零售商品定价管理难题,提出上下文感知多智能体框架自动化构建Lines and Ladders价格分类,3智能体系统在Lines任务F1达0.83,在多品类数据上表现优异且已投入生产。
Comments 8 pages. Accepted in the Main Conference of IEEE ICMLA 2026