arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2821 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2821 篇

1906.08464 2019-06-21 cs.RO cs.AI cs.LG cs.SY eess.SY 57%

A Hierarchical Architecture for Sequential Decision-Making in Autonomous Driving using Deep Reinforcement Learning

Majid Moghadam, Gabriel Hugh Elkaim

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments Appears in ICML 2019 workshop on Real-world Sequential Decision Making: Reinforcement Learning and Beyond. Source code available in: https://github.com/MajidMoghadam2006

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.10621 2019-05-28 cs.AI 57%

Dynamic Epistemic Logic with ASP Updates: Application to Conditional Planning

Pedro Cabalar, Jorge Fandinno, Luis Fariñas del Cerro

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.04899 2019-01-16 cs.CL 57%

Conversational Intent Understanding for Passengers in Autonomous Vehicles

Eda Okur, Shachi H Kumar, Saurav Sahay, Asli Arslan Esme, Lama Nachman

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.09245 2018-11-02 cs.AI 57%

A Review on Learning Planning Action Models for Socio-Communicative HRI

Ankuj Arora, Humbert Fiorino, Damien Pellier, Sylvie Pesty

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Workshop on Affect, Artifcial Compagnon and Interaction, 2016

Journal ref Workshop on Affect, Artifcial Compagnon and Interaction, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.07225 2018-10-18 cs.RO cs.AI cs.LG 57%

Integrating kinematics and environment context into deep inverse reinforcement learning for predicting off-road vehicle trajectories

Yanfu Zhang, Wenshan Wang, Rogerio Bonatti, Daniel Maturana, Sebastian Scherer

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments CoRL 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.02647 2018-10-08 cs.AI q-bio.NC 57%

Hybrid Active Inference

André Ofner, Sebastian Stober

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.03916 2018-09-12 cs.AI 57%

Detecting Intentions of Vulnerable Road Users Based on Collective Intelligence

Maarten Bieshaar, Günther Reitberger, Stefan Zernetsch, Bernhard Sick, Erich Fuchs, Konrad Doll

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 20 pages, published at Automatisiertes und vernetztes Fahren (AAET), Braunschweig, Germany, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1607.07762 2018-08-06 cs.AI cs.LG cs.RO stat.AP stat.ML 57%

Focused Model-Learning and Planning for Non-Gaussian Continuous State-Action Systems

Zi Wang, Stefanie Jegelka, Leslie Pack Kaelbling, Tomás Lozano-Pérez

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.07035 2018-05-21 cs.CG cs.AI cs.GR cs.RO 57%

Automated Process Planning for Hybrid Manufacturing

Morad Behandish, Saigopal Nelaturi, Johan de Kleer

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments Special Issue on symposium on Solid and Physical Modeling (SPM'2018)

Journal ref Journal of Computer-Aided Design, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.03116 2018-02-12 cs.CL 57%

Zero-Resource Neural Machine Translation with Multi-Agent Communication Game

Yun Chen, Yang Liu, Victor O. K. Li

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments Published at AAAI-18

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.05376 2017-02-20 cs.AI cs.DM stat.ML 57%

Towards a Unified Taxonomy of Biclustering Methods

Dmitry I. Ignatov, Bruce W. Watson

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments http://ceur-ws.org/Vol-1552/

Journal ref Russian and South African Workshop on Knowledge Discovery Techniques Based on Formal Concept Analysis (RuZA 2015), November 30 - December 5, 2015, Stellenbosch, South Africa, In CEUR Workshop Proceedings, Vol. 1552, p. 23-39

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.01472 2016-09-07 cs.CY cs.AI 57%

OpenTripPlanner, OpenStreetMap, General Transit Feed Specification: Tools for Disaster Relief and Recovery

Chelcie Narboneta, Kardi Teknomo

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 6 pages, Narboneta, C. G. and Teknomo, K. (2014) OpenTripPlanner, OpenStreetMap, General Transit Feed Specification: Tools for Disaster Relief and Recovery, Proceeding of the 7th IEEE International Conference Humanoid, Nanotechnology, Information Technology Communication and Control, Environment and Management (HNICEM) 12-16 November 2014 Hotel Centro, Puerto Princesa, Palawan, Philippines

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.03845 2016-08-15 cs.RO cs.AI 57%

Traversing Environments Using Possibility Graphs for Humanoid Robots

Michael X. Grey, Aaron D. Ames, C. Karen Liu

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments Submitted to the International Workshop on the Algorithmic Foundations of Robotics (2016)

详情

展开后加载摘要…

URL PDF HTML 收藏
1405.5164 2014-05-21 cs.CV 57%

Multi-ellipses detection on images inspired by collective animal behavior

Erik Cuevas, Maurici Gonzalez, Daniel Zaldivar, Marco Perez

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments 21 pages

Journal ref Neural Computing and Applications, 24(5), (2014), 1019-1033

详情

展开后加载摘要…

URL PDF HTML 收藏
1301.0216 2013-01-03 cs.AI 57%

Applying Strategic Multiagent Planning to Real-World Travel Sharing Problems

Jan Hrnčíř, Michael Rovatsos

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 7th International Workshop on Agents in Traffic and Transportation, AAMAS, 2012

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13357 2026-07-24 cs.HC 版本更新 56%

TANDE: Disentangling Verbal and Nonverbal Backchannels in Emotional AI-Avatar Conversations with Young Adults

TANDE:在与年轻人的情感人工智能-虚拟化身对话中区分言语和非言语反馈渠道

Ann-Kareen Gedeus, Jack Good, Nadine Wagener, Angelique Taylor

专题命中 多模态Agent :multimodal(abstract,comments)

AI总结 研究在与年轻人的情感人工智能-虚拟化身对话中反馈渠道模式的影响,引入TANDE这个由LLM驱动的ECA,通过实验探讨其对融洽关系、同理心和参与度的作用及性别差异,得出相关设计启示。

Comments This paper has been accepted for publication at the 28th ACM International Conference on Multimodal Interaction (ICMI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27158 2026-08-28 cs.LG 新提交 50%

Diffusion Policies for Short-Horizon Planning in Robot Crowd Navigation

用于机器人人群导航短程规划的扩散策略

Wendong Li, Jochen Garcke

机构 * Institute for Numerical Simulation(数值模拟研究所) University of Bonn(波恩大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 针对机器人人群导航中现有方法难以表征多样化短期避障策略的问题,提出PDPO框架,生成短程动作块并引入边界约束,在基准测试中取得更好导航成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24749 2026-08-26 cs.SE 新提交 50%

From Natural Language Requirements to Graphical User Interfaces: Automated Prototyping and Verification with Pretrained Language Models

从自然语言需求到图形用户界面:基于预训练语言模型的自动原型设计与验证

Kristian Kolthoff

专题命中 多模态Agent :multimodal(abstract)

AI总结 本研究针对将自然语言需求转化为GUI原型及GUI应用需求验证的两大挑战,提出基于预训练语言模型的自动化方法,可显著减少相关手动工作量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22702 2026-08-25 cs.HC 新提交 50%

AffAdapt: AFFect-driven ADAPTive AI Personas for Seamless Conversations

AffAdapt:面向流畅对话的情感驱动型自适应AI角色

Nishanth Chidambaram, Kaustubh Paliwal, Kayla Hom, Shaoze Zhou, Chen Chen, Manas Satish Bedmutha, Nadir Weibel

专题命中 多模态Agent :multimodal(abstract)

AI总结 研究提出AffAdapt框架,整合多模块构建AI角色交互循环,实现流畅人机对话,在高风险对话场景验证其有效性,指出相关挑战并说明其应用场景。

Comments 3 pages, 2 figures, Adjunct Proceedings of the 39th Annual ACM Symposium on User Interface Software and Technology (UIST '26 Adjunct), Detroit, MI, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06607 2026-08-24 cs.CR 版本更新 50%

Your Harness is Not Secure: Benchmarking Real-world Threat of Command Line Interface Agent

你的工具不安全:评测命令行界面智能体的现实威胁

Weidi Luo, Qiming Zhang, Tianyu Lu, Xiaogeng Liu, Bin Hu, Hung-Chun Chiu, Siyuan Ma, Yizhe Zhang, Xusheng Xiao, Yinzhi Cao, Zhen Xiang, Chaowei Xiao

专题命中 多模态Agent :multimodal(abstract)

AI总结 该研究针对CLI智能体滥用风险,推出与MITRE ATT&CK对齐的AdvCLI基准测试,评估7款CLI智能体后发现其常越过拒绝边界,可完成部分恶意OS级任务,为开发安全机制提供测试平台。

Comments Accepted by EMNLP 2026 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19964 2026-08-21 cs.LG 新提交 50%

G-MARK: Grounded Multi-Agent Reasoning for Cooperative Driving via Knowledge Graphs

G-MARK:基于知识图谱的协同驾驶接地多智能体推理

Bhavya Gupta, Onat Gungor, Tajana Rosing

机构 * University of California, San Diego(加利福尼亚大学圣迭戈分校) West Virginia University(西弗吉尼亚大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 提出 G-MARK 框架,通过知识图谱实现协同驾驶多智能体推理,提升遮挡推理与控制选择性能,减小通信负载,效果优于现有基线。

Comments Accepted for oral presentation at the 25th IEEE International Conference on Machine Learning and Applications (ICMLA'26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18450 2026-08-20 cs.LG 新提交 50%

Adaptive Multi-Agent Feature Selection for Personalized Fall Risk Prevention

面向个性化跌倒风险预防的自适应多智能体特征选择

Chang Liu, Ladda Thiamwong, Yanjie Fu, Rui Xie

机构 * University of Central Florida(中佛罗里达大学) Arizona State University(亚利桑那州立大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 针对老年人跌倒风险识别的静态方法无法适配动态个性化风险因素,本文提出PAFIR框架,将自适应特征选择建模为强化学习问题,在PEER试验数据上验证其能更有效捕捉特征模式,实现动态个性化跌倒风险预防。

Comments 38 pages, 10 figures, 12 tables. Accepted at Machine Learning for Healthcare (MLHC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16843 2026-08-18 cs.RO 新提交 50%

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

基于基础模型的具身智能体的安全性:攻击面、攻击、防御与评估

Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao

机构 * Wuhan University(武汉大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 该研究以信任边界为核心,针对基于基础模型的具身智能体安全,划分了五个层级与十二个攻击面,分析了58种攻击、61种防御的现状,指出部分领域研究不足并提出开放挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15198 2026-08-18 stat.ML cs.LG physics.comp-ph 新提交 50%

Identifying parameter couplings and uncertainties of mixed-noise stochastic systems via full-covariance Gaussian mixture network

通过全协方差高斯混合网络识别混合噪声随机系统的参数耦合与不确定性

Xiaolong Wang, Xiangwen Hao, Jing Feng, Yuanyuan Liu, Yong Xu

机构 * School of Science, Xi’an University of Posts and Telecommunications(西安邮电大学理学院)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 研究针对混合噪声随机系统参数识别的难点,提出PENN-GMD神经网络,采用全协方差高斯混合分布,经五个数值示例验证可准确恢复似然分布、捕捉参数耦合,为复杂随机系统参数识别提供实用工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18518 2026-08-18 cs.LG stat.ME stat.ML 版本更新 50%

Measuring the Prevalence of Policy Violating Content with ML Assisted Sampling and LLM Labeling

利用机器学习辅助抽样和大语言模型标注测量违规内容的普及率

Attila Dobi, Aravindh Manickavasagam, Benjamin Thompson, Xiaohan Yang, Faisal Farooq

机构 * Pinterest

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种基于机器学习和大语言模型的系统,用于高效测量违反政策内容的普及率,通过概率抽样和多模态标注提升测量准确性。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22983 2026-08-18 cs.RO 版本更新 50%

Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives

基础模型时代中的具身机器人操作:规划与学习视角

Shuanghao Bai, Wenxuan Song, Jiayi Chen, Yuheng Ji, Zhide Zhong, Jin Yang, Han Zhao, Wanqi Zhou, Zhe Li, Pengxiang Ding, Cheng Chi, Chang Xu, Xiaolong Zheng, Donglin Wang, Haoang Li, Shanghang Zhang, Badong Chen

机构 * Xi’an Jiaotong Univeristy(西安交通大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Chinese Academy of Sciences(中国科学院) Westlake University(西湖大学) Zhejiang University(浙江大学) University of Sydney(悉尼大学) BAAI(百度人工智能研究院) Peking University(北京大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文探讨了基础模型时代机器人操作的规划与学习方法,分析了高层推理与低层控制的统一框架,并提出了未来研究方向。

Comments This work is a re-architected core derived from the full survey (arXiv:2510.10903), refined to highlight the most central themes and representative studies

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10903 2026-08-18 cs.RO 版本更新 50%

Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey

面向机器人操作的统一理解:一项综合调查

Shuanghao Bai, Wenxuan Song, Jiayi Chen, Yuheng Ji, Zhide Zhong, Jin Yang, Han Zhao, Wanqi Zhou, Wei Zhao, Zhe Li, Pengxiang Ding, Cheng Chi, Haoang Li, Chang Xu, Xiaolong Zheng, Donglin Wang, Shanghang Zhang, Badong Chen

专题命中 多模态Agent :multimodal(abstract)

AI总结 本调查针对机器人操作这一具身智能核心挑战,提出方法的统一分类与瓶颈分类,为机器人操作研究提供了全面的路线图与结构化参考。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14131 2026-08-17 cs.SE 新提交 50%

LegacyWorld: Atomicity-Aware Evaluation of GUI Agents for Legacy Workflows

LegacyWorld:面向遗留工作流的GUI智能体的原子性感知评估

Thilo Reintjes, Sivajeet Chand, Derui Zhu, Sushant Kumar Pandey, Alexander Pretschner

专题命中 多模态Agent :multimodal(abstract)

AI总结 该研究开发LegacyUse框架以自动化遗留工作流,构建含28个Windows GUI工作流的基准,采用原子性评估6款计算机使用智能体,发现不同操作特征并提出相关首要要求。

Comments Accepted for publication in the Industry Track of the 42nd IEEE International Conference on Software Maintenance and Evolution (ICSME 2026), 14-18 September 2026, Benevento, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12763 2026-08-14 cs.CE 新提交 50%

ARIES-Mission2: A Zero-Shot Vision-Language-Action Framework for Fast Large-Scale Aerial Mission Generation

ARIES-Mission2:用于快速大规模空中任务生成的零样本视觉-语言-动作框架

Junhao Wei, Yanxiao Li, Haochen Li, Yifu Zhao, Dexing Yao, Baili Lu, Zikun Li, Yapeng Wang, Sio-Kei Im, Dingcheng Yang, Xu Yang

专题命中 多模态Agent :multimodal(abstract)

AI总结 ARIES-Mission2是解耦视觉语义感知与路径优化的零样本VLA框架,在UAV基准上,其飞行距离更短、任务生成速度更快,且TSP模块可扩展性佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10252 2026-08-12 stat.ME stat.AP 新提交 50%

Joint return levels of maximum temperature and minimum relative humidity by combining copulas with an extreme value framework for bimodal data

结合Copula与适用于双峰数据的极值框架,分析最高气温与最小相对湿度的联合重现水平

Beatriz G da Cruz Albernaz, Cira E G Otiniano, Carolyne Soares de Brito, Fidel E C Morales, Enzo Porto Brasil

专题命中 多模态Agent :multimodal(abstract)

AI总结 本研究结合Copula与适用于双峰数据的极值框架,分析巴西利亚最高气温与最小相对湿度的联合重现水平,为当地气候风险应对提供关键依据。

Comments 19 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏