Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
将单体基础模型转化为具身体验的多智能体架构以实现人机协作
Nan Sun, Bo Mao, Yongchang Li, Chenxu Wang, Di Guo, Huaping Liu
机构
*
Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学)
;
School of Artificial Intelligence, Beijing University of Posts and Telecommunications(人工智能学院,北京邮电大学)
CommentsPublished at AAAI/ACM AIES 2025. Presented at NeurIPS 2025 Workshop Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling
Journal refProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8(1), 2025, 343-354
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
FOM-Nav:用于目标导航的前沿-对象地图
Thomas Chabal, Shizhe Chen, Jean Ponce, Cordelia Schmid
机构
*
Inria(法国国家信息与自动化技术研究院)
;
Department of Computer Science, École normale supérieure (ENS-PSL, CNRS, Inria)(高等师范学院计算机科学系)
;
Courant Institute of Mathematical Sciences and Center for Data Science, New York University(纽约大学Courant数学科学研究所和数据科学中心)
机构
*
Department of computer science, University of Central Florida, Orlando, USA(计算机科学系,中央佛罗里达大学)
;
Department of Bioengineering, University of Pennsylvania, Philadelphia, PA, USA(生物工程系,宾夕法尼亚大学)
;
Department of electrical engineering, Columbia university, New York, NY, USA(电气工程系,哥伦比亚大学)
;
Technical University of Applied Sciences Regensburg, Regensburg, Germany(应用科学技术大学(雷根斯堡))
;
Department of Surgery, University of Calgary, Calgary, Alberta, Canada(外科系,卡尔加里大学)
;
University College of Nabi Akram, Tabriz, Iran(纳比阿克兰大学)
;
School of Electrical Engineering, Iran University of Science and Technology, Tehran, Iran(电气工程学院,伊朗科学技术大学)
专题命中
幻觉与鲁棒性
:vision language model(title);vision-language model(abstract);visual question answering(abstract);分类 cs.CV
Scaling Down to Scale Up: Towards Operationally-Efficient and Deployable Clinical Models via Cross-Modal Low-Rank Adaptation for Medical Vision-Language Models
缩小规模以扩大规模:通过跨模态低秩适应实现操作高效且可部署的临床模型
Thuraya Alzubaidi, Farhad R. Nezami, Muzammil Behzad
机构
*
King Fahd University of Petroleum(国王法赫德石油与矿物大学)
;
Institute for Medical Engineering(医学工程研究所)
;
Science, Massachusetts Institute of Technology, US(科学,麻省理工学院,美国)
;
Harvard Medical School, Harvard University, US(哈佛医学院,哈佛大学,美国)
;
SDAIA-KFUPM Joint Research Center for Artificial Intelligence, Saudi Arabia(SDAIA-KFUPM人工智能联合研究中心,沙特阿拉伯)
Closing the Gap: Data-Centric Fine-Tuning of Vision Language Models for the Standardized Exam Questions
弥合差距:面向标准化考试题目的视觉语言模型数据驱动微调
Egemen Sert, Şeyda Ertekin
机构
*
organization= Department of Computer Engineering, Middle East Technical University (METU) , city= Ankara , country= Türkiye
;
organization= METU-DTX Digital Transformation \& Innovation Centre, METU , city= Ankara , country= Türkiye
专题命中
VLM训练与架构
:vision language model(title,abstract);分类 cs.CV、cs.AI
机构
*
Morgan Stanley(摩根士丹利)
;
Clemson University(克莱姆森大学)
;
Arizona State University(亚利桑那州立大学)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
University of Arizona(亚利桑那大学)
;
University of Notre Dame(圣母大学)
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
面向渲染的强化学习用于矢量图形生成
Juan A. Rodriguez, Haotian Zhang, Abhay Puri, Aarash Feizi, Rishav Pramanik, Pascal Wichmann, Arnab Mondal, Mohammad Reza Samsami, Rabiul Awal, Perouz Taslakian, Spandana Gella, Sai Rajeswar, David Vazquez, Christopher Pal, Marco Pedersoli
机构
*
ServiceNow Research(ServiceNow研究机构)
;
Mila
;
ÉTS Montréal(蒙特利尔ÉTS)
;
Polytechnique Montréal(蒙特利尔Polytechnique)
;
Columbia University(哥伦比亚大学)
;
Stony Brook University(石溪大学)
;
Apple(苹果公司)
;
Google Research(谷歌研究)
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
;
McGill University(麦吉尔大学)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Nanyang Technological University(南洋理工大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Sichuan University(四川大学)
;
National University of Singapore(新加坡国立大学)
;
Shenzhen University(深圳大学)