On-device Large Multi-modal Agent for Human Activity Recognition
用于人体活动识别的设备端大多模态智能体
机构 * The Ohio State University(俄亥俄州立大学)
专题命中 多模态Agent :multi-modal(title,abstract)
AI总结 本文提出了一种用于人体活动识别的设备端大多模态智能体,结合大语言模型提升性能与可解释性,实现高分类准确率和用户友好交互。
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
用于人体活动识别的设备端大多模态智能体
机构 * The Ohio State University(俄亥俄州立大学)
专题命中 多模态Agent :multi-modal(title,abstract)
AI总结 本文提出了一种用于人体活动识别的设备端大多模态智能体,结合大语言模型提升性能与可解释性,实现高分类准确率和用户友好交互。
多模态移动形变机器人(M4)的可 traversability 自主导航
机构 * SiliconSynapse Lab(硅合成实验室)
专题命中 多模态Agent :multi-modal(title,abstract)
AI总结 本研究提出了一种基于LiDAR的多模态移动形变机器人M4的可 traversability 自主导航框架,通过学习地形分析生成节能路径,提升地形适应能力。
Comments Master's thesis
利用多智能体分歧进行多模态推理中的工具招募
机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) ; Nanyang Technological University(南洋理工大学) ; The University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 多模态Agent :multimodal(title);分类 cs.CV、cs.CL、cs.AI
AI总结 DART通过多智能体分歧识别有用视觉工具,提升多模态推理中的工具调用效果。
Comments Code: https://github.com/nsivaku/dart
ResponsibleRobotBench: 使用多模态大语言模型评估负责任的机器人操控
机构 * University of Hamburg(汉堡大学) ; Agile Robots SE(敏捷机器人公司) ; Technical University of Munich(慕尼黑技术大学) ; Hong Kong Polytechnic University(香港理工大学)
专题命中 多模态Agent :multi-modal(title);multimodal(abstract)
AI总结 ResponsibleRobotBench通过多模态大语言模型评估机器人操控的责任性,涵盖23个多阶段任务,强调安全性、泛化能力和物理可靠性。
Comments https://sites.google.com/view/responsible-robotbench
CogDrive: 基于认知的多模态预测-规划融合用于安全自主性
机构 * Singapore-MIT Alliance for Research and Technology (SMART), Singapore(新加坡-麻省理工联合研究技术联盟) ; Department of Urban Studies and Planning, Massachusetts Institute of Technology, USA(麻省理工学院城市研究与规划系) ; School of Vehicle and Mobility, Tsinghua University, China(清华大学车辆与移动系统学院) ; Department of Mechanical Engineering, National University of Singapore, Singapore(新加坡国立大学机械工程系) ; Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University, China(同济大学交通工程重点实验室)
专题命中 多模态Agent :multimodal(title,abstract)
AI总结 CogDrive通过结合认知多模态预测与安全导向规划,实现了在复杂交通中的安全自主性,提升了轨迹预测和适应性行为。
Comments 25 pages, 6 figures
专题命中 多模态Agent :multimodal(title,abstract)
Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript
专题命中 多模态Agent :multimodal(title,abstract)
Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript
专题命中 多模态Agent :multimodal(title,abstract)
专题命中 多模态Agent :multimodal(title,abstract)
Journal ref In Proceedings of The 19th IEEE International Conference on Bioinformatics and Biomedicine (BIBM 2025)
机构 * Department of Artificial Intelligence, Fraunhofer Heinrich Hertz Institute(人工智能部门,弗劳恩霍夫海因里希·赫兹研究所) ; Department of Nuclear Medicine, Charité—Universitätsmedizin Berlin(核医学部门,柏林查理医院) ; Department of Hepatology and Gastroenterology, Charité—Universitätsmedizin Berlin(肝病与胃肠病学部门,柏林查理医院) ; Division of Interventional Radiology, Department of Radiology, Memorial Sloan Kettering Cancer Center(介入放射学部门,纪念斯隆-凯特琳癌症中心) ; Department of Endocrinology and Metabolism, Charité—Universitätsmedizin Berlin(内分泌与代谢学部门,柏林查理医院) ; Clinic for Gastroenterology, Hepatology and Infectious Diseases, University Hospital Düsseldorf, Medical Faculty of Heinrich Heine University Düsseldorf(胃肠病、肝病和传染病诊所,杜塞尔多夫大学医院,海因里希·海涅大学医学部) ; Department of Electrical Engineering and Computer Science, Technische Universität Berlin(电气工程与计算机科学部门,柏林技术大学) ; Berlin Institute of Health at Charité – Universitätsmedizin Berlin(柏林查理医院健康研究所)
专题命中 多模态Agent :multimodal(title,abstract)
机构 * Georgetown University(乔治城大学)
专题命中 多模态Agent :multimodal(title,abstract)
Comments Accepted for 2025 ACM SIGSPATIAL conference
专题命中 多模态Agent :multi-modal(title,abstract)
Comments Abstracted submitted in the Proceedings of the IISE Annual Conference & Expo 2025
Journal ref Proceedings of the IISE Annual Conference & Expo 2025
专题命中 多模态Agent :multi-modal(title,abstract)
机构 * University of Glasgow(格拉斯哥大学) ; University of Glasgow School of Computer Science(格拉斯哥大学计算机科学学院)
专题命中 多模态Agent :multimodal(title,abstract)
Comments Accepted by ACM Multimedia 2025, camera-ready version
专题命中 多模态Agent :multimodal(title,abstract)
机构 * Johns Hopkins University(约翰霍普金斯大学)
专题命中 多模态Agent :multimodal(title,abstract)
Comments 8 Pages, 7 Figures
机构 * AI Division, School of Engineering, Westlake University(西拉丘学院人工智能系,西湖大学) ; Zhejiang University(浙江大学) ; Oxford University(牛津大学)
专题命中 多模态Agent :multimodal(title,abstract)
Comments 9 pages, 5 figures. This paper was withdrawn from the IJCAI 2025 proceedings due to the lack of participation in the conference and presentation
专题命中 多模态Agent :multi-modal(title,abstract)
Comments Change the method and experimentation
机构 * Department of Encephalopathy, Chengdu Pidu District Hospital of Traditional Chinese Medicine(脑病科,成都_pidu区中医医院) ; Department of Tuina, Chengdu Pidu District Hospital of Traditional Chinese Medicine(推拿科,成都_pidu区中医医院) ; Wuhan Hospital of Integrated Traditional Chinese and Western Medicine, Affiliated to Hubei University of Chinese Medicine(中西医结合医院,湖北中医药大学附属医院) ; Department of Electrical and Computer Engineering, Carnegie Mellon University(电气与计算机工程系,卡内基梅隆大学)
专题命中 多模态Agent :multi-modal(title,abstract)
机构 * Delft University of Technology(代尔夫特理工大学)
专题命中 多模态Agent :multi-modal(title,abstract)
Comments Accepted as an oral presentation at the 29th IAVSD. August 18-22, 2025. Shanghai, China
专题命中 多模态Agent :multimodal(title,abstract)
Comments 18 pages, 4 figures
机构 * Zhixuan Shen, Haonan Luo, Kexun Chen, Fengmao Lv, Tianrui Li(作者)
专题命中 多模态Agent :multimodal(title,abstract)
Comments 16 pages, 10 figures, Extended Version of accepted AAAI 2025 Paper
专题命中 多模态Agent :multimodal(title,abstract)
Comments 11 pages, 10 figures, Accepted to IEEE ISMAR 2025 (TVCG)
机构 * National University of Singapore (NUS)(新加坡国立大学) ; The Cambridge Centre for Advanced Research and Education in Singapore (CARES)(新加坡剑桥高级研究与教育中心) ; China Southwest Architectural Design and Research Institute Co., Ltd. (CSWADI)(中国西南建筑规划设计研究院有限公司)
专题命中 多模态Agent :multi-modal(title);multimodal(abstract)
Comments 42 pages, 32 figures, submitted to Environment and Planning B: Urban Analytics and City Science
专题命中 多模态Agent :multimodal(title,abstract)
Comments Accepted to the IEEE Visualization Conference (VIS'25). 11 pages, 6 figures
机构 * Aalto University(阿alto大学) ; TU Wien(维也纳技术大学)
专题命中 多模态Agent :multimodal(title,abstract)
Comments 9 pages
Journal ref The Future of Human-Robot Synergy in Interactive Environments: The Role of Robots at the Workplace @ CHIWork 2025
机构 * College of Informatics, Harbin Institute of Technology(信息学院,哈尔滨工业大学) ; Key Laboratory of Forest and Grassland Fire Risk Prevention, Ministry of Emergency Management, China Fire and Rescue Institute(森林和草原火灾风险预防重点实验室,应急管理部,中国消防救援学院) ; Southern University of Science and Technology(南方科技大学)
专题命中 多模态Agent :multimodal(title,abstract)
Comments 8 pages, 5 figures,submitted to IEEE wcm
机构 * Chalmers University of Technology(查尔姆斯理工大学)
专题命中 多模态Agent :multimodal(title,abstract)
Comments Published in IEEE Robotics and Automation Letters (RA-L)
专题命中 多模态Agent :multimodal(title,abstract)
Comments 11 Pages
机构 * Chang’an University(长安大学) ; Agency for Science, Technology and Research (A*STAR)(科技研究局) ; University of California, Davis(加州大学戴维斯分校) ; CCNY Robotics Lab, The City College of New York(纽约城市学院机器人实验室)
专题命中 多模态Agent :multimodal(title,abstract)