3D BAT: A Semi-Automatic, Web-based 3D Annotation Toolbox for Full-Surround, Multi-Modal Data Streams
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态Agent :multi-modal(title);分类 cs.CV
专题命中 多模态Agent :multimodal(title);分类 cs.CV
专题命中 多模态Agent :multimodal(title);分类 cs.CV
具身形态塑造多模态婴儿模型中的翻滚行为
机构 * Frankfurt Institute for Advanced Studies(法兰克福高等研究院) ; Goethe University Frankfurt(法兰克福大学) ; University of New South Wales(新南威尔士大学)
专题命中 多模态Agent :multimodal(title,comments)
AI总结 通过虚拟婴儿MIMo学习仰卧到俯卧翻滚,研究婴儿运动发展中的具身形态变化如何影响行为,发现与真实婴儿一致的发育趋势和协调模式。
Comments 7 pages, 7 figures. Accepted at the 2026 IEEE ICDL Conference. Cite as: L. Philipp, F. M. López, and J. Triesch, "Embodiment Shapes Rolling Behavior in a Multimodal Infant Model", in 2026 IEEE International Conference on Development and Learning (ICDL). IEEE, 2026, pp. 1-7
前沿大语言模型在空间意象推理中的局限性
机构 * Institute of Mathematics and Statistics – University of São Paulo(数学统计研究所 – 圣保罗大学)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI
AI总结 本研究通过引入外部“意象模块”辅助3D模型旋转任务,发现即使外包整体3D状态维护,前沿模型仍缺乏基础视觉空间原语,导致准确率最高仅62.5%。
Comments 25 pages. v2: Title updated; added a section on object/spatial imagery and propositional reasoning; added new experimental results for the single-object rotation probe
快门之前:3D场景中美学的且可执行的人像摄影规划
机构 * The Hong Kong Polytechnic University(香港理工大学)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI
AI总结 提出在3D场景中生成人像姿态、相机、照明和曝光方案的方法,通过构建摄影场景图实现美学引导的规划,生成视觉上引人注目且几何与光度可行的人像。
具身人工智能的安全性:风险、攻击与防御综述
机构 * Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院) ; City University of Hong Kong(香港城市大学) ; Jilin University(吉林大学) ; Singapore Management University(新加坡管理大学) ; Deakin University(德肯大学) ; Tongji University(同济大学) ; Nanyang Technological University(南洋理工大学) ; Chinese Academy of Sciences(中国科学院) ; The University of Melbourne(墨尔本大学) ; Johns Hopkins University(约翰霍普金斯大学)
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI
AI总结 本文综述了具身AI在感知、认知、规划、行动及交互全流程中的安全风险、攻击与防御方法,提出了多层次分类体系,并指出了多模态感知融合脆弱性、规划不稳定及人机交互可信度等关键挑战。
Comments Survey paper; 75 pages, 4 figures, 18 tables; v2 expands embodied-specific coverage of agentic threats, World Action Model threats, and contextual risk mitigation, with over 100 new references added. Project page: https://x-zheng16.github.io/Awesome-Embodied-AI-Safety/
StarVLA:一种积木式代码库,用于视觉-语言-动作模型开发
机构 * Von Neumann Institute, HKUST(香港科技大学冯·诺依曼研究所)
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI
AI总结 StarVLA通过模块化架构、可重用训练策略和统一评估接口,解决VLA方法碎片化问题,提升可复现性和跨架构兼容性。
Comments Open-source VLA infra, Technical Report
HippoCamp:在个人电脑上对上下文代理进行基准测试
机构 * S-Lab, Nanyang Technological University, Singapore(新加坡南洋理工大学S-Lab)
专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 HippoCamp是一个新的基准,用于评估代理在多模态文件管理中的能力,通过用户为中心的环境建模个体用户档案并搜索大规模个人文件进行上下文感知推理,揭示了当前代理在真实环境中的局限性。
Comments Project Page: https://hippocamp-ai.github.io/
FAPE-IR:面向全场景图像修复的频率感知规划与执行框架
机构 * Tianjin University(天津大学) ; University of Macau(澳门大学) ; City University of Hong Kong(香港城市大学) ; Institute of Artificial Intelligence (TeleAI)(人工智能研究所(TeleAI))
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 本文提出FAPE-IR框架,通过频率感知规划与执行模块,结合多模态大语言模型和LoRA-MoE架构,实现统一且可解释的全场景图像修复,实验显示其在七项任务中表现优异。
多智能体系统实现从化学文献中灵活的信息提取
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 本研究提出了一种基于多模态大语言模型的多智能体系统,实现了从化学文献中高效提取化学信息,F1分数达76.27%,显著提升信息提取效率。
ACE-Brain-0:空间智能作为通用具身化体系的共享框架
机构 * Shanghai Jiao Tong University(上海交通大学) ; Nanyang Technological University(南洋理工大学) ; The Chinese University of Hong Kong(香港中文大学) ; The University of Hong Kong(香港大学) ; University of Science(科学技术大学) ; Fudan University(复旦大学) ; Xiamen University(厦门大学) ; East China Normal University(华东师范大学) ; Wuhan University(武汉大学) ; Sun Yat-sen University(中山大学)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL
AI总结 ACE-Brain-0 通过空间智能作为共享框架,统一了自动驾驶、机器人和 UAVs 的具身化任务,采用 SSR 范式和 GRPO 方法实现跨领域泛化和领域精通的平衡。
Comments Code: https://github.com/ACE-BRAIN-Team/ACE-Brain-0 Hugging Face: https://huggingface.co/ACE-Brain/ACE-Brain-0-8B
MADIAVE:多智能体辩论用于隐式属性值提取
机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI
AI总结 MADIAVE通过多智能体辩论框架提升隐式属性值提取的准确性和鲁棒性,适用于多模态电子商务场景。
Comments Accepted by EACL 2026 (Findings)
MEDVISTAGYM:一种通过工具集成强化学习进行医学图像思考的可扩展训练环境
机构 * Virginia Tech(弗吉尼亚理工大学) ; UT Southwestern Medical Center(德克萨斯大学西南医学中心) ; Georgia Institute of Technology(佐治亚理工学院) ; Cisco(思科公司)
专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI
AI总结 MedVistaGym通过工具集成强化学习提升医学图像分析的代理训练效果。
VULCAN:工具增强的多智能体用于迭代3D物体排列
机构 * Stanford University(斯坦福大学) ; Google(谷歌) ; New York University(纽约大学)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 VULCAN通过引入MCP API、视觉工具和多智能体框架,提升了3D物体排列任务中MLLMs的视觉接地能力与迭代处理效率。
机构 * University of Milan-Bicocca(米兰-比科卡大学) ; University of Naples Federico II(那不勒斯费德里科二世大学) ; Oversonic Robotics(Oversonic机器人公司) ; University of Essex(埃塞克斯大学) ; TUM School of Social Sciences and Technology(慕尼黑技术大学社会科学与技术学院)
专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL、cs.AI
Comments Accepted at IAS19
机构 * Shanghai Jiao Tong University(上海交通大学)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI
Comments NAACL 2025 Demo Track [code] https://github.com/OpenDFM/MobA [dataset] https://huggingface.co/datasets/OpenDFM/MobA-MobBench
机构 * School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI、cs.MM
专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI
机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) ; Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心) ; Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究院)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI
Comments Accepted by IEEE CASM
机构 * Information Processing Lab, University of Washington, USA(华盛顿大学信息处理实验室) ; Electronics and Telecommunications Research Institute, South Korea(韩国电子电信研究院) ; National Center for High-performance Computing, Taiwan(台湾高性能计算国家研究中心)
专题命中 多模态Agent :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments 1st Place Solution of the 9th AI City Challenge Track 3
机构 * Tsinghua University(清华大学) ; University of Science and Technology of China(中国科学技术大学) ; Hefei University of Technology(合肥工业大学)
专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);分类 cs.AI、cs.MM
专题命中 多模态Agent :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments Paper accepted to CVPR 2025
专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);分类 cs.CL、cs.AI
专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI
Comments 24 pages, 16 tables, 17 figures
专题命中 多模态Agent :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.CL
Comments 22 pages, 11 figures, 10 Tables
专题命中 多模态Agent :multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI
专题命中 多模态Agent :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments To appear at CVPR 2020. The first two authors contributed equally to this manuscript. Code: https://github.com/weituo12321/PREVALENT
检索前先复用:诊断具身多模态策略测试时增强的余量与互补性
机构 * KAIST(韩国科学技术院)
专题命中 多模态Agent :multimodal(title)
AI总结 该研究提出通过可恢复余量与检索互补性两个因素,为具身多模态冻结 VLA 策略的测试时增强选择采样或检索干预,在 LIBERO 上成功提升最高 21.0 个百分点,且可迁移至其他机器人与环境。
Comments Accepted to ECCV 2026 workshop