Online,Target-Free LiDAR-Camera Extrinsic Calibration via Cross-Modal Mask Matching
专题命中 多模态Agent :cross-modal(title,abstract);分类 cs.CV、cs.AI
Comments accepted to IEEE Trans. on Intelligent Vehicles (T-IV)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态Agent :cross-modal(title,abstract);分类 cs.CV、cs.AI
Comments accepted to IEEE Trans. on Intelligent Vehicles (T-IV)
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Project released at: https://github.com/ChenYi99/EgoPlan
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to ACL 2024 (main). Code and data is released at https://github.com/MinorJerry/WebVoyager
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments Findings of ACL 2024
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments Accepted by ACL 2024
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments Project website: https://github.com/cocacola-lab/MineLand
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 15 pages, 7 figures
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments 10 pages, 3 figures, 2 tables
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Accept by Lrec-Coling 2024
专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CL、cs.AI
Comments 14 pages, 10 figures
专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CV、cs.CL
Comments The datasets were incomplete as they did not include all the necessary copyrights
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Accepted to ICASSP 2024
专题命中 多模态Agent :multimodal(title);cross-modal(abstract);分类 cs.CL、cs.AI
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Work in progress
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments EMNLP 2023
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Accepted to NeurIPS 2023. Project webpage: https://sites.google.com/view/2023arp
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments NILLI at EMNLP 2022
专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.CL、cs.MM
Comments This manuscript has been accepted at ACM MM 2020
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments See https://devendrachaplot.github.io/projects/EMML for demo videos
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI
Journal ref Proceedings of the 2018 EMNLP Workshop SCAI: The 2nd International Workshop on Search-Oriented Conversational AI, pages 59-66, Brussels, Belgium, October 2018
机构 * Player2
专题命中 多模态Agent :multimodal(title);multi-modal(abstract,comments);分类 cs.AI
Comments International Conference on Computer Vision Workshop on Multi-Modal Reasoning for Agentic Intelligence
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV
Comments ACM Intl. Conf. on Multimodal Interaction (ICMI). arXiv admin note: substantial text overlap with arXiv:2004.07173
VDC-Agent:当视频详细描述器通过代理自我反思而自我进化
机构 * Xi’an Jiaotong University(西安交通大学) ; Kuaishou Technology(快手科技) ; Shenzhen University of Advanced Technology(深圳先进技术大学)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 VDC-Agent通过自我反思机制实现视频详细描述的自我进化,利用自动生成的(描述,评分)对提升描述准确性与评分表现。
Comments Accepted to ECCV 2026. Project Page: https://vdcagent.github.io
学习选择视觉上下文示例
机构 * University of Cincinnati(辛辛那提大学) ; University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
专题命中 多模态Agent :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本文提出LSD方法,通过强化学习构建最优演示集,提升多模态大语言模型在视觉回归任务中的表现,揭示了学习选择在视觉上下文学习中的必要性。
Comments 21 pages, 12 figure, accepted to Computer Vision and Pattern Recognition Conference (CVPR) 2026 Findings Track
Journal ref In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9455-9465) 2026
可验证几何问题求解:求解器驱动的自动形式化与定理提出
机构 * Beijing Normal University(北京师范大学)
专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 提出SD-GPS框架,通过求解器驱动的自动形式化和可验证定理提出,解决几何问题求解中神经符号方法的瓶颈,在Geometry3K和PGPS9K上超越现有方法。
HATS:用于多臂数据收集的人-智能体遥操作系统
机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) ; Nanyang Technological University(南洋理工大学) ; Imperial College London(帝国理工学院)
专题命中 多模态Agent :MLLM(summary_cn,abstract)
AI总结 提出HATS系统,由单操作员借助MLLM智能体控制两主臂和两辅助臂,实现高效多臂数据收集,性能媲美双人专家团队。
面向多模态大语言模型的符合规则的视觉空间规划
机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所) ; Yinwang Intelligent Technology Co., Ltd(银湾智能科技有限公司)
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI
AI总结 该研究针对多模态大语言模型的规则遵循空间规划问题,构建了RuleMaze基准,提出语言-逻辑-函数混合方法和解耦多模态规划(DMP),提升了规则遵循度与规划成功率。