机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Lingnan University(岭南大学)
;
The Third Affiliated Hospital of Kunming Medical University, Yunnan Cancer Hospital, Peking University Cancer Hospital Yunnan(昆明医科大学第三附属医院、云南癌症医院、北京大学肿瘤医院云南分院)
;
Guangzhou First People's Hospital, South China University of Technology(广州第一人民医院、华南理工大学)
;
The Third Affiliated Hospital, Sun Yat-Sen University(中山大学第三附属医院)
;
HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(香港科技大学深圳-香港协同创新研究院)
机构
*
Department of Biomedical Engineering, Hefei University of Technology(合肥工业大学生物医学工程系)
;
Faculty of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院)
;
School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学技术学院)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
AffordGrasp:跨模态扩散用于感知意识抓取合成
Xiaofei Wu, Yi Zhang, Yumeng Liu, Yuexin Ma, Yujiao Shi, Xuming He
机构
*
ShanghaiTech University(上海科技大学)
;
Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程技术研究中心)
;
University of Science and Technology of China(中国科学技术大学)
Junli Wang, Yinan Zheng, Xueyi Liu, Zebin Xing, Pengfei Li, Guang Li, Kun Ma, Guang Chen, Hangjun Ye, Zhongpu Xia, Long Chen, Qichao Zhang
机构
*
SKL-MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Xiaomi EV(小米汽车)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学智能产业研究院)
A Structured Framework for Evaluating and Enhancing Interpretive Capabilities of Multimodal LLMs in Culturally Situated Tasks
一种评估和增强多模态大语言模型在文化情境任务中解释能力的结构框架
Haorui Yu, Ramon Ruiz-Dolz, Qiufeng Yi
机构
*
DJCAD, University of Dundee, United Kingdom(邓迪大学DJCAD部门)
;
ARG-tech, SSEN, University of Dundee, United Kingdom(邓迪大学)
;
School of Computer Science, University of Birmingham, United Kingdom(伯明翰大学计算机科学学院)
PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
PPGFlowECG: 基于跨模态编码的潜在修正流用于PPG引导的ECG生成和心血管疾病检测
Xiaocheng Fang, Jiarui Jin, Haoyu Wang, Che Liu, Jieyi Cai, Yujie Xiao, Guangkun Nie, Bo Liu, Shun Huang, Hongyan Li, Shenda Hong
机构
*
National Institute of Health Data Science, Peking University, China(北京大学国家健康数据科学研究院)
;
School of Intelligence Science and Technology, Peking University, China(北京大学智能科学与技术学院)
;
Data Science Institute, Imperial College London, UK(伦敦帝国理工学院数据科学研究院)
;
University of Chinese Academy of Sciences, China(中国科学院大学)
机构
*
Stanford University(斯坦福大学)
;
University of California, San Diego(加州大学圣地亚哥分校)
;
Wuhan University of Science and Technology(武汉科技大学)
;
Hunan University(湖南大学)
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
Zheng Wei, Mingchen Li, Zeqian Zhang, Ruibin Yuan, Pan Hui, Huamin Qu, James Evans, Maneesh Agrawala, Anyi Rao
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
University of Chicago(芝加哥大学)
;
Stanford University(斯坦福大学)
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
Ruibing Hou, Mingshuang Luo, Hongyu Pan, Hong Chang, Shiguang Shan
机构
*
Key Laboratory of Intelligent Information Processing, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS)(智能信息处理重点实验室,计算技术研究所(ICT),中国科学院(CAS))
;
University of the Chinese Academy of Sciences(中国科学院大学)
Rich Vehicle Routing Problem in Disaster Management enabling Temporally-causal Transhipments across Multi-Modal Transportation Network
Santanu Banerjee, Goutam Sen, Siddhartha Mukhopadhyay
机构
*
Department of Industrial and Systems Engineering (ISE), Indian Institute of Technology (IIT) Kharagpur(工业与系统工程系,印度理工学院Kharagpur分校)
专题命中
多模态生成
:multi-modal(title);分类 cs.AI
CommentsMajor changes in version II: 1) Supplementary is now a separate document, 2) Algorithm steps have been updated with pseudocode in the Heuristic, 3) Explanation of the MILP formulation construction is further detailed in a supplementary section
DepMicroDiff: Diffusion-Based Dependency-Aware Multimodal Imputation for Microbiome Data
Rabeya Tus Sadia, Qiang Cheng
机构
*
Department of Computer Science University of Kentucky(计算机科学系 哥伦比亚大学)
;
Department of Computer Science, Institute for Biomedical Informatics University of Kentucky(计算机科学系 生物医学信息学研究所 哥伦比亚大学)