Text-guided Feature Disentanglement for Cross-modal Gait Recognition
文本引导的跨模态步态识别特征解耦
Zhiyang Lu, Ming Cheng
机构
*
Fujian Key Laboratory of Urban Intelligent Sensing and Computing, Xiamen University(福建城市智能感知与计算重点实验室,厦门大学)
;
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算重点实验室,中华人民共和国教育部,厦门大学)
SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning
SVSR:一种用于多模态推理的自我验证与自我修正范式
Zhe Qian, Nianbing Su, Zhonghua Wang, Hebei Li, Zhongxing Xu, Yueying Li, Fei Luo, Zhuohan Ouyang, Yanbiao Ma
机构
*
South China Agricultural University(华南农业大学)
;
University of Glasgow(格拉斯哥大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Monash University(莫纳什大学)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Defense Technology(国防科技大学)
;
Renmin University of China(中国人民大学)
;
South China Normal University(华南师范大学)
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
;
University College London(伦敦大学学院)
;
Guangzhou University(广州大学)
;
Wuhan University(武汉大学)
;
Nanjing University(南京大学)
M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language Models
M3DocDep: 多模态、多页、多文档依赖分块方法基于大视觉-语言模型
Joongmin Shin, Jeongbae Park, Jaehyung Seo, Heuiseok Lim
机构
*
Human-inspired AI Research, Korea University(韩国大学人智AI研究所)
;
Computer Science and Engineering, Konkuk University(konkuk大学计算机科学与工程系)
;
Department of Computer Science and Engineering, Korea University(韩国大学计算机科学与工程系)
机构
*
Peking University Shanghai AI Laboratory(北京大学上海人工智能实验室)
;
Nanjing University(南京大学)
;
Peking University(北京大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Zhongguancun Academy Beijing Key Laboratory of Data Intelligence and Security(中关村北京数据智能与安全重点实验室)
PestVL-Net: Enabling Multimodal Pest Learning via Fine-grained Vision-Language Interaction
PestVL-Net: 通过细粒度视觉-语言交互实现多模态害虫学习
Xueheng Li, Tao Hu, Ke Cao, Runsheng Qi, Huixin Zhang, Rui Li, Jie Zhang, Chengjun Xie
机构
*
Institute of Intelligent Machines, Hefei Institutes of Physical Science, Chinese Academy of Sciences(智能机器研究所,合肥物理科学研究院,中国科学院)
;
University of Science and Technology of China(中国科学技术大学)
;
Zhongke Hefei Institute of Technology Innovation Engineering(中科合肥技术创新工程研究院)
Bias-constrained multimodal intelligence for equitable and reliable clinical AI
具有偏见约束的多模态智能用于公平且可靠的临床AI
Cheng Li, Weijian Huang, Jiarun Liu, Hao Yang, Qi Yang, Song Wu, Ye Li, Hairong Zheng, Shanshan Wang
机构
*
Paul C. Lauterbur Research Center for Biomedical Imaging(Paul C. Lauterbur生物医学成像研究中心)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)
;
Pengcheng Laboratory(鹏城实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Department of Radiology, Beijing Chaoyang Hospital, Capital Medical University(首都医科大学北京朝阳医院放射科)
;
Department of Urology, South China Hospital, Medical School, Shenzhen University(深圳大学南方医院泌尿科)
;
Huawei Technologies Co., Ltd.(华为技术有限公司)
Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models
通过上下文对齐的视觉-语言模型实现负责任的多模态医疗推理
Sumra Khan, Sagar Chhabriya, Aizan Zafar, Sheeraz Arif, Amgad Muneer, Anas Zafar, Shaina Raza, Rizwan Qureshi
机构
*
Salim Habib University(萨利姆·哈比卜大学)
;
Institute of Business Administration Sukkur(苏库尔工商管理学院)
;
University of Central Florida(中佛罗里达大学)
;
The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)
;
Toronto Metropolitan University(多伦多都会大学)
;
Vector Institute(向量研究所)
机构
*
Hangzhou City University(杭州城市大学)
;
Hong Kong Baptist University(香港浸会大学)
;
Zhejiang University(浙江大学)
;
The First Affiliated Hospital, Zhejiang University School of Medicine(浙江大学医学院附属第一医院)
;
City University of Hong Kong(香港城市大学)
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
理解并缓解多模态推理链模型中的幻觉
Ji Ma, Wei Suo, Peng Wang, Yanning Zhang
机构
*
School of Computer Science and Ningbo Institute, Northwestern Polytechnical University, China(西北工业大学计算机学院和宁波研究院)
;
National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, China(国家空天地海一体化大数据应用技术国家工程实验室)
机构
*
MoE Key Lab of BIPC, University of Science and Technology of China(脑信息处理实验室,中国科学技术大学)
;
Opus AI Research(Opus AI研究院)
;
Southeast University(东南大学)
;
Nanyang Technological University(南洋理工大学)
MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management
MARCUS:一种用于心脏诊断和管理的代理式多模态视觉-语言模型
Jack W O'Sullivan, Mohammad Asadi, Lennart Elbe, Akshay Chaudhari, Tahoura Nedaee, Francois Haddad, Michael Salerno, Li Fe-Fei, Ehsan Adeli, Rima Arnaout, Euan A Ashley
机构
*
Division of Cardiology, Department of Medicine, Stanford University(斯坦福大学心脏病学系)
;
Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系)
;
Department of Medicine, Radiology, and Pediatrics, UCSF(旧金山大学医学系、放射学与儿科学系)
;
Bakar Institute, UCSF(Bakar研究所,旧金山大学)
;
UCSF–UC Berkeley Joint Program in Computational Precision Health(旧金山大学-伯克利计算精准健康联合计划)
;
Department of Radiology, Stanford University(斯坦福大学放射学系)
;
Department of Psychiatry and Behavioral Sciences, Stanford University(斯坦福大学精神病学与行为科学系)
;
Department of Computer Science, Stanford University(斯坦福大学计算机科学系)
;
Department of Electrical Engineering, Stanford University(斯坦福大学电气工程系)
;
Department of Biology, Stanford University(斯坦福大学生物学系)