The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
机构 * University of Virginia(弗吉尼亚大学) ; Adobe(Adobe公司)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Journal ref CVPR 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * University of Virginia(弗吉尼亚大学) ; Adobe(Adobe公司)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Journal ref CVPR 2025
机构 * University of Waterloo(滑铁卢大学) ; Vector Institute, Toronto(多伦多向量研究所)
专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
机构 * ARTORG Center for Biomedical Engineering Research, University of Bern, Switzerland(ARTORG生物医学工程研究中心,伯尔尼大学,瑞士) ; Shanghai Jiao Tong University, China(上海交通大学,中国) ; Kaiko.AI, Switzerland(Kaiko.AI,瑞士) ; Dept. of Radiation Oncology, Inselspital, Bern University Hospital(放射肿瘤科,因斯普尔茨医院,伯尔尼大学医院)
专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.AI
专题命中 多模态评测 :multimodal(abstract);分类 cs.CL
专题命中 多模态评测 :multimodal(abstract);分类 cs.CV
Comments 16 pages manuscript, 5 figures, 9 pages supplementary material
专题命中 多模态评测 :multimodal(abstract)
机构 * UMR IRIT University Toulouse Capitole(IRIT大学)
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI
机构 * Michigan State University(密歇根州立大学) ; Amazon(亚马逊) ; Northeastern University(东北大学) ; Northwestern University(西北大学) ; University of Central Florida(佛罗里达中央大学)
专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI、cs.MM
专题命中 多模态Agent :multimodal(abstract)
Comments Accepted in 6th World Conference on Artificial Intelligence: Advances and Applications (WCAIAA 2025)
机构 * School of Computer Science and Information Engineering, Hefei University of Technology, China(计算机科学与信息工程学院,合肥工业大学,中国) ; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China(人工智能研究所,合肥综合性国家科学中心,中国)
专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);image-text(abstract)
Comments 12 pages, 6 figures
机构 * EPFL(瑞士联邦理工学院) ; University of Basel(巴塞尔大学) ; HSLU(苏黎世联邦理工学院)
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments NeurIPS 2025 camera-ready
机构 * OPPO AI Center(OPPO人工智能中心)
专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
机构 * Allen Institute for AI(艾伦人工智能研究所) ; University of Chicago(芝加哥大学) ; Stony Brook University(石溪大学)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2025 Findings ("Text or Pixels? Evaluating Efficiency and Understanding of LLMs with Visual Text Inputs")
机构 * Radar Technology Research Institute, School of Information and Electronics, Beijing Institute of Technology(雷达技术研究院,信息电子学院,北京理工大学) ; School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) ; Innovative Equipment Research Institute, Beijing Institute of Technology(创新装备研究院,北京理工大学) ; Key Laboratory of Electronic and Information Technology in Satellite Navigation (Beijing Institute of Technology), Ministry of Education(卫星导航电子信息技术重点实验室(北京理工大学),教育部) ; Beijing Racobit Electronic Information Technology Co., Ltd.(北京瑞科比特电子信息技术有限公司)
专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments Submitted to Pattern Recognition
机构 * School of New Media and Communication, Tianjin University(新媒体与传播学院,天津大学) ; Baidu(百度) ; Shanghai Jiao Tong University(上海交通大学) ; School of Computer Science and Engineering, Beihang University(计算机科学与工程学院,北航)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract)
专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.MM
机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,中国) ; University of Chinese Academy of Sciences (CAS), China(中国科学院大学(中国科学院))
专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI
机构 * University of Washington(华盛顿大学) ; University of Copenhagen(哥本哈根大学)
专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV
机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) ; Guangdong Provincial Key Laboratory of Novel Security Intelligence Technologies(广东省新型安全智能技术重点实验室)
专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI
Comments EMNLP 2025 Main Conference
专题命中 其他多模态 :multimodal(abstract)
Comments 11 pages, 7 figures