Controlling Multimodal LLMs via Reward-guided Decoding
Oscar Mañas, Pierluca D'Oro, Koustuv Sinha, Adriana Romero-Soriano, Michal Drozdzal, Aishwarya Agrawal
机构
*
Mila - Quebec AI Institute(魁北克AI研究院)
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
Meta FAIR
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG
An end-to-end-trained vision-language model for native-language prostate pathology report generation
用于生成本土语言前列腺病理报告的端到端训练视觉-语言模型
Christian Grashei, Fabian Gülhan, Maximilian Legnar, Fabian Stögbauer, Cleo-Aron Weis, Carolin Mogler, Peter Schüffler
机构
*
Technical University of Munich(慕尼黑工业大学)
;
Munich Data Science Institute(慕尼黑数据科学研究所)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
University Hospital Heidelberg(海德堡大学医院)
;
Heidelberg University(海德堡大学)
;
Interdisciplinary Center for Scientific Computing (IWR)(跨学科科学计算中心(IWR))
AutoSchema: Live Schema Grounding for Agentic Text-to-Sparql over Heterogeneous Knowledge Graphs
AutoSchema:面向异构知识图谱的智能体文本转SPARQL的实时模式接地
Yiming Zhang, Koji Tsuda
机构
*
The University of Tokyo(东京大学)
;
National Institute for Materials Science(国立材料科学研究所)
;
RIKEN Center for Advanced Intelligence Project(理化学研究所高级智能项目中心)
Grounding Large Language Models as Generalizable Policies in Network Control
大语言模型作为网络优化的通用策略
Duo Wu, Linjia Kang, Zhimin Wang, Fangxin Wang, Wei Zhang, Chongbo Sun, Xuefeng Tao, Wei Yang, Le Zhang, Wenwu Zhu, Peng Cui, Zhi Wang
机构
*
Bytedance(字节跳动)
;
Shenzhen International Graduate School(深圳国际研究生院)
;
Tsinghua University(清华大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Department of Computer Science and Technology(计算机科学与技术系)
Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting
基于语义三维高斯溅射的开放词汇移动操作具身多模态定位
Huosen Ou, Dongni Song, Yuncong Wang, Tao Zhou, Yiding Ji
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Midea Group(美的集团)
;
The Hong Kong University of Science and Technology(香港科技大学)
JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures
JEPA-DNA:通过联合嵌入预测架构夯实基因组基础模型
Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
机构
*
Applied AI Architecture, NVIDIA, Israel(NVIDIA应用人工智能架构,以色列)
;
Worldwide Field Ops, NVIDIA, Israel(NVIDIA全球现场运营,以色列)
;
Developer Programs, NVIDIA, Israel(NVIDIA开发者计划,以色列)
;
Cancer Research Center and Wohl Institute of Translational Medicine, Sheba Medical Center, Tel Hashomer, Israel(癌症研究中心和Wohl转化医学研究所,Sheba医疗中心,Tel Hashomer,以色列)
;
Windreich Department of AI and Human Health, Icahn School of Medicine at Mount Sinai, New York, USA(AI与人类健康风reich部门,Mount Sinai医学中心,纽约,美国)