Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
合成血管和病理增强视觉-语言模型推理
Chenjun Li, Cheng Wan, Laurin Lux, Alexander Berger, Richard B. Rosen, Martin J. Menten, Johannes C. Paetzold
机构
*
Cornell University(康奈尔大学)
;
Weill Cornell Medicine(韦尔·康奈尔医学)
;
Technical University of Munich(慕尼黑技术大学)
;
New York Eye and Ear Infirmary of Mount Sinai(圣文森特医院)
;
Cornell Tech(康奈尔科技)
机构
*
Shenzhen Stomatology Hospital (Pingshan) of Southern Medical University(南方医科大学深圳口腔医院(平山))
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
State Key Laboratory of Membrane Biology, Beijing Key Laboratory of Cardiometabolic Molecular Medicine, Institute of Molecular Medicine, National Biomedical Imaging Center, School of Future Technology, Peking University(北京大学膜生物学国家重点实验室、北京心代谢分子医学重点实验室、分子医学研究院、国家生物医学成像中心、未来技术学院)
;
Freedom AI
;
Division of Applied Oral Sciences & Community Dental Care Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院应用口腔科学与社区牙科护理系)
;
Beijing Institute of Collaborative Innovation(北京协同创新研究院)
;
National Health Data Institute, Shenzhen(深圳国家健康数据研究院)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Shenzhen Institute of Big Data(深圳大数据研究院)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Benchmarking the Generality of Vision-Language-Action Models
对视觉-语言-动作模型通用性的基准测试
Pranav Guruprasad, Sudipta Chowdhury, Harsh Sikka, Mridul Sharma, Helen Lu, Sean Rivera, Aryan Khurana, Hangliang Ren, Yangyue Wang
机构
*
Manifold Research
;
Metarch AI
;
Georgia Tech(佐治亚理工学院)
;
Tufts University(塔夫茨大学)
;
Northeastern University(东北大学)
;
Birla Institute of Technology and Science, Pilani(比拉理工学院,帕利尼)
;
Institute for Research and Innovation in Intelligent Systems (IRIIS)(智能系统研究与创新研究所)
专题命中
GUI与屏幕智能体
:vision language model(abstract);grounding(abstract);分类 cs.LG
Surveillance Video-Based Traffic Accident Detection Using Transformer Architecture
基于监控视频的交通事故检测使用变换器架构
Tanu Singh, Pranamesh Chakraborty, Long T. Truong
机构
*
Department of Civil Engineering, Indian Institute of Technology Kanpur(印度理工学院坎浦尔分校土木工程系)
;
School of Computing, Engineering and Mathematical Sciences, La Trobe University(拉特罗布大学计算科学与工程数学科学学院)
专题命中
VLM训练与架构
:vision language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI
TV2TV: A Unified Framework for Interleaved Language and Video Generation
TV2TV:一种用于交错语言和视频生成的统一框架
Xiaochuang Han, Youssef Emad, Melissa Hall, John Nguyen, Karthik Padthe, Liam Robbins, Amir Bar, Delong Chen, Michal Drozdzal, Maha Elbayad, Yushi Hu, Shang-Wen Li, Sreya Dutta Roy, Jakob Verbeek, XuDong Wang, Marjan Ghazvininejad, Luke Zettlemoyer, Emily Dinan