Moving Pictures of Thought: Extracting Visual Knowledge in Charles S. Peirce's Manuscripts with Vision-Language Models
Carlo Teo Pedretti, Davide Picca, Dario Rodighiero
机构
*
Department of Classics, University Sapienza of Rome(罗马大学萨皮恩扎文学院)
;
Department of Language and Communication Sciences, University of Lausanne(洛桑大学语言与沟通科学系)
;
Campus Fryslân, University of Groningen(格罗宁根大学弗里斯兰校区)
专题命中
视觉推理
:vision-language model(title);VLM(abstract);visual language model(abstract);分类 cs.AI、cs.LG
Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
Yunqi Hong, Johnson Kao, Liam Edwards, Nein-Tzu Liu, Chung-Yen Huang, Alex Oliveira-Kowaleski, Cho-Jui Hsieh, Neil Y. C. Lin
机构
*
Computer Science Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校计算机科学系)
;
Mechanical and Aerospace Engineering Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校机械与航空航天工程系)
;
Department of Pathology, Tri-Service General Hospital, National Defense Medical Center, Taipei, Taiwan(台湾国防医学院三军总医院病理部)
;
Department of Pathology, National Taiwan University Hospital, Taipei, Taiwan(台湾国立台湾大学医院病理部)
;
Department of Pathology, David Geffen School of Medicine, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校大卫·Geffen医学院病理部)
;
Bioengineering Department, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校生物工程系)
;
Institute for Quantitative and Computational Biosciences, University of California, CA, USA(加州大学定量与计算生物科学研究所)
专题命中
视觉推理
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
Wenhao Zhou, Hao Zheng, Rong Zhao
机构
*
Center for Brain-Inspired Computing Research (CBICR)(脑启发计算研究中心)
;
Department of Precision Instruments(精密仪器系)
;
IDG/McGovern Institute for Brain Research(IDG/麦戈文脑研究学院)
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
Jeong Hun Yeo, Sangyun Chung, Sungjune Park, Dae Hoe Kim, Jinyoung Moon, Yong Man Ro
机构
*
Integrated Vision and Language Lab., School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(整合视觉与语言实验室,电气工程学院,韩国科学技术院(KAIST))
;
Visual Intelligence Research Section, Superintelligence Creative Research Laboratory, Electronics and Telecommunications Research Institute (ETRI)(视觉智能研究部,超智能创意研究实验室,电子电信研究院)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
ViSS-R1: Self-Supervised Reinforcement Video Reasoning
Bo Fang, Yuxin Song, Qiangqiang Wu, Haoyuan Sun, Wenhao Wu, Antoni B. Chan
机构
*
City University of Hong Kong(香港城市大学)
;
Baidu Inc.(百度公司)
;
Tsinghua University(清华大学)
;
The University of Sydney(悉尼大学)
专题命中
视觉推理
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
CommentsOur paper was initially titled "Video-SSR1: Self-Supervised Reinforcement Video Reasoning." Upon noticing its close resemblance to the title of a recently released paper, we have decided to rename our work as "ViSS-R1."