Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models
Hangzhou He, Lei Zhu, Kaiwen Li, Xinliang Zhang, Jiakui Hu, Ourui Fu, Zhengjian Yao, Yanye Lu
机构
*
Department of Biomedical Engineering, College of Future Technology, Peking University(生物医学工程系,未来技术学院,北京大学)
;
Institute of Medical Technology, Peking University Health Science Center, Peking University(医学技术研究所,北京大学医学部,北京大学)
;
National Biomedical Imaging Center, College of Future Technology, Peking University(国家生物医学成像中心,未来技术学院,北京大学)
Automated Procedural Analysis via Video-Language Models for AI-assisted Nursing Skills Assessment
Shen Chang, Dennis Liu, Renran Tian, Kristen L. Swartzell, Stacie L. Klingler, Amy M. Nagle, Nan Kong
机构
*
Weldon School of Biomedical Engineering, Purdue University(普渡大学生物医学工程学院)
;
Department of Industrial and Operations Engineering, University of Michigan(密歇根大学工业与运作工程系)
;
Edward P. Fitts Department of Industrial and Systems Engineering, North Carolina State University(北卡罗来纳州立大学工业与系统工程系)
;
School of Nursing, Purdue University(普渡大学护理学院)
The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
Abdelrahman Sadallah, Tim Baumgärtner, Iryna Gurevych, Ted Briscoe
机构
*
NLP Department, Mohamed Bin Zayed University of Artificial Intelligence(马尔代夫比兹艾兹大学人工智能学院自然语言处理系)
;
Ubiquitous Knowledge Processing Lab, Department of Computer Science(计算机科学系通用知识处理实验室)
;
Hessian Center for AI (hessian.AI), TU Darmstadt(图尔努尔德马斯特大学海斯塞人工智能中心)
机构
*
College of Intelligence and Computing, Tianjin University, Tianjin, China(智能与计算学院,天津大学,天津,中国)
;
PipeChina Institute of Science and Technology, Tianjin, China(中石油科技研究院,天津,中国)
OTAS: Open-vocabulary Token Alignment for Outdoor Segmentation
Simon Schwaiger, Stefan Thalhammer, Wilfried Wöber, Gerald Steinbauer-Wagner
机构
*
Graz University of Technology, Faculty of Computer Science and Biomedical Engineering, Institute of Software Engineering and Artificial Intelligence(格拉茨技术大学,计算机科学与生物医学工程学院,软件工程与人工智能研究所)
;
University of Applied Sciences Technikum Wien, Faculty of Industrial Engineering, Research Group Digital Manufacturing, Automation and Robotics(应用科学大学技术学院,工业工程学院,数字制造、自动化与机器人研究组)
;
University of Natural Resources and Life Sciences, Department of Integrative Biology and Biodiversity Research, Institute for Integrative Nature Conservation Research(自然资源与生命科学大学,整合生物学与生物多样性研究部门,整合自然保护研究 institute)
机构
*
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
State Key Laboratory of Cognitive Intelligence, iFLYTEK(认知智能国家重点实验室)
专题命中
文档图表理解
:vision language model(title);vision-language model(abstract);分类 cs.CV、cs.AI
Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models
Jun Ling, Yao Qi, Tao Huang, Shibo Zhou, Yanqin Huang, Jiang Yang, Ziqi Song, Ying Zhou, Yang Yang, Heng Tao Shen, Peng Wang
机构
*
School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
;
Research Center for Scientific Data Hub, Zhejiang Lab, Hangzhou, China(浙江实验室科学数据中心研究中心)
;
School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
专题命中
文档图表理解
:multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
机构
*
University of Chinese Academy of Sciences (UCAS)(中国科学院大学)
;
New Laboratory of Pattern Recognition (NLPR), CASIA(中国科学院自动化所模式识别新实验室)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(中国科学院多模态人工智能系统国家重点实验室)
;
Hong Kong Institute of Science & Innovation, CASIA(中国科学院香港创新科学研究院)
;
PolyU
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
Songsheng Wang, Rucheng Yu, Zhihang Yuan, Chao Yu, Feng Gao, Yu Wang, Derek F. Wong
机构
*
NICS-EFC Lab, Department of Electronic Engineering, Tsinghua University(清华大学电子工程系NICS-EFC实验室)
;
Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系)
;
Infinigence AI
;
Tsinghua University(清华大学)
;
Zhongguancun Academy(中关村学院)
专题命中
GUI与屏幕智能体
:visual language model(abstract);分类 cs.AI、cs.LG
Comments13 pages, 5 figures, Accepted by EMNLP 2025 (main conference)
机构
*
National Engineering Research Center for Big Data Technology and System(大数据技术与系统国家工程研究中心)
;
Services Computing Technology and System Lab(服务计算技术与系统实验室)
;
Cluster and Grid Computing Lab(集群与网格计算实验室)
;
Hubei Engineering Research Center on Big Data Security(湖北省大数据安全工程研究中心)
;
Hubei Key Laboratory of Distributed System Security(湖北省分布式系统安全重点实验室)
;
School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学信息科学与工程学院)
;
School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)
;
Department of Computer Science, City University of HongKong(香港城市大学计算机科学系)
;
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)
Video-to-BT: Generating Reactive Behavior Trees from Human Demonstration Videos for Robotic Assembly
Xiwei Zhao, Yiwei Wang, Yansong Wu, Fan Wu, Teng Sun, Zhonghua Miao, Sami Haddadin, Alois Knoll
机构
*
Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所)
;
Technical University of Munich(慕尼黑技术大学)
;
Aalto University(艾尔沃斯大学)
;
Shanghai University(上海大学)
;
Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
Neural Antidote: Class-Wise Prompt Tuning for Purifying Backdoors in CLIP
Jiawei Kong, Hao Fang, Sihang Guo, Chenxi Qing, Kuofeng Gao, Bin Chen, Shu-Tao Xia, Ke Xu
机构
*
Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)
;
School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)