TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
TimeScope: 向长视频中面向任务的时序定位迈进
Xiangrui Liu, Minghao Qin, Yan Shu, Zhengyang Liang, Yang Tian, Chen Jason Zhang, Bo Zhao, Zheng Liu
机构
*
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院)
;
University of Trento(特伦多大学)
;
Singapore Management University(新加坡管理大学)
HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
HalluShift++: 通过内部表示转移弥合语言与视觉,解决多模态大语言模型中的层级幻觉
Sujoy Nath, Arkaprabha Basu, Sharanya Dasgupta, Swagatam Das
机构
*
Netaji Subhash Engineering College (NSEC)(奈尔贾伊·萨布哈工程学院)
;
TCG Crest
;
Electronics and Communication Sciences Unit (ECSU)(电子与通信科学单位)
;
Indian Statistical Institute(印度统计研究所)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Subgoal Graph-Augmented Planning for LLM-Guided Open-World Reinforcement Learning
子目标图增强的规划用于LLM引导的开放世界强化学习
Shanwei Fan, Bin Zhang, Zhiwei Xu, Yingxuan Teng, Siqi Dai, Lin Cheng, Guoliang Fan
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
School of Artificial Intelligence, Shandong University(山东大学人工智能学院)