arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

International Journal of Computer Vision · 期刊 · Computer Vision

至 收录 556
2505.05474 2026-08-13 cs.CV 版本更新

3D Scene Generation: A Survey

3D场景生成:一项综述

Haozhe Xie, Beichen Wen, Zhaoxi Chen, Fangzhou Hong, Ziwei Liu

机构 * S-Lab, Nanyang Technological University, Singapore(南洋理工大学S实验室)

AI总结 本综述系统梳理3D场景生成的四大范式方法,分析其技术基础与挑战,展望物理感知生成等方向,为该领域发展提供参考。

Comments Accepted by IJCV. Project Page: https://github.com/hzxie/Awesome-3D-Scene-Generation

URL PDF HTML 收藏
2608.10954 2026-08-12 cs.CV cs.AI 新提交

Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

复杂城市场景中基于证据的可信多模态推理与评估基准

Zhaoyang Wei, Bowen Jiang, Xumeng Han, Jiashu Li, Xuehui Yu, Yuling Liu, Guorong Li, Zhenjun Han, Jianbin Jiao

机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) Tencent CDG(腾讯云与智慧产业事业群) Institute of Information Engineering, CAS(中国科学院信息工程研究所)

AI总结 针对复杂城市场景中多模态大语言模型的推理可靠性问题,提出AD2-Bench基准与EGVOR模型,提升了不利条件下的多模态推理稳定性。

Comments Accepted by IJCV

URL PDF HTML 收藏
2507.21606 2026-08-12 cs.CV 版本更新

Exploring Decoupled Spatio-Temporal Consistency Learning and Self-Prompting Evolution for Self-Supervised Tracking

探索用于自监督跟踪的解耦时空一致性学习与自提示演化

Yaozong Zheng, Bineng Zhong, Qihua Liang, Ning Li, Haiying Xia, Shuxiang Song, Rongrong Ji

AI总结 本研究提出SSTrack++自监督跟踪模型,通过弱到强自监督框架、解耦时空一致性策略等,在十个基准数据集上超越SOTA自监督跟踪方法,显著缩小与全监督跟踪器的性能差距。

Comments IJCV 2026

URL PDF HTML 收藏
2607.14821 2026-07-17 cs.CV 新提交

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification

模糊模态边界:从单模态到多模态行人重识别的统一综述

Xiao Wang, Bing Wang, Bin Yang, Cuiqun Chen, Xin Xu, Mang Ye

机构 * Wuhan University of Science and Technology(武汉科技大学) Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System(湖北省智能信息处理与实时工业系统重点实验室) Wuhan University(武汉大学) Anhui University(安徽大学)

AI总结 综述行人重识别领域从单模态向多模态发展的转变,系统回顾关键跨模态任务,研究多模态融合ReID,提出基于Transformer的可见光-红外ReID基线框架,并概述未来研究方向。

Comments 61 pages, 6 figures, journal

Journal ref International Journal of Computer Vision.2026

URL PDF HTML 收藏
2603.27948 2026-07-13 cs.CV 版本更新

RehearsalNeRF: Decoupling Intrinsic Neural Fields of Dynamic Illuminations for Scene Editing

RehearsalNeRF:解耦动态光照的内在神经场以实现场景编辑

Changyeon Won, Hyunjun Jung, Jungu Cho, Seonmi Park, Chi-Hoon Lee, Hae-Gon Jeon

机构 * Gwangju Institute of Science and Technology(光州科学技术院) Yonsei University(延世大学) CJ Corporation(CJ集团)

AI总结 RehearsalNeRF通过利用稳定光照下的场景数据,解耦动态光照下的场景辐射,提升动态光照下的视图合成和场景编辑性能。

Comments Accepted to the International Journal of Computer Vision (IJCV). Changyeon Won and Hyunjun Jung contributed equally to this work

URL PDF HTML 收藏
2312.09076 2026-07-13 cs.CV 版本更新

ProSGNeRF: Progressive Dynamic Neural Scene Graph with Frequency Modulated Foundation Model in Urban Scenes

ProSGNeRF:城市场景中基于频率调制基础模型的渐进式动态神经场景图

Tianchen Deng, Yanbo Wang, Yejia Liu, Chenpeng Su, Jingchuan Wang, Danwei Wang, Shao-Yuan Lo, Weidong Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) National Taiwan University(国立台湾大学)

AI总结 针对大规模城市场景中快速移动对象建模及相机自我运动处理难题,提出ProSGNeRF,通过渐进式场景图网络架构学习局部场景表示,利用基础模型网络编码潜码并引入频率调制模块,提升视图合成等能力。

Comments Accepted by IJCV 2026

URL PDF HTML 收藏
2607.07438 2026-07-09 cs.CV 新提交

Two-Stage Multi-Modal Fusion with Adaptive Alignment for Action Quality Assessment

用于动作质量评估的具有自适应对齐的两阶段多模态融合

Kanglei Zhou, Ruizhi Cai, Xinning Wang, Yijian Zheng, Liyuan Wang, Jianguo Li, Xiaohui Liang

机构 * Tsinghua University(清华大学) Beihang University(北京航空航天大学) Children’s Hospital, Capital Institute of Pediatrics(首都儿科研究所附属儿童医院) Zhongguancun Laboratory(中关村实验室)

AI总结 研究动作质量评估问题,提出两阶段多模态融合框架DualAlign,通过自适应对齐解决跨模态错位等挑战,引入MM-JDM数据集,实验表明该框架在多模态评估中表现出色,相比现有方法有显著提升且在特定条件下稳健。

Comments Accepted to IJCV

URL PDF HTML 收藏
2607.03696 2026-07-07 cs.CV 新提交

IPDiff: Diffusion-driven ORSI Salient Object Detection with Information Reconstruction and Multi-Prior Guidance

IPDiff:基于信息重构和多先验引导的扩散驱动光学遥感图像显著目标检测

Gongyang Li, Zhen Bai, Runmin Cong, Dan Zeng, Weisi Lin, Xiao-Ping Zhang

机构 * School of Communication and Information Engineering, Shanghai University(上海大学通信与信息工程学院) Department of Medical Equipment, the First Affiliated Hospital of Zhengzhou University(郑州大学第一附属医院医学装备部) School of Control Science and Engineering, Shandong University(山东大学控制科学与工程学院) School of Computer Science and Engineering, Nanyang Technological University(南洋理工大学计算机科学与工程学院) Shenzhen Key Laboratory of Ubiquitous Data Enabling, Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院深圳泛在数据赋能重点实验室)

AI总结 提出IPDiff,一种基于独特动态优化策略的扩散驱动光学遥感图像显著目标检测方法,通过提取多先验信息,在去噪网络中迭代去噪生成显著图,并经混合损失函数监督训练,性能优于46种先进方法。

Comments 22 pages, 8 figures, accepted by IJCV 2026

URL PDF HTML 收藏
2203.14308 2026-06-24 cs.CV

Temporal Transductive Inference for Few-Shot Video Object Segmentation

时间传递推断用于少样本视频对象分割

Mennatullah Siam, Konstantinos G. Derpanis, Richard P. Wildes

机构 * Electrical Engeering and Computer Science, York University, ON, Canada(电气工程与计算机科学系,约克大学,加拿大)

AI总结 本文提出一种简单有效的时序传递推断方法,利用未标记视频帧中的时间一致性,通过全局和局部时间约束提升时间一致性并减少过拟合,实验显示在YouTube-VIS上平均交并比优于现有元学习方法2.8%。

Comments IJCV submission under review

URL PDF HTML 收藏
2603.05963 2026-06-23 cs.CV cs.AI 版本更新

Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models

骨架到图像编码:通过视觉预训练模型实现骨架表示学习

Siyuan Yang, Jun Liu, Hao Cheng, Chong Wang, Shijian Lu, Hedvig Kjellstrom, Weisi Lin, Alex C. Kot

机构 * KTH Royal Institute of Technology(皇家理工学院) Lancaster University(兰卡斯特大学) Hebei University of Technology(河北工业大学) Nanyang Technological University(南洋理工大学) Shenzhen MSU-BIT University(深圳MSU-BIT大学) VinUniversity(文大学)

AI总结 提出S2I编码,将骨架序列转换为图像状数据,首次利用视觉预训练模型进行自监督骨架表示学习,统一异构骨架格式,在多个数据集上验证了有效性。

Comments Submitted to IJCV, under review

URL PDF HTML 收藏
2412.05976 2026-06-23 cs.CV 版本更新

LightOcc: Lightweight Spatial Embedding for Efficient Vision-based 3D Occupancy Prediction

LightOcc: 用于高效基于视觉的3D占用预测的轻量级空间嵌入

Jinqing Zhang, Yanan Zhang, Wenrui Cai, Qingjie Liu, Yunhong Wang

机构 * State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室) Beihang University(北京航空航天大学) School of Computer Science and Information Engineering(计算机科学与信息工程学院) Hefei University of Technology(合肥工业大学) Hangzhou Innovation Institute(杭州创新研究院)

AI总结 针对现有占用预测方法使用体素特征导致计算和内存开销大的问题,提出轻量级空间嵌入(Lightweight Spatial Embedding),通过单通道占用和空间到通道机制压缩高度信息,结合跨视图交互和边缘感知嵌入,在多个基准上实现最先进性能并显著提升效率。

Comments Accepted by International Journal of Computer Vision (IJCV), 2026

URL PDF HTML 收藏
2603.26551 2026-06-17 cs.CV cs.AI 版本更新

Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones

超越MACs:面向视觉骨干网络的硬件高效架构设计

Moritz Nottebaum, Matteo Dunnhofer, Christian Micheloni

机构 * Machine Learning and Perception Lab, University of Udine(乌迪大学机器学习与感知实验室) Centre for Vision Research, York University(约克大学视觉研究中心)

AI总结 针对MACs指标在边缘设备上的不足,提出基于硬件效率洞察的LowFormer骨干网络,通过轻量级Lowtention模块实现显著加速。

Comments Accepted at International Journal of Computer Vision (IJCV)

Journal ref Int J Comput Vis 134, 295 (2026)

URL PDF HTML 收藏
2401.14381 2026-06-17 cs.LG math.DG 版本更新

Manifold GCN: Diffusion-based Convolutional Neural Network for Manifold-valued Graphs

Manifold GCN:基于扩散的流形值图卷积神经网络

Martin Hanik, Gabriele Steidl, Christoph von Tycowicz

机构 * BIFOLD—Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究院) Technical University Berlin(柏林技术大学) Zuse Institute Berlin(柏林泽尼茨研究所)

AI总结 提出两种适用于黎曼流形特征图的图神经网络层:基于流形值图扩散方程的扩散层和受向量神经元启发的切向多层感知器,两者在节点置换和流形等距下等变,在更广泛问题上优于任务特定网络。

Comments Extended ADNI experiment

Journal ref International Journal of Computer Vision, Volume 134, article number 315 (2026)

URL PDF HTML 收藏
2606.15570 2026-06-16 cs.CV 新提交

An Extensive Benchmark for Single-round and Multi-round Instruction-based Image Editing

单轮与多轮指令式图像编辑的广泛基准

Yiwei Ma, Ke Ye, Weihuang Lin, Jiayi Ji, Xiaoshuai Sun, Tat-Seng Chua, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室) National University of Singapore(新加坡国立大学)

AI总结 提出I2EBench2.0基准,通过16个单轮和7个多轮维度评估指令式图像编辑模型,结合用户研究确保与人类判断一致,并基于八种模型分析提供研究指导。

Comments Accepted by International Journal of Computer Vision (IJCV), 2026

URL PDF HTML 收藏
2604.18866 2026-06-16 cs.CV 版本更新

HMR-Net: Hierarchical Modular Routing for Cross-Domain Object Detection in Aerial Images

HMR-Net: 用于航拍图像跨域目标检测的层次化模块化路由

Pourya Shamsolmoali, Masoumeh Zareapoor, Michael Felsberg, Nick Pears, Huiyu Zhou, Yue Lu

机构 * Department of Computer Science, University of York(约克大学计算机科学系) SEIEE, Shanghai Jiao Tong University(上海交通大学SEIEE) Computer Vision Laboratory, Linkoping University(林哈姆大学计算机视觉实验室) School of Computing and Mathematical Sciences, University of Leicester(莱斯特大学计算与数学科学学院) SCEE, East China Normal University(华东师范大学SCEE)

AI总结 提出层次化模块化路由框架,通过领域路由和场景路由实现跨数据集和复杂场景下的结构化专业化,并利用条件专家模块支持零样本新类别检测。

Comments Submitted to IJCV September 2025

URL PDF HTML 收藏
2508.07797 2026-06-16 cs.CV 版本更新

Power Battery Detection

动力电池检测

Xiaoqi Zhao, Peiqian Cao, Chenyang Yu, Zonglei Feng, Lihe Zhang, Hanqi Liu, Jiaming Zuo, Youwei Pang, Jinsong Ouyang, Weisi Lin, Georges El Fakhri, Huchuan Lu, Xiaofeng Liu

机构 * Yale University, USA(耶鲁大学,美国) Dalian University of Technology, China(大连理工大学,中国) Volkswagen Automotive Co., Ltd(大众汽车有限公司) X3000 Inspection Co., Ltd(X3000检测有限公司) Nanyang Technological University, Singapore(南洋理工大学,新加坡)

AI总结 针对动力电池X射线图像中极板端点定位任务,提出首个大规模基准PBD5K和点级分割模型MDCNeXt,通过多维度结构线索与状态空间模块提升检测精度。

Comments Accepted by International Journal of Computer Vision (IJCV). Code: https://github.com/NTU-AI4X/X-ray-PBD

URL PDF HTML 收藏
2605.17773 2026-06-11 cs.CV

PlantPose: Universal Plant Skeleton Estimation via Tree-constrained Graph Generation

PlantPose: 通过树约束图生成实现通用植物骨架估计

Xinpeng Liu, Hiroaki Santo, Yosuke Toda, Fumio Okura

机构 * Graduate School of Information Science and Technology, Osaka University(大阪大学信息科学与技术研究生院) Phytometrics(Phytometrics公司) Institute of Transformative Bio-Molecules, Nagoya University(名古屋大学变革生物分子研究所)

AI总结 本文提出PlantPose,一种通过树约束图生成实现通用植物骨架估计的方法,通过结合学习基于图生成和传统图算法,提高模型的泛化能力,并在多个领域实现了鲁棒且准确的植物骨架估计。

Comments International Journal of Computer Vision, 2026

URL PDF HTML 收藏
2604.04554 2026-06-09 cs.CV cs.RO 版本更新

Relational Epipolar Graphs for Robust Relative Camera Pose Estimation

基于关系epipolar图的鲁棒相对相机姿态估计

Prateeth Rao, Sachit Rao

机构 * International Institute of Information Technology(国际信息科技研究所)

AI总结 本文提出基于epipolar图的关系推断方法,用于估计相对相机姿态,通过图操作估计旋转、平移和本质矩阵,提升对密集噪声和大基线变化的鲁棒性。

Comments 21 pages, 11 figures, 11 Tables, Submitted to IJCV

URL PDF HTML 收藏
2606.06338 2026-06-05 cs.CV

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

StoryVideoQA: 通过大规模、多类型和自动生成的数据集扩展深度视频理解

Zhengqian Wu, Zhixian Liu, Aodong Chen, Jingyang Zhang, Ruizhe Li, Hanlin Ge, Zhongyuan Wang, Chunxia Xiao, Chao Liang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) National Engineering Research Center for Multimedia Software(多媒体软件国家工程研究中心) Hubei Key Laboratory of Multimedia and Network Communication Engineering(湖北省多媒体与网络通信工程重点实验室)

AI总结 提出StoryVideoQA数据集和PlotTree方法,通过多智能体协作框架自动生成大规模深度视频理解问答对,并利用层次化情节结构提升复杂故事线推理能力。

Comments Accepted by IJCV 2026

Journal ref International Journal of Computer Vision (2026)

URL PDF HTML 收藏
2410.04960 2026-06-05 cs.CV

On Efficient Variants of Segment Anything Model: A Survey

关于高效分段任何模型的变体:一项调查

Xiaorui Sun, Jun Liu, Heng Tao Shen, Xiaofeng Zhu, Ping Hu

机构 * School of Computer Science and Engineering(计算机科学与工程学院) School of Computing and Communications(计算与通信学院) School of Computer Science and Technology(计算机科学与技术学院)

AI总结 本文综述了高效分段任何模型变体的研究,探讨了提升效率的同时保持准确性的核心技术和方法,并评估了不同硬件上的性能。

Comments IJCV

URL PDF HTML 收藏
1604.08714 2026-06-04 math.NA cs.NA

Iterative Multiplicative Filters for Data Labeling

迭代乘法滤波器用于数据标注

Ronny Bergmann, Jan Henrik Fitschen, Johannes Persch, Gabriele Steidl

AI总结 本文提出了一种新的迭代乘法过滤算法,用于标签分配矩阵的监督数据分割,通过高效分配标签提升数据处理效果。

Journal ref International Journal of Computer Vision, 2017, 123, p. 435-53

URL PDF HTML 收藏
2410.21361 2026-06-02 cs.CV cs.LG

Domain Adaptation with a Single Vision-Language Embedding

基于单一视觉-语言嵌入的域适应

Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick Pérez, Raoul de Charette

机构 * Inria(法国国家信息与自动化研究所) Kyutai(Kyutai公司)

AI总结 提出一种利用单一视觉-语言(VL)嵌入进行域适应的框架,通过提示/照片驱动的实例归一化(PIN)挖掘多种视觉风格,实现零样本和单样本无监督域适应,在语义分割任务上优于基线方法。

Comments International Journal of Computer Vision (IJCV 2026)

URL PDF HTML 收藏
2511.04711 2026-05-27 cs.CR cs.AI cs.LG

SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking

SWAP:通过顺序水印实现软提示的版权审计

Wenyuan Yang, Yichen Sun, Changzheng Chen, Zhixuan Chu, Jiaheng Zhang, Yiming Li, Dacheng Tao

机构 * Sun Yat-sen University(中山大学) Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

AI总结 针对软提示的版权保护问题,提出一种基于顺序水印的审计方法SWAP,通过将水印嵌入到更复杂的输出分布顺序空间中,实现无害且鲁棒的版权验证。

Comments This paper has been accepted by the International Journal of Computer Vision (IJCV), 2026. The first two authors contributed equally to this work. 28 pages

URL PDF HTML 收藏
2504.15404 2026-05-26 cs.CV

Context Aware Grounded Teacher for Source Free Object Detection

上下文感知的接地教师用于无源目标检测

Tajamul Ashraf, Rajes Manna, Partha Sarathi Purkayastha, Tavaheed Tariq, Janibul Bashir

机构 * Department of Computer Vision(计算机视觉系) MBZUAI Microsoft Research India(微软印度研究院) GAASH Research Lab(GAASH研究实验室) NIT Srinagar(斋普尔理工学院)

AI总结 针对无源目标检测中类别不平衡导致的上下文偏差和噪声伪标签问题,提出一种基于关系上下文模块和语义增强的偏差感知框架Grounded Teacher,通过关系正则化和语义增强提升少数类检测性能。

Comments Accepted in International Journal of Computer Vision (IJCV); Project Webpage: https://tajamul21.github.io/Grounded_Teacher/

URL PDF HTML 收藏
2403.11247 2026-05-14 cs.CV cs.RO

Compact 3D Gaussian Splatting For Dense Visual SLAM

紧凑的3D高斯散射用于密集视觉SLAM

Tianchen Deng, Chang Nie, Shuhong Liu, Wenhua Wu, Jianfei Yang, Shenghai Yuan, Jiuming Liu, Danwei Wang, Hesheng Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) The University of Tokyo(东京大学) Harvard University(哈佛大学) University of Cambridge(剑桥大学)

AI总结 本文提出紧凑的3D高斯散射SLAM系统,通过滑动窗口遮罩策略和几何代码本压缩高斯参数,提升训练和渲染速度,保持高质量场景表示。

Comments Accepted by IJCV 2026

URL PDF HTML 收藏
2501.03717 2026-05-13 cs.CV cs.AI cs.GR

Materialist: Physically Based Editing Using Single-Image Inverse Rendering

Materialist: 基于物理的单图像逆渲染中的材质编辑

Lezhong Wang, Duc Minh Tran, Ruiqi Cui, Thomson TG, Anders Bjorholm Dahl, Siavash Arjomand Bigdeli, Jeppe Revall Frisvad, Manmohan Chandraker

机构 * Technical University of Denmark(丹麦技术大学) University of California San Diego(加州大学圣地亚哥分校)

AI总结 Materialist提出一种基于物理的单图像逆渲染方法,利用神经网络预测初始材质属性并优化,实现材质编辑、物体插入和光照调整,同时引入高效的光线追踪折射方法进行透明度编辑。

Comments More Comprehensive IJCV Camera-Ready Version. Project website: https://lez-s.github.io/materialist_project/

Journal ref International Journal of Computer Vision (IJCV), 134(6), 267 (2026)

URL PDF HTML 收藏
2408.06747 2026-05-11 cs.CV

ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation

ReCLIP++: 学习校正CLIP的偏差以实现无监督语义分割

Jingyun Wang, Guoliang Kang

机构 * Beihang University(北航大学)

AI总结 本文提出ReCLIP++,通过显式建模和校正CLIP中的偏差以提升无监督语义分割性能,设计了参考提示和位置嵌入投影来分别编码类别偏好和空间偏好偏差,并通过矩阵乘法生成偏差logit图,再通过元素级减法校正logits,最后利用Gumbel-Softmax操作生成分割掩码。

Comments Extended version of our CVPR 24 paper, accepted by IJCV 2025

URL PDF HTML 收藏
2605.05910 2026-05-08 cs.CV

Plug-and-play Class-aware Knowledge Injection for Prompt Learning with Visual-Language Model

即插即用的类感知知识注入用于具有视觉-语言模型的提示学习

Junhui Yin, Nan Pu, Xinyu Zhang, Lingfeng Yang, Lin Wu, Xiaojie Wang, Zhun Zhong

机构 * University of Science and Technology Beijing(北京科技大学) Hefei University of Technology(合肥工业大学) University of Auckland(奥克兰大学) Nanjing University of Science and Technology(南京理工大学) Swansea University(斯旺西大学) Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 本文提出CAKI框架,通过类特定提示生成和查询-键提示匹配,补充现有方法中的类特定知识,提升基类和新类的性能。

Comments Accepted by International Journal of Computer Vision

URL PDF HTML 收藏
2604.24380 2026-04-28 cs.CL

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency

大型视觉语言模型的结构剪枝:对剪枝动态、恢复和数据效率的全面研究

Yiran Huang, Lukas Thede, Massimiliano Mancini, Wenjia Xu, Zeynep Akata

机构 * Technical University of Munich, Germany(慕尼黑技术大学,德国) Munich Center for Machine Learning, Germany(慕尼黑机器学习中心,德国) Helmholtz Munich, Germany(海德堡-慕尼黑赫尔姆霍尔茨研究中心,德国) University of Tübingen, Tübingen AI Center, Germany(图宾根大学,图宾根人工智能中心,德国) University of Trento, Italy(特伦托大学,意大利) Beijing University of Posts and Telecommunications, China(北京邮电大学,中国)

AI总结 本文研究了通过结构化剪枝压缩现有大型视觉语言模型的方法,发现宽度剪枝在低资源环境下表现更优,结合监督微调和隐藏状态蒸馏可实现高效恢复,仅需5%的数据即可保持95%性能。

Comments Accepted at International Journal of Computer Vision (IJCV) 2026

URL PDF HTML 收藏
2410.05970 2026-04-28 cs.CV cs.AI cs.CL

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

PDF-WuKong:一种用于高效长PDF阅读的大型多模态模型,采用端到端稀疏采样

Xudong Xie, Hao Yan, Liang Yin, Yang Liu, Jing Ding, Minghui Liao, Yuliang Liu, Wei Chen, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Huawei Inc.(华为公司)

AI总结 本文提出PDF-WuKong,一种针对长PDF文档的多模态大语言模型,通过端到端稀疏采样提升多模态问答效率,实验表明其在长文档理解任务中优于其他模型,F1得分平均高出8.6%。

Comments Accepted by International Journal of Computer Vision (IJCV)

URL PDF HTML 收藏