arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26172 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 4484 篇

2603.00565 2026-03-03 cs.CV cs.AI cs.CR 62%

MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs

MIDAS: 多图像分散与语义重建用于对抗多模态大语言模型

Yilian Liu, Xiaojun Jia, Guoshun Nan, Jiuyang Lyu, Zhican Chen, Tao Guan, Shuyuan Luo, Zhongyi Zhai, Yang Liu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学) Guilin University of Electronic Technology(桂林电子科技大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 MIDAS通过多图像分散与语义重建技术,提升对抗多模态大语言模型的劫持性能,达到81.46%的平均攻击成功率。

Journal ref The Fourteenth International Conference on Learning Representations(2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06292 2026-03-03 cs.CV cs.AI 62%

ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations

ChainMPQ: 交错文本图像推理链用于缓解关系幻觉

Yike Wu, Yiwei Wang, Yujun Cai

机构 * University of Queensland(昆士兰大学) University of California, Merced(加州大学梅尔德分校) Ant Group(蚂蚁集团)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 ChainMPQ通过交错文本和图像推理链,利用积累的记忆减少大型视觉语言模型的关系幻觉,提升关系推理能力。

Comments Accepted by ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00207 2026-03-03 cs.CV cs.AI 62%

VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models

VisRef: 通过思考进行视觉再聚焦以提高多模态大推理模型的测试时间扩展

Soumya Suvra Ghosal, Youngeun Kim, Zhuowei Li, Ritwick Chaudhry, Linghan Xu, Hongjing Zhang, Jakub Zablocki, Yifan Xing, Qin Zhang

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Amazon(亚马逊) Physion Labs(Physion实验室)

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 VisRef通过视觉再聚焦提升多模态大推理模型测试时间扩展性能,有效解决视觉依赖任务中推理性能下降问题。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00115 2026-03-03 physics.soc-ph cs.AI cs.CV 62%

Multimodal Modular Chain of Thoughts in Energy Performance Certificate Assessment

多模态模块化思维链在能源性能证书评估中的应用

Zhen Peng, Peter J. Bentley

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种基于多模态模块化思维链的低成本EPC预评估方法,通过结构化提示实现对EPC评分的序数结构捕捉,实验表明其在数据稀缺环境下具有显著优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23945 2026-03-02 cs.CV cs.AI cs.MM 62%

PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning

PointCoT: 一种用于显式3D几何推理的多模态基准

Dongxu Zhang, Yiding Sun, Pengcheng Li, Yumou Liu, Hongqiang Lin, Haoran Xu, Xiaoxuan Mu, Liang Lin, Wenbiao Yan, Ning Yang, Chaowei Fang, Juanjuan Zhao, Jihua Zhu, Conghui He, Cheng Tan

机构 * Xi'an Jiaotong University(西安交通大学) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Institute of Automation, CASIA(中国科学院自动化研究所) Taiyuan University of Technology(太原理工大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 PointCoT通过显式链式推理提升3D几何推理能力,提出多模态基准和双流架构,实现对3D点云的高精度理解与推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21048 2026-03-02 cs.CV cs.AI 62%

Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning

Veritas:通过模式感知推理实现通用的深度伪造检测

Hao Tan, Jun Lan, Zichang Tan, Ajian Liu, Chuanbiao Song, Senyuan Shi, Huijia Zhu, Weiqiang Wang, Jun Wan, Zhen Lei

机构 * School of Advanced Interdisciplinary Sciences (SAIS), University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Ant Group(蚂蚁集团) Shenzhen Institute of Advanced Technology (SIAT), Chinese Academy of Sciences(中国科学院深圳先进技术研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 视觉推理 :MLLM(abstract);分类 cs.CV、cs.AI

AI总结 Veritas通过模式感知推理,基于多模态大语言模型实现通用深度伪造检测,提升对未知伪造技术和数据领域的检测能力。

Comments ICLR 2026 Oral. Project: https://github.com/EricTan7/Veritas

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00043 2026-02-26 cs.CV cs.AI cs.CL 62%

VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning

VOILA:对多模态大语言模型的感知理解与类比推理能力评估

Nilay Yilmaz, Maitreya Patel, Yiran Lawrence Luo, Tejas Gokhale, Chitta Baral, Suren Jayasuriya, Yezhou Yang

机构 * Arizona State University(亚利桑那州立大学) University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 VOILA通过类比映射方法评估多模态大语言模型的感知理解和抽象推理能力,发现其在图像间关系理解上存在不足,但通过多步提示策略可提升性能。

Comments Accepted at ICLR 2025. Code and data: https://github.com/nlylmz/Voila

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01874 2026-02-25 cs.CV cs.AI 62%

CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving

CogFlow:通过知识内化连接感知与推理以解决视觉数学问题

Shuhang Chen, Yunqiu Xu, Junjie Xie, Aojun Lu, Tao Feng, Zeying Huang, Ning Zhang, Yi Sun, Yi Yang, Hangjie Yuan

机构 * Zhejiang University(浙江大学) Intelligent Learning(智能学习) Sichuan University(四川大学) Tsinghua University(清华大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 CogFlow通过知识内化连接感知与推理,提升视觉数学问题解决能力。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13742 2026-02-24 cs.CV cs.AI 62%

DL$^3$M: A Vision-to-Language Framework for Expert-Level Medical Reasoning through Deep Learning and Large Language Models

DL$^3$M: 一种通过深度学习和大语言模型实现专家级医学推理的视觉-语言框架

Md. Najib Hasan, Imran Ahmad, Sourav Basak Shuvo, Md. Mahadi Hasan Ankon, Sunanda Das, Nazmul Siddique, Hui Wang

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 DL$^3$M通过结合深度学习和大语言模型,实现专家级医学推理,但当前LLMs在高风险医疗决策中仍不可靠。

Comments This work was submitted without the consent of my current adviser. Additionally, it overlaps with my unpublished research work. In order to avoid potential academic and authorship conflicts, I am requesting withdrawal of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13444 2026-02-24 cs.CV cs.AI 62%

VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning

VideoMind: 一种用于时序 grounded 视频推理的链式 LoRA 代理

Ye Liu, Kevin Qinghong Lin, Chang Wen Chen, Mike Zheng Shou

机构 * The Hong Kong Polytechnic University(香港理工大学) National University of Singapore(新加坡国立大学)

专题命中 视觉推理 :grounding(abstract);分类 cs.CV、cs.AI

AI总结 VideoMind 提出了一种基于角色的链式 LoRA 机制,用于提升视频时序 grounded 推理的效率与灵活性。

Comments ICLR 2026 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05705 2026-02-18 cs.CV cs.AI cs.CL 62%

Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale

长 grounded 思考:大规模合成视觉问题和推理链

David Acuna, Chao-Han Huck Yang, Yuntian Deng, Jaehun Jung, Ximing Lu, Prithviraj Ammanabrolu, Hyunwoo Kim, Yuan-Hong Liao, Yejin Choi

机构 * nvidia(NVIDIA公司) uoft(多伦多大学) uwaterloo(滑铁卢大学)

专题命中 视觉推理 :VLM(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种大规模合成视觉问题和推理链的框架,通过生成高质量数据集提升多模态推理性能,验证了其在视觉、文本和音频任务中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15460 2026-02-18 cs.LG cs.CV 62%

On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks

在简单视觉规划任务中多模态大语言模型推理的分布外泛化

Yannic Neuhaus, Nicolas Flammarion, Matthias Hein, Francesco Croce

机构 * Tübingen AI Center -- University of Tübingen(图宾根人工智能中心 -- 图宾根大学) EPFL(瑞士联邦理工学院) ELLIS Institute Finland -- Aalto University(芬兰ELLIS研究所 -- 阿尔托大学)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.LG

AI总结 研究多模态大语言模型在简单视觉规划任务中推理能力的分布外泛化,发现结合多种文本格式的推理方法效果最佳,纯文本模型表现优于图像输入模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10551 2026-02-17 cs.CV cs.AI 62%

C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning

C^2ROPE: 3D 大多模态模型推理中的因果连续旋转位置编码

Guanting Ye, Qiyan Zhao, Wenhao Yu, Xiaofeng Zhang, Jianmin Ji, Yanyong Zhang, Ka-Veng Yuen

机构 * State Key Laboratory of Internet of Things for Smart City, University of Macau(物联网智能城市国家重点实验室,澳门大学) Department of Automation, Shanghai Jiaotong University(上海交通大学自动化系) Institute of Advanced Technology, University of Science and Technology of China(中国科学技术大学先进技术研究院) School of Computer Science and Technology, USTC(中国科学技术大学计算机科学与技术学院) School of Artificial Intelligence and Data Science, USTC(中国科学技术大学人工智能与数据科学学院)

专题命中 视觉推理 :visual question answering(abstract);分类 cs.CV、cs.AI

AI总结 C^2ROPE通过引入空间-时间连续位置编码和切比雪夫因果掩码,解决3D多模态模型中视觉特征连续性和因果关系建模问题。

Comments Accepted in ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14589 2026-02-17 cs.AI cs.CL cs.LG 62%

MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs

MATEO:一种多模态基准,用于LVLMs中的时间推理和规划

Gabriel Roccabruna, Olha Khomyn, Giuseppe Riccardi

机构 * Signals and Interactive Systems Lab, University of Trento, Italy(特伦托大学信号与交互系统实验室) University of Trento(特伦托大学) Amazon(亚马逊)

专题命中 视觉推理 :vision language model(abstract);分类 cs.AI、cs.LG

AI总结 MATEO是一个多模态基准,用于评估和提升大型视觉语言模型在时间推理和规划方面的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11678 2026-02-13 cs.AI cs.CV 62%

Beyond Pixels: Vector-to-Graph Transformation for Reliable Schematic Auditing

超越像素:用于可靠图示审计的向量到图转换

Chengwei Ma, Zhen Tian, Zhou Zhou, Zhixian Xu, Xiaowei Zhu, Xia Hua, Si Shi, F. Richard Yu

机构 * Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(人工智能与数字经济广东实验室) Guangdong Power Grid Co., Ltd.(广东电网公司) Shanghai University(上海大学) Carleton University(卡尔顿大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出向量到图转换方法,通过将CAD图转换为属性图,提升工程图示审计的准确性,克服基于像素方法的局限性。

Comments 4 pages, 3 figures. Accepted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11024 2026-02-12 cs.CV cs.AI 62%

Chain-of-Look Spatial Reasoning for Dense Surgical Instrument Counting

密集手术器械计数的链式观察空间推理

Rishikesh Bhyri, Brian R Quaranto, Philip J Seger, Kaity Tung, Brendan Fox, Gene Yang, Steven D. Schwaitzberg, Junsong Yuan, Nan Xi, Peter C W Kim

机构 * State University of New York at Buffalo(纽约州立大学布法罗分校)

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 本文提出Chain-of-Look框架,通过结构化视觉链提升密集手术器械计数的准确性,并引入邻近损失函数和SurgCount-HD数据集,实验证明其在复杂场景中的优越性能。

Comments Accepted to WACV 2026. This version includes additional authors who contributed during the rebuttal phase

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23050 2026-02-12 cs.LG cs.AI 62%

Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding

通过对比嵌入链理解LVLMs的语言先验

Lin Long, Changdae Oh, Seongheon Park, Sharon Li

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.AI、cs.LG

AI总结 通过对比嵌入链分析,揭示LVLMs中视觉信息整合的关键层及影响响应生成的强度量化方法。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07689 2026-02-10 cs.CV cs.AI 62%

Process-of-Thought Reasoning for Videos

视频过程推理

Jusheng Zhang, Kaitong Cai, Jian Wang, Yongsen Zheng, Kwok-Yan Lam, Keze Wang

机构 * Sun Yat-sen University, China(中山大学) Nanyang Technological University, Singapore(南洋理工大学) Snap Inc.(Snap公司)

专题命中 视觉推理 :grounding(abstract);分类 cs.CV、cs.AI

AI总结 视频过程推理框架通过结构化推理步骤提升视频理解的准确性与可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07625 2026-02-10 cs.CV cs.AI 62%

AD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured Reasoning

AD-MIR:通过结构化推理弥合从感知到说服的广告视频理解鸿沟

Binxiao Xu, Junyu Feng, Xiaopeng Lin, Haodong Li, Zhiyuan Feng, Bohan Zeng, Shaolin Lu, Ming Lu, Qi She, Wentao Zhang

机构 * Peking University, Beijing, China(北京大学) Tsinghua University, Beijing, China(清华大学) Xi'an Jiaotong University, Xi'an, China(西安交通大学) South China University of Technology, Guangzhou, China(华南理工大学)

专题命中 视觉推理 :grounding(abstract);分类 cs.CV、cs.AI

AI总结 AD-MIR通过结构化推理框架,结合语义检索与精确关键词匹配,有效解码广告意图,提升广告视频理解的准确性与说服策略分析能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07061 2026-02-10 cs.LG cs.AI 62%

TACIT: Transformation-Aware Capturing of Implicit Thought

TACIT:面向隐式思维的变换感知捕捉

Daniel Nobrega

机构 * Independent Researcher(独立研究者)

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.AI、cs.LG

AI总结 TACIT提出一种基于扩散的Transformer模型,通过像素空间中的校正流实现可解释的视觉推理,在迷宫求解中展示了隐式思维的同步涌现特性。

Comments 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07054 2026-02-10 cs.LG cs.CV cs.HC 62%

AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization

AVERE: 通过偏好优化提升音频视觉情感推理

Ashutosh Chaubey, Jiacheng Pang, Maksim Siniukov, Mohammad Soleymani

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.LG

AI总结 AVERE通过偏好优化技术提升音频视觉情感推理性能,解决情感与无关线索的虚假关联和幻觉问题,提升多模态模型在情感理解中的表现。

Comments Accepted as a conference paper at ICLR 2026. Project page: https://avere-iclr.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08585 2026-02-09 cs.AI cs.CV 62%

Simulating the Visual World with Artificial Intelligence: A Roadmap

用人工智能模拟视觉世界:一条路线图

Jingtong Yue, Ziqi Huang, Zhaoxi Chen, Xintao Wang, Pengfei Wan, Ziwei Liu

机构 * Robotics Institute Carnegie Mellon University(卡内基梅隆大学机器人研究所) S-Lab Nanyang Technological University(南洋理工大学S实验室) Kling Team Kuaishou Technology(快手科技 Kling 团队)

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种基于隐式世界模型和视频渲染器的视频基础模型,用于模拟物理动态、智能体交互和任务规划,涵盖四代视频生成技术,应用于机器人、自动驾驶和互动游戏等领域。

Comments Project page: https://world-model-roadmap.github.io/ Github Repo: https://github.com/ziqihuangg/Awesome-From-Video-Generation-to-World-Model

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04144 2026-02-05 cs.AI cs.LG 62%

OMG-Agent: Toward Robust Missing Modality Generation with Decoupled Coarse-to-Fine Agentic Workflows

OMG-Agent:迈向鲁棒缺失模态生成的解耦粗到细代理工作流

Ruiting Dai, Zheyu Wang, Haoyu Yang, Yihan Liu, Chengzhi Wang, Zekun Zhang, Zishan Huang, Jiaman Cen, Lisi Mo

机构 * University of Electronic Science and Technology of China(电子科技大学)

专题命中 视觉推理 :MLLM(abstract);分类 cs.AI、cs.LG

AI总结 OMG-Agent通过解耦粗到细的代理工作流,有效解决多模态数据缺失问题,提升生成的鲁棒性和保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03414 2026-02-04 cs.CV cs.AI 62%

Socratic-Geo: Synthetic Data Generation and Geometric Reasoning via Multi-Agent Interaction

Socratic-Geo:通过多智能体交互实现合成数据生成与几何推理

Zhengbo Jiao, Shaobo Wang, Zifan Zhang, Wei Wang, Bing Zhao, Hu Wei, Linfeng Zhang

机构 * AI DATA, Alibaba Group Holding Limited(阿里数据,阿里巴巴集团控股有限公司) EPIC Lab, Shanghai Jiao Tong University(上海交通大学EPIC实验室) Shanghai University of Finance(上海财经大学) Wuhan University(武汉大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 Socratic-Geo通过多智能体交互实现合成数据生成与几何推理,利用教师代理和求解代理动态耦合数据合成与模型学习,提升图像生成和推理能力。

Comments 18pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17873 2026-02-04 cs.CV cs.AI 62%

SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model

SurgVidLM:迈向多粒度外科视频理解的大型语言模型

Guankun Wang, Junyi Wang, Wenjin Mo, Long Bai, Kun Yuan, Ming Hu, Jinlin Wu, Junjun He, Yiming Huang, Nicolas Padoy, Zhen Lei, Hongbin Liu, Nassir Navab, Hongliang Ren

机构 * The Chinese University of Hong Kong(香港中文大学) Sun Yat-sen University(中山大学) University of Strasbourg(斯特拉斯堡大学) Technical University of Munich(慕尼黑技术大学) Monash University(墨尔本大学) Centre for Artificial Intelligence and Robotics, HKISI-CAS(人工智能与机器人中心,HKISI-CAS) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 SurgVidLM通过多粒度分析提升外科视频理解能力,结合全局与局部机制实现更精确的手术流程解析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01884 2026-02-03 cs.AI cs.LG 62%

Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models

熵引导的数据高效训练用于多模态推理奖励模型

Shidong Yang, Tongwen Huang, Hao Wen, Yong Wang, Li Chen, Xiangxiang Chu

机构 * School of Software, Tsinghua University(清华大学软件学院) AMAP, Alibaba Group(阿里巴巴集团AMAP)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出熵引导训练方法,通过熵指导数据筛选和训练策略提升多模态推理奖励模型的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01477 2026-02-03 cs.LG cs.AI 62%

Achieving Time Series Reasoning Requires Rethinking Model Design, Tasks Formulation, and Evaluation

实现时间序列推理需要重新思考模型设计、任务制定和评估

Yaxuan Kong, Yiyuan Yang, Shiyu Wang, Chenghao Liu, Yuxuan Liang, Ming Jin, Stefan Zohren, Dan Pei, Yan Liu, Qingsong Wen

机构 * University of Oxford, UK(牛津大学) Hong Kong University of Science(香港科学大学) Griffith University, Australia(格里菲斯大学) Tsinghua University, China(清华大学) University of Southern California, USA(南加州大学)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI、cs.LG

AI总结 本文指出时间序列推理需重新审视模型设计、任务制定和评估,提出统一框架以提升现实应用中的鲁棒性、可解释性和决策相关性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16648 2026-02-02 cs.AI cs.CL cs.LG 62%

FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs

FESTA:用于多模态大语言模型信任评估的功能等效采样

Debarpan Bhattacharya, Apoorva Kulkarni, Sriram Ganapathy

机构 * Indian Institute of Science(印度科学研究院) University of Maryland College Park(马里兰大学学院公园分校)

专题命中 视觉推理 :multimodal large language model(abstract);分类 cs.AI、cs.LG

AI总结 FESTA通过功能等效采样技术提升多模态大语言模型的预测选择性性能,实现33.3%和29.6%的改进。

Comments Accepted in the Findings of EMNLP, 2025

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11558 2026-01-29 cs.CV cs.AI cs.CL 62%

DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs

DaMO:一种数据高效的多模态协调器,用于视频LLM的时序推理

Bo-Cheng Chiu, Jen-Jee Chen, Yu-Chee Tseng, Feng-Chi Chen, An-Zi Yen

机构 * College of Artificial Intelligence, National Yang Ming Chiao Tung University(人工智能学院,阳明交通大学) Institute of Population Health Sciences, National Health Research Institutes(人口健康科学研究所,国家健康研究院) Department of Computer Science, National Yang Ming Chiao Tung University(计算机科学系,阳明交通大学)

专题命中 视觉推理 :grounding(abstract);分类 cs.CV、cs.AI

AI总结 DaMO是一种专为视频LLM设计的数据高效多模态协调器,通过时序感知Fuseformer和四阶段训练范式提升时序推理能力,实现更精确的多模态理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20303 2026-01-29 cs.CV cs.AI 62%

Physically Guided Visual Mass Estimation from a Single RGB Image

基于单个RGB图像的物理引导视觉质量估计

Sungjae Lee, Junhan Jeong, Yeonjoo Hong, Kwang In Kim

机构 * Graduate School of Artificial Intelligence, POSTECH, South Korea(人工智能研究生院,POSTECH,韩国) Department of Electrical Engineering, POSTECH, South Korea(电气工程系,POSTECH,韩国)

专题命中 视觉推理 :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种基于物理引导的单图像质量估计方法,通过融合几何、语义和外观信息,实现对物体质量的准确预测。

详情

展开后加载摘要…

URL PDF HTML 收藏