arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2602.20731 2026-02-25 cs.CV cs.AI cs.LG 62%

Communication-Inspired Tokenization for Structured Image Representations

受通信启发的结构化图像表示分词

Aram Davtyan, Yusuf Sahin, Yasaman Haghighi, Sebastian Stapf, Pablo Acuaviva, Alexandre Alahi, Paolo Favaro

机构 * Computer Vision Group, University of Bern, Switzerland(伯恩大学计算机视觉组) VITA Lab, EPFL, Switzerland(苏黎世联邦理工学院VITA实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 COMiT通过受通信启发的分词框架,学习结构化离散视觉分词序列,提升图像重建和语义理解能力。

Comments Project website: https://araachie.github.io/comit/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13647 2026-02-25 cs.RO cs.AI cs.CV 62%

An Efficient LiDAR-Camera Fusion Network for Multi-Class 3D Dynamic Object Detection and Trajectory Prediction

一种高效的激光雷达-摄像头融合网络用于多类3D动态物体检测和轨迹预测

Yushen He, Lei Zhao, Tianchen Deng, Zipeng Fang, Weidong Chen

机构 * Institute of Medical Robotics and Department of Automation, Shanghai Jiao Tong University, Key Laboratory of System Control and Information Processing, Ministry of Education(医学机器人研究所和自动化系,上海交通大学,系统控制与信息处理重点实验室,教育部)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种高效的激光雷达-摄像头融合网络,用于多类3D动态物体检测与轨迹预测,通过UniMT和RTMCT模型实现高精度检测与多样化轨迹预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19367 2026-02-24 cs.AI cs.CV 62%

Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation Spaces

时间序列、视觉与语言:探讨对比表示空间中对齐的极限

Pratham Yashwante, Rose Yu

机构 * University of California San Diego, USA(加州大学圣地亚哥分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 研究探讨了时间序列、视觉和语言在对比表示空间中的对齐问题,发现模型规模越大对齐性越强,但对齐是不对称的,图像可作为中介,文本和视觉信息的密度影响对齐效果。

Comments 24 Figures, 12 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05992 2026-02-24 cs.CV cs.AI 62%

Exploring Partial Multi-Label Learning via Integrating Semantic Co-occurrence Knowledge

探索通过整合语义共现知识的半多标签学习

Xin Wu, Fei Teng, Yue Feng, Kaibo Shi, Zhuosheng Lin, Ji Zhang, James Wang

机构 * School of Computing and Artificial Intelligence, Southwest Jiaotong University(计算机与人工智能学院,西南交通大学) Engineering Research Center of Sustainable Urban Intelligent Transportation, Ministry of Education(可持续城市智能交通工程研究中心,教育部) School of Engineering, Swinburne University of Technology(工程学院,斯winburne大学) School of Electronic and Information Engineering, Wuyi University(电子与信息工程学院,五邑大学) College of Electrical Engineering, Sichuan University(电气工程学院,四川大学) School of Computer Science, Chengdu University(计算机科学学院,成都大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出SCINet框架,通过整合语义共现知识,提升半多标签学习的准确性与效果。

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03407 2026-02-17 cs.GR cs.AI cs.CV cs.LG 62%

Multi-Spectral Gaussian Splatting with Neural Color Representation

多光谱高斯点云渲染与神经颜色表示

Lukas Meyer, Josef Grün, Maximilian Weiherer, Bernhard Egger, Marc Stamminger, Linus Franke

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出MS-Splatting框架,通过神经颜色表示实现多光谱3D高斯点云渲染,提升多光谱和单光谱渲染质量,应用于植被指数渲染。

Comments for project page, see https://meyerls.github.io/ms_splatting

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09651 2026-02-12 cs.CV cs.AI cs.LG 62%

Geospatial Representation Learning: A Survey from Deep Learning to The LLM Era

地理空间表示学习:从深度学习到大语言模型时代的一次综述

Xixuan Hao, Yutian Jiang, Xingchen Zou, Jiabo Liu, Yifang Yin, Song Gao, Flora Salim, Tianrui Li, Yuxuan Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) The University of New South Wales(新南威尔士大学) Southwest Jiaotong University(西南交通大学) Institute for Infocomm Research (I$^2$R), A*STAR(信息通信研究院(I$^2$R),A*STAR)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文综述了从深度学习到大语言模型时代的地理空间表示学习,探讨了其方法、应用及未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06973 2026-02-10 cs.CL cs.AI cs.LG 62%

Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models

视觉渲染能否绕过分词?探究基于像素的语言模型中脚本-分词器不一致问题

Lucky Susanto, Musa Izzanardi Wijanarko, Khumaisa Nur'aini, Farid Adilazuarda, Alham Fikri Aji, Derry Tanti Wijaya

机构 * Monash University Indonesia(墨尔本大学印尼分校) MBZUAI Boston University(波士顿大学) University of Edinburgh(爱丁堡大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本研究探讨了基于像素的语言模型中视觉渲染是否能绕过分词约束,发现重新引入文本分词器加剧了分词不一致问题,自定义分词器在性能上表现更优。

Comments Submitted to ARR January

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04904 2026-02-06 cs.LG cs.AI cs.MM eess.IV 62%

DCER: Dual-Stage Compression and Energy-Based Reconstruction

DCER:双阶段压缩与基于能量的重建

Yiwen Wang, Jiahao Qin

机构 * Yiwen Wang(无) Jiahao Qin(无)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI、cs.MM

AI总结 DCER通过双阶段压缩和基于能量的重建解决多模态融合中的噪声和缺失模态问题,实现鲁棒性提升。

Comments 13 pages, 2 figures, 8 tables. Submitted to ICML 2026. Code will be available on GitHub

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18533 2026-02-04 cs.CV cs.CL cs.CR 62%

Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models

重新审视视觉语言模型安全微调中的瓶颈

Yi Ding, Lijun Li, Bing Cao, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Purdue University(普渡大学) Tianjin University(天津大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出多图像安全数据集,通过增强视觉推理能力,提升模型在安全关键任务中的性能与通用能力。

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18179 2026-02-04 cs.CL cs.AI 62%

Problem Solved? Information Extraction Design Space for Layout-Rich Documents using LLMs

问题已解决?利用LLMs处理布局丰富文档的信息提取设计空间

Gaye Colakoglu, Gürkan Solmaz, Jonathan Fürst

机构 * Zurich University of Applied Sciences(苏黎世应用科学大学) NEC Laboratories Europe(NEC欧洲实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文通过LayIE-LLM测试套件研究了利用LLMs处理布局丰富文档的信息提取设计空间,证明通用LLMs在优化配置下可媲美专用模型,提供低成本无微调方案。

Comments accepted at EMNLP'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09125 2026-02-04 cs.CV cs.AI 62%

HAAP: Vision-context Hierarchical Attention Autoregressive with Adaptive Permutation for Scene Text Recognition

HAAP: 基于自适应排列的视觉-上下文分层注意力自回归模型

Honghui Chen, Yuhang Qiu, Jiabao Wang, Pingping Chen, Nam Ling

机构 * College of Physics and Information Engineering, Fuzhou University(福州大学物理与信息工程学院) Faculty of Engineering, Monash University(莫纳什大学工程学院) Department of Computer Science and Engineering, Santa Clara University(圣克拉拉大学计算机科学与工程系)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 HAAP通过自适应排列和跨模态分层注意力机制,提升场景文本识别的准确性和效率。

Comments 12 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07593 2026-02-03 cs.RO cs.AI cs.CV cs.SY eess.IV eess.SY 62%

Vision-Proprioception Fusion with Mamba2 in End-to-End Reinforcement Learning for Motion Control

基于Mamba2的视觉-本体感知融合在端到端强化学习中的运动控制

Xiaowen Tao, Yinuo Wang, Jinzhao Zhou

机构 * School of Computer Science and Statistics, Trinity College Dublin(都柏林三一学院计算机科学与统计学系) Faculty of Engineering and Information Technology, University of Technology Sydney(新南威尔士大学理工学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出基于SSD-Mamba2的视觉-本体感知融合框架,通过端到端强化学习提升运动控制的效率和安全性。

Comments 6 figures and 8 tables. This paper has been accepted by Advanced Engineering Informatics

Journal ref Advanced Engineering Informatics, vol. 71, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19112 2026-01-28 cs.AI cs.MM cs.SD 62%

Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation

具有情绪先验蒸馏的不确定性感知3D情感说话面孔合成

Nanhan Shen, Zhilei Liu

机构 * School of Artificial Intelligence, Tianjin University, Tianjin, China(人工智能学院,天津大学,天津,中国)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI、cs.MM

AI总结 UA-3DTalk通过引入情绪先验蒸馏和不确定性感知模块,提升3D情感说话面孔合成的音频-视觉对齐和渲染质量。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16788 2026-01-26 cs.CV cs.AI 62%

REL-SF4PASS: Panoramic Semantic Segmentation with REL Depth Representation and Spherical Fusion

REL-SF4PASS:基于REL深度表示和球形融合的全景语义分割

Xuewei Li, Xinghan Bao, Zhimin Chen, Xi Li

机构 * School of Electronic and Information Engineering, Shanghai DianJi University(电子信息学院,上海电机大学) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 REL-SF4PASS通过REL深度表示和球形动态多模态融合方法,提升全景语义分割的性能和鲁棒性。

Comments submitted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02368 2026-01-26 cs.CV cs.AI 62%

MoE-Enhanced Multi-Domain Feature Selection and Fusion for Fast Map-Free Trajectory Prediction

增强型多域特征选择与融合的无地图轨迹预测

Wenyi Xiong, Jian Chen, Ziheng Qi, Wenhua Chen

机构 * School of Mechanical Engineering, Zhejiang University(浙江大学机械工程学院) Guangdong Provincial Key Laboratory of Fully Actuated System Control Theory and Technology, School of Automation and Intelligent Manufacturing, Southern University of Science and Technology(南方科技大学自动化与智能制造学院) Leapmotor Technology(Leapmotor科技) Department of Aeronautical and Automotive Engineering, Loughborough University(伦敦大学洛桑学院航空与汽车工程系)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种无地图轨迹预测方法,通过多域特征选择与融合,提升轨迹预测的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06504 2026-01-21 cs.CV cs.AI cs.RO 62%

Method of UAV Inspection of Photovoltaic Modules Using Thermal and RGB Data Fusion

基于热成像与RGB数据融合的光伏板无人机检测方法

Andrii Lysyi, Anatoliy Sachenko, Pavlo Radiuk, Mykola Lysyi, Oleksandr Melnychenko, Diana Zahorodnia

机构 * Khmelnytskyi National University(赫梅利尼茨基国立大学) Department of Informatics and Teleinformatics(信息学与电信学系) Kazimierz Pulaski University of Ra-dom(卡齐米日·波尼亚图夫斯基大学) National Academy of the State Border Service of Ukraine named after Bogdan Khmelnitsky(以博丹·赫梅利尼茨基命名的乌克兰国家边境服务学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出了一种基于热成像与RGB数据融合的无人机光伏板检测方法,通过多模态系统实现自动化检测,提升电站安全性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11311 2026-01-15 eess.IV cs.AI cs.CV cs.LG 62%

Large-scale modality-invariant foundation models for brain MRI analysis: Application to lesion segmentation

大规模模态不变基础模型用于脑MRI分析:应用于病变分割

Petros Koutsouvelis, Matej Gazda, Leroy Volmer, Sina Amirrajab, Kamil Barbierik, Branislav Setlak, Jakub Gazda, Peter Drotar

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种大规模模态不变基础模型,用于提升脑MRI中病变分割的性能,通过自监督学习预训练并保留细粒度模态特定特征。

Comments Submitted to IEEE ISBI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08811 2026-01-14 cs.CV cs.AI 62%

Reasoning Matters for 3D Visual Grounding

推理在3D视觉定位中至关重要

Hsiang-Wei Huang, Kuang-Ming Chen, Wenhao Chai, Cheng-Yen Yang, Jen-Hao Cheng, Jenq-Neng Hwang

机构 * University of Washington(华盛顿大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种自动合成3D视觉定位数据的管道,并引入了在仅使用1.6%训练数据下表现优于现有方法的Reason3DVG-8B模型,证明了推理在3D视觉定位中的重要性。

Comments 2025 CVPR Workshop on 3D-LLM/VLA: Bridging Language, Vision and Action in 3D Environments

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08684 2026-01-14 cs.AI cs.CV 62%

MEMEWEAVER: Inter-Meme Graph Reasoning for Sexism and Misogyny Detection

MEMEWEAVER:跨迷因图推理用于性别歧视和性别歧视检测

Paolo Italiani, David Gimeno-Gomez, Luca Ragazzi, Gianluca Moro, Paolo Rosso

机构 * Department of Computer Science and Engineering, University of Bologna(博洛尼亚大学计算机科学与工程系) PRHLT Research Center, Universtitat Politècnica de València(巴塞罗那理工大学PRHLT研究中心) ValgrAI - Valencian Graduate School and Research Network of Artificial Intelligence, Spain(西班牙瓦伦西亚人工智能研究生院与研究网络)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MemeWeaver通过跨迷因图推理机制,有效检测性别歧视和性别歧视,优于现有基线方法。

Comments Accepted at EACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07871 2026-01-14 q-bio.QM cs.AI cs.CV cs.LG 62%

Imaging-anchored Multiomics in Cardiovascular Disease: Integrating Cardiac Imaging, Bulk, Single-cell, and Spatial Transcriptomics

心血管疾病中的成像锚定多组学:整合心脏成像、批量、单细胞和空间转录组学

Minh H. N. Le, Tuan Vinh, Thanh-Huy Nguyen, Tao Li, Bao Quang Gia Le, Han H. Huynh, Monika Raj, Carl Yang, Min Xu, Nguyen Quoc Khanh Le

机构 * International Ph.D. Program in Medicine, College of Medicine, Taipei Medical University, Taipei, Taiwan AIBioMed Research Group, Taipei Medical University, Taipei, Taiwan Medical Sciences Division, University of Oxford, Oxford, United Kingdom Computational Biology Department, School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA Department of Computer Science, Emory University, Atlanta, GA, USA Department of Chemistry, Emory University, Atlanta, GA, USA International Master Program for Translational Science, College of Medical Science Technology, Taipei Medical University, Taipei 110, Taiwan In-Service Master Program in Artificial Intelligence in Medicine, College of Medicine, Taipei Medical University, Taipei, Taiwan Translational Imaging Research Center, Taipei Medical University Hospital, Taipei, Taiwan

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出通过整合心脏成像与多组学数据,推动心血管疾病研究的多模态融合方法,提升疾病诊断和治疗的精准性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08667 2026-01-13 cs.CV cs.MM 62%

Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization

通过分层双阶段优化学习通用且高效的图像水印技术

Ke Liu, Xuanhan Wang, Qilong Zhang, Lianli Gao, Jingkuan Song

机构 * Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(深圳先进研究院,电子科学与技术大学) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.MM

AI总结 本文提出分层双阶段优化框架,通过学习通用且高效的图像水印技术,提升水印提取准确率并保持低延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06835 2026-01-13 cs.CV cs.AI 62%

OSCAR: Optical-aware Semantic Control for Aleatoric Refinement in Sar-to-Optical Translation

OSCAR:面向随机细化的语义控制光学感知SAR到光学翻译

Hyunseo Lee, Sang Min Kim, Ho Kyung Shin, Taeheon Kim, Woo-Jeoung Nam

机构 * Kyungpook National University(庆北国立大学) Korea Aerospace Research Institute(韩国航空航天研究机构)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 OSCAR通过跨模态语义对齐、语义引导生成指导和不确定性感知目标,提升SAR到光学图像翻译的语义一致性和感知质量。

Comments main 15 pages, supplementary 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05470 2026-01-12 cs.CV cs.CL 62%

ROAP: A Reading-Order and Attention-Prior Pipeline for Optimizing Layout Transformers in Key Information Extraction

ROAP:一种基于阅读顺序和注意力优先的流水线,用于优化布局变换器在关键信息提取中的应用

Tingwei Xie, Jinxin He, Yonghong Song

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 ROAP通过优化布局变换器的注意力分布,提升关键信息提取的性能,解决阅读顺序建模和视觉干扰问题。

Comments 10 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16942 2026-01-08 cs.SI cs.AI cs.CV 62%

S2Vec: Self-Supervised Geospatial Embeddings for the Built Environment

S2Vec: 自监督的建成环境地理嵌入

Shushman Choudhury, Elad Aharoni, Chandrakumari Suvarna, Iveel Tsogsuren, Abdul Rahman Kreidieh, Chun-Ta Lu, Neha Arora

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 S2Vec通过自监督学习生成通用的地理空间嵌入,适用于建成环境特征的表示,并在社会经济任务中表现优异,同时支持多模态融合提升性能。

Journal ref ACM Transactions on Spatial Algorithms and Systems 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03490 2026-01-08 cs.CV cs.AI 62%

CroBIM-U: Uncertainty-Driven Referring Remote Sensing Image Segmentation

CroBIM-U: 基于不确定性的遥感图像分割

Yuzhe Sun, Zhe Dong, Haochen Jiang, Tianzhu Liu, Yanfeng Gu

机构 * School of Electronics and Information Engineering, Harbin Institute of Technology(电子与信息工程学院,哈尔滨工业大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 CroBIM-U通过不确定性引导框架提升遥感图像分割的鲁棒性和几何精度,采用不确定性图和即插即用模块实现自适应推理与局部细化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03460 2026-01-08 cs.CV cs.AI 62%

FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder

FROST-Drive: 可扩展且高效的端到端驾驶方法,采用冻结的视觉编码器

Zeyu Dong, Yimin Zhu, Yu Wu, Yu Sun

机构 * Stony Brook University(石溪大学) Rutgers University(罗格斯大学) Sunrise Technology Inc.(Sunrise技术公司)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 FROST-Drive通过冻结预训练视觉编码器,结合适配器和解码器,实现高效端到端驾驶,优于全微调方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24826 2026-01-01 cs.CV cs.AI 62%

Video and Language Alignment in 2D Systems for 3D Multi-object Scenes with Multi-Information Derivative-Free Control

用于多物体3D场景的2D系统中视频与语言对齐的多信息无导数控制

Jason Armitage, Rico Sennnrich

机构 * University of Zurich(苏黎世大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种无导数优化方法,用于在多物体3D场景中实现视频与语言的对齐,通过在线适应物体遮挡和区分特征来提升跨模态任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18262 2025-12-30 cs.RO cs.AI cs.CV cs.HC cs.LG 62%

ReSemAct: Advancing Fine-Grained Robotic Manipulation via Semantic Structuring and Affordance Refinement

ReSemAct:通过语义结构化和效用细化推进细粒度机器人操作

Chenyu Su, Weiwei Shang, Chen Qian, Fei Zhang, Shuang Cong

机构 * Department of Automation, University of Science and Technology of China(自动化系,中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 ReSemAct 通过语义结构化和效用细化方法,在细粒度机器人操作中实现更精确的效用目标生成与动态环境适应。

Comments Code and videos: https://github.com/scy-v/ReSemAct and https://resemact.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22217 2025-12-30 cs.CV cs.AI 62%

VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition

VLM-PAR:一种用于行人属性识别的视觉语言模型

Abdellah Zakaria Sellam, Salah Eddine Bekhouche, Fadi Dornaika, Cosimo Distante, Abdenour Hadid

机构 * Department of Innovation Engineering(创新工程系) University of Salento, Italy(意大利萨伦托大学) Institute of Applied Sciences and Intelligent Systems - CNR(应用科学与智能系统研究所 - CNR) University of the Basque Country UPV/EHU(巴斯克国家大学UPV/EHU) IKERBASQUE, Basque Foundation for Science(伊基塔斯克巴塞克基金会) Sorbonne University Abu Dhabi(索邦大学阿布扎比分校)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 VLM-PAR通过整合大规模视觉语言预训练与跨模态细化,提升行人属性识别在类别不平衡和泛化挑战中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22188 2025-12-30 cs.CV cs.AI 62%

HookMIL: Revisiting Context Modeling in Multiple Instance Learning for Computational Pathology

HookMIL: 重新审视多实例学习中的上下文建模以用于计算病理学

Xitong Ling, Minxi Ouyang, Xiaoxiao Li, Jiawen Li, Ying Chen, Yuxuan Sun, Xinrui Chen, Tian Guan, Xiaoping Liu, Yonghong He

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) School of Informatics, Xiamen University(厦门大学信息学院) School of Engineering, Westlake University(西湖大学工程学院) Zhongnan Hospital, Wuhan University(武汉大学中南医院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 HookMIL通过引入可学习的钩子标记和多样性损失,提升多实例学习在计算病理学中的效率与可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏