arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-30 至 2026-01-30 共收录 75 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 3 篇

2510.21296 2026-01-30 cs.LG 67%

An Evidence-Based Post-Hoc Adjustment Framework for Anomaly Detection Under Data Contamination

基于证据的后验调整框架用于数据污染下的异常检测

Sukanya Patra, Souhaib Ben Taieb

机构 * Department of Computer Science, University of Mons(蒙斯大学计算机科学系)

专题命中 图文多模态 :multimodal(abstract);multimodal foundation model(abstract)

AI总结 EPHAD提出一种基于证据的后验调整框架,通过测试时收集的证据提升异常检测在数据污染下的性能。

Comments Accepted in the Thirty-ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21220 2026-01-30 cs.CV 57%

LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models

LAMP: 通过预训练模型学习通用对抗扰动以实现多图像任务

Alvi Md Ishmam, Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Chris Thomas

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 LAMP通过预训练模型学习通用对抗扰动,针对多图像多模态大语言模型实现高效的黑盒攻击,提升了多任务攻击成功率。

Comments Accepted in main technical track AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21794 2026-01-30 cs.LG 50%

Knowledge Vector Weakening: Efficient Training-free Unlearning for Large Vision-Language Models

知识向量削弱:高效无训练卸载方法用于大视觉-语言模型

Yejin Kim, Dongjun Hwang, Sungmin Cha, Junsuk Choe

机构 * Sogang University(ソガン大学) New York University(纽约大学)

专题命中 图文多模态 :multimodal(abstract)

AI总结 KVW提出了一种无需训练的高效卸载方法,通过削弱模型中被激活的知识向量,有效防止模型利用有害知识,提升计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2601.21181 2026-01-30 cs.AI 89%

MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models

MAD:用于减轻多模态大语言模型中跨模态幻觉的模态自适应解码

Sangyun Chung, Se Yeon Kim, Youngchae Chee, Yong Man Ro

机构 * Integrated Vision Language Lab, KAIST, South Korea(韩国科学技术院集成视觉语言实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(title,abstract);audio-visual(abstract);分类 cs.AI

AI总结 MAD通过自适应加权对比解码分支,有效减少多模态大语言模型中的跨模态幻觉问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21740 2026-01-30 cs.MM cs.SD 83%

MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding

MIDI-LLaMA:一种用于符号音乐理解的指令遵循多模态大语言模型

Meng Yang, Jon McCormack, Maria Teresa Llano, Wanchao Su, Chao Lei

机构 * SensiLab, Monash University, Australia(Monash大学) University of Sussex, Brighton, United Kingdom(Sussex大学) School of Computing and Information Systems, The University of Melbourne, Australia(墨尔本大学计算机与信息系统学院)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.MM

AI总结 MIDI-LLaMA通过结合MusicBERT和Llama-3-8B,实现了对符号音乐的指令遵循理解,显著提升了音乐描述和语义对齐能力。

Comments Accepted for publication at International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13576 2026-01-30 eess.AS cs.AI cs.SD eess.IV 81%

End-to-end audio-visual learning for cochlear implant sound coding simulations in noisy environments

端到端音频视觉学习用于在噪声环境中的人工耳蜗声音编码模拟

Meng-Ping Lin, Enoch Hsin-Ho Huang, Shao-Yi Chien, Yu Tsao

机构 * Graduate Institute of Electronics Engineering(电子工程研究所) the Department of Electrical Engineering, National Taiwan University, Taipei 106319, Taiwan(国立台湾大学电子工程系) Research Center for Information Technology Innovation Technology, Academia Sinica, Taipei 115201, Taiwan(中科院资讯科技创新研究中心) Department of Electrical Engineering, Chung Yuan Christian University, Taoyuan 320314, Taiwan(Chung Yuan Christian University 电子工程系)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI、eess.AS

AI总结 本文提出了一种端到端的人工耳蜗声音编码系统,通过整合音频视觉信息提升噪声环境中的语音可懂度和信噪比。

Comments 7 pages, 2 figures

Journal ref JASA Express Lett. 6 (2026) 015202

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12142 2026-01-30 eess.AS cs.MM cs.RO 62%

Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving

Listen, Look, Drive: 通过用户意识的VLA基于自主驾驶的音频指令耦合

Ziang Guo, Feng Yang, Xuefeng Zhang, Jiaqi Guo, Kun Zhao, Yixiao Zhou, Peng Lu, Sifa Zheng, Zufeng Zhang

机构 * SuZhou Automotive Research Institute of Tsinghua University(清华大学苏州汽车研究院) Department of Electrical and Electronic Engineering, The University of Hong Kong(香港大学电子与电气工程系) Hyundai Motor Advanced Tech. R&D Center School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS

AI总结 EchoVLA通过结合音频指令与视觉信息,提升自动驾驶对用户意图和情绪的感知能力,显著降低误差和碰撞率。

Comments Accepted by IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13244 2026-01-30 cs.SD cs.AI cs.MM 62%

MotionBeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding

MotionBeat: 通过具身体验对比学习和小节等价接触感知编码实现运动对齐的音乐表示

Xuanchen Wang, Heng Wang, Weidong Cai

机构 * School of Computer Science, The University of Sydney, Australia(计算机科学学院,悉尼大学)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、cs.MM

AI总结 MotionBeat通过具身体验对比学习和小节等价接触感知编码,实现运动对齐的音乐表示学习,提升音乐与舞蹈生成及多任务应用性能。

Comments 5 pages, 1 figure, accepted by ICASSP 2026. demo page: https://motionbeat2025.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20230 2026-01-30 cs.CL cs.HC 57%

Unit-Based Agent for Semi-Cascaded Full-Duplex Dialogue Systems

基于单元的代理用于半级联全双工对话系统

Haoyuan Yu, Yuxuan Chen, Minjie Cai

机构 * Hunan University(湖南大学) Gongdao Technology(公道科技) Jilin University(吉林大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

AI总结 本文提出基于单元的半级联全双工对话系统,利用多模态大语言模型和辅助模块实现高效对话处理,实验显示其在挑战赛中表现优异。

Comments ICASSP 2026 (Grant Challenge). https://github.com/yu-haoyuan/fd-badcat

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01559 2026-01-30 cs.SD 50%

LLM2Fx-Tools: Tool Calling For Music Post-Production

LLM2Fx-Tools: 音乐后期制作中的工具调用

Seungheon Doh, Junghyun Koo, Marco A. Martínez-Ramírez, Woosung Choi, Wei-Hsiang Liao, Qiyu Wu, Juhan Nam, Yuki Mitsufuji

机构 * KAIST(韩国科学技术院) Sony AI(索尼人工智能) Sony Group Corporation(索尼集团)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 LLM2Fx-Tools通过LLM实现音频效果模块的工具调用,生成可执行的音频效果序列,提升音乐后期制作的可解释性和可控性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 7 篇

2512.10408 2026-01-30 cs.CV 83%

MultiHateLoc: Towards Temporal Localisation of Multimodal Hate Content in Online Videos

MultiHateLoc:迈向在线视频中多模态仇恨内容的时间定位

Qiyue Sun, Tailin Chen, Yinghui Zhang, Yuchen Zhang, Jiangbei Yue, Jianbo Jiao, Zeyu Fu

机构 * Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系) Institute for Analytics and Data Science, University of Essex(埃塞克斯大学分析与数据科学研究所) School of Computer Science, University of Birmingham(伯明翰大学计算机科学学院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 MultiHateLoc提出了一种弱监督多模态仇恨内容时间定位框架,通过动态跨模态融合和模态感知MIL目标实现细粒度帧级预测。

Comments In Proceedings of the ACM Web Conference 2026 (WWW 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19139 2026-01-30 cs.LG cs.DC cs.ET 82%

Native LLM and MLLM Inference at Scale on Apple Silicon

在Apple Silicon上大规模原生LLM和MLLM推理

Wayner Barrios

机构 * Wiqonn Technologies(Wiqonn技术公司)

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract)

AI总结 vllm-mlx在Apple Silicon上实现高效LLM和MLLM推理,通过原生优化提升文本模型吞吐量并减少多模态延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22039 2026-01-30 cs.CV 79%

Understanding Multimodal Complementarity for Single-Frame Action Anticipation

理解单帧动作预测中的多模态互补性

Manuel Benavent-Lledo, Konstantinos Bacharidis, Konstantinos Papoutsakis, Antonis Argyros, Jose Garcia-Rodriguez

机构 * Department of Computer Technology, University of Alicante(阿尔瓦雷斯大学计算机技术系) Institute of Computer Science, FORTH(福蒂研究所) Computer Science Department, University of Crete(克里特大学计算机科学系) Department of Management, Science and Technology, Hellenic Mediterranean University(希腊地中海大学管理、科学与技术系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究通过单帧动作预测框架AAG+,探索多模态互补性对动作预测的影响,验证单帧信息在动作预测中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04379 2026-01-30 cs.CV cs.AI cs.CL cs.HC 67%

Can Large Language Models Capture Video Game Engagement?

大语言模型能否捕捉视频游戏参与度?

David Melhart, Matthew Barthet, Georgios N. Yannakakis

机构 * Institute of Digital Games, University of Malta Msida, Malta(马耳他大学数字游戏研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文研究了大语言模型在多模态输入下预测视频游戏参与度的能力,发现尽管LLMs在多个领域表现优异,但其在连续情绪标注上仍无法超越人类注释。

Comments This work has been submitted to the IEEE for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21670 2026-01-30 cs.CV cs.AI cs.LG physics.comp-ph 62%

MORPH: PDE Foundation Models with Arbitrary Data Modality

MORPH:具有任意数据模态的PDE基础模型

Mahindra Singh Rautela, Alexander Most, Siddharth Mansingh, Bradley C. Love, Alexander Scheinker, Diane Oyen, Nathan Debardeleben, Earl Lawrence, Ayan Biswas

机构 * Computing and Artificial Intelligence Division (CAI), Los Alamos National Laboratory, NM, US(计算与人工智能部门(CAI),洛斯阿拉莫斯国家实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MORPH是一种基于卷积视觉变压器的PDE基础模型,通过多模态处理和高效注意机制实现对异质科学数据的高效学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22153 2026-01-30 cs.RO cs.CV 57%

DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation

DynamicVLA: 一种用于动态物体操控的视觉-语言-动作模型

Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 DynamicVLA通过整合时间推理和闭环适应,提出了一种用于动态物体操控的视觉-语言-动作模型,提升了响应速度和泛化能力。

Comments Project Page: https://www.infinitescript.com/project/dynamic-vla/ GitHub: https://github.com/hzxie/DynamicVLA

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21193 2026-01-30 cs.CV 57%

Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval

生成回忆,密集重排序:学习多视图语义ID以实现高效的文本到视频检索

Zecheng Zhao, Zhi Chen, Zi Huang, Shazia Sadiq, Tong Chen

机构 * The University of Queensland(昆士兰大学) The University of Southern Queensland(南部昆士兰大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

AI总结 GRDR通过生成回忆和密集重排序方法,提升文本到视频检索的效率和准确性,减少存储并加速检索过程。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2601.22055 2026-01-30 cs.CL 83%

$G^2$-Reader: Dual Evolving Graphs for Multimodal Document QA

$G^2$-Reader: 双重演化图式用于多模态文档问答

Yaxin Du, Junru Song, Yifan Zhou, Cheng Wang, Jiahao Gu, Zimeng Chen, Menglan Chen, Wen Yao, Yang Yang, Ying Wen, Siheng Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Intelligent Game and Decision Laboratory(智能游戏与决策实验室)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 $G^2$-Reader通过双重图式解决多模态文档问答中的结构破坏和检索失效问题,实现66.21%的高准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09523 2026-01-30 cs.CV cs.AI cs.CY cs.LG 62%

MuseCL: Predicting Urban Socioeconomic Indicators via Multi-Semantic Contrastive Learning

MuseCL:通过多语义对比学习预测城市社会经济指标

Xixian Yong, Xiao Zhou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 MuseCL通过多语义对比学习融合视觉与文本信息,提升城市社会经济预测的精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22294 2026-01-30 cs.CV 57%

A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation

一种用于大规模3D检索和受控4D生成的三级对齐框架

Philip Xu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 Uni4D通过三级对齐框架实现大规模3D检索与可控4D生成,提升多模态动态理解与应用

Comments arXiv admin note: Author list truncated. This submission has been withdrawn by arXiv administrators as authors were added without their knowledge or consent

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03517 2026-01-30 cs.LG 50%

Understanding Self-Supervised Learning via Gaussian Mixture Models

通过高斯混合模型理解自监督学习

Parikshit Bansal, Ali Kavis, Sujay Sanghavi

机构 * UT Austin(得克萨斯大学)

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 本文通过高斯混合模型分析自监督学习,证明对比学习能有效找到最优低维子空间,优于传统谱技术,并验证了多模态对比学习的理论结果。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 15 篇

2507.13993 2026-01-30 eess.IV cs.AI cs.CV 81%

OrthoInsight: Rib Fracture Diagnosis and Report Generation Based on Multi-Modal Large Models

OrthoInsight:基于多模态大模型的肋骨骨折诊断与报告生成

Ningyong Wu, Jiangbo Zhang, Wenhong Zhao, Jinzhi Wang, Chenzhan Yu, Zhigang Xiu, Duwei Dai, Ziyu Xu, Yongli Yang

机构 * Organizational Management Department, School of Management, Xi’an Jiaotong University(管理学院组织管理部,西安交通大学) West China Longquan Hospital, Sichuan University(四川大学西部临床医学院) School of Electronic Science and Engineering, Xi’an Jiaotong University(西安交通大学电子科学与工程学院) Systems Engineering Institute, Xi’an Jiaotong University(西安交通大学系统工程研究院) Institute of Medical Artificial Intelligence, the Second Affiliated Hospital of Xi’an Jiaotong University(西安交通大学第二附属医院医学人工智能研究所) School of Human Settlements and Civil Engineering, Xi’an Jiaotong University(西安交通大学人居环境与土木工程学院) School of Life Science and Technology, Xi’an Jiaotong University(西安交通大学生命科学与技术学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 OrthoInsight通过多模态大模型实现肋骨骨折的自动诊断与报告生成,结合CT图像分析与医学知识图谱,提升诊断效率和临床实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21821 2026-01-30 cs.CV 79%

MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods

MMFineReason: 通过以开放数据为中心的方法缩小多模态推理差距

Honglin Lin, Zheng Liu, Yun Zhu, Chonghan Qin, Juekai Lin, Xiaoran Shang, Conghui He, Wentao Zhang, Lijun Wu

机构 * Shanghai Artificial Intelligence Laboratory, OpenDataLab(上海人工智能实验室、OpenDataLab) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) The University of Hong Kong(香港大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 MMFineReason通过大规模多模态推理数据集和基于难度的过滤策略,提升模型推理能力与参数效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21076 2026-01-30 cs.AI 79%

Multi-modal Imputation for Alzheimer's Disease Classification

多模态缺失数据填补用于阿尔茨海默病分类

Abhijith Shaji, Tamoghna Chattopadhyay, Sophia I. Thomopoulos, Greg Ver Steeg, Paul M. Thompson, Jose-Luis Ambite

机构 * Information Sciences Institute(信息科学研究所) University of Southern California(美国南加州大学) University of California(加州大学) Stevens Neuroimaging and Informatics Institute(史蒂文斯神经影像与信息学研究所)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 本文提出利用条件去噪扩散概率模型填补缺失的DWI扫描,以提高多模态深度学习模型在阿尔茨海默病三类分类中的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21950 2026-01-30 cs.LG 78%

Embracing Aleatoric Uncertainty in Medical Multimodal Learning with Missing Modalities

在医疗多模态学习中拥抱概率不确定性与缺失模态

Linxiao Gong, Yang Liu, Lianlong Sun, Yulai Bi, Jing Liu, Xiaoguang Zhu

机构 * HKUST (GZ)(香港科技大学) Tongji University(同济大学) University of Rochester(罗切斯特大学) Meta Fudan University(复旦大学) The University of British Columbia(不列颠哥伦比亚大学) University of California, Davis(加州大学戴维斯分校)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文提出AUM框架,通过建模单模态概率不确定性来应对医疗多模态学习中的缺失模态问题,在死亡预测任务中取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21129 2026-01-30 cs.RO 78%

WheelArm-Sim: A Manipulation and Navigation Combined Multimodal Synthetic Data Generation Simulator for Unified Control in Assistive Robotics

WheelArm-Sim: 一种结合 manipulation 和 navigation 的多模态合成数据生成模拟器,用于辅助机器人中的统一控制

Guangping Liu, Tipu Sultan, Vittorio Di Giorgio, Nick Hawkins, Flavio Esposito, Madi Babaiasl

机构 * Aerospace and Mechanical Engineering Department, Saint Louis University(圣路易斯大学航空航天与机械工程系) Computer Science Department, Saint Louis University(圣路易斯大学计算机科学系)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 WheelArm-Sim 是一种用于辅助机器人统一控制的多模态合成数据生成模拟器,通过集成轮椅和机械臂控制,为数据驱动的机器学习模型提供支持。

Comments Accepted to IEEE International Symposium on Medical Robotics (ISMR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20911 2026-01-30 cs.CV cs.AI 62%

Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs

非马尔可夫多轮对话图像生成与历史条件化大语言模型

Haochen Zhang, Animesh Sinha, Felix Juefei-Xu, Haoyu Ma, Kunpeng Li, Zhipeng Fan, Meng Dong, Xiaoliang Dai, Tingbo Hou, Peizhao Zhang, Zecheng He

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出非马尔可夫多轮对话图像生成方法,通过历史条件化框架和数据构建策略提升多轮一致性与指令遵循性,同时保持单轮编辑能力。

Comments 19 pages, 19 figures, plan for TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03147 2026-01-30 cs.HC cs.CL cs.LG cs.SD eess.AS 62%

A conversational gesture synthesis system based on emotions and semantics

基于情感和语义的对话手势合成系统

Thanh Hoang-Minh

机构 * Department of Information Technology, VNUHCM -- University of Science(越南科学大学信息科技系)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、eess.AS

AI总结 本文提出DeepGesture,一种基于扩散的 gesture 合成系统,通过多模态信号生成具有情感和语义条件的手势,提升数字人类的自然表达能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02507 2026-01-30 cs.CV cs.RO 57%

Keeping it Local, Tiny and Real: Automated Report Generation on Edge Computing Devices for Mechatronic-Based Cognitive Systems

保持本地化、小巧和现实:面向机电认知系统的边缘计算设备自动化报告生成

Nicolas Schuler, Lea Dewald, Jürgen Graf

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种基于边缘计算设备的自动化报告生成方法,利用本地模型实现多模态传感器数据处理,以提升机电认知系统在不同环境中的评估效率与隐私保护。

Comments 6 pages, 4 figures, 1 table; accepted for MECATRONICS-REM 2025 International Conference, PARIS, FRANCE December 3-5 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21669 2026-01-30 cs.LG cs.AI 57%

Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling

预期回报导致强化学习中的结果层面模式崩溃及如何通过逆概率缩放修复它

Abhijeet Sinha, Sundari Elango, Dianbo Liu

机构 * National University of Singapore, Singapore(新加坡国立大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出逆概率缩放方法,通过修正预期回报目标来解决强化学习中的结果层面模式崩溃问题,有效提升多模式策略优化的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏