arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-21 至 2025-11-21 共收录 45 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2511.15661 2025-11-21 cs.CV cs.AI cs.CL cs.LG 67%

VisPlay: Self-Evolving Vision-Language Models from Images

VisPlay: 从图像中自我进化视觉-语言模型

Yicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang, Yonghui Yang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Washington University in St. Louis(华盛顿大学圣路易斯分校) University of Maryland(马里兰大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 VisPlay通过自我进化强化学习框架,利用未标注图像数据提升视觉-语言模型的推理能力,实现多模态智能的可扩展发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16423 2025-11-21 cs.AI cs.CL 62%

TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language Models

TOFA: 无需训练的一次性联邦适应用于视觉-语言模型

Li Zhang, Zhongxuan Han, XiaoHua Feng, Jiaming Zhang, Yuyuan Li, Linbo Jiang, Jianan Lin, Chaochao Chen

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 TOFA提出一种无需训练的一次性联邦适应方法,通过视觉和文本管道提取任务相关表示,以高效适应视觉-语言模型在联邦学习中的应用。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16454 2025-11-21 cs.CV 57%

LLaVA$^3$: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs

LLaVA$^3$:像立体主义画家一样表示3D场景以提升VLM的3D场景理解

Doriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot, Loïc Barthe

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

AI总结 LLaVA$^3$通过立体主义方法提升VLM对3D场景的理解能力,利用多视角2D图像无需微调实现更优的3D场景理解。

Comments Accepted at AAAI'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09540 2025-11-21 cs.CV 57%

vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMs

vMFCoOp:在统一超球面流形上实现平衡的提示生物医学VLMs

Minye Shao, Sihan Guo, Xinrun Li, Xingyu Miao, Haoran Duan, Yang Long

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 vMFCoOp通过统一超球面流形上的vMF分布估计,实现LLM与CLIP主干间的语义对齐,提升生物医学VLMs的提示效果和少样本分类性能。

Comments Accepted as an Oral Presentation at AAAI 2026 Main Technical Track (this version is not peer-reviewed; it is the extended version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16170 2025-11-21 cs.CV 57%

Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective

通过注意力重新分配的目标重聚焦实现开放词汇语义分割:可解释性视角

Jiahao Li, Yang Lu, Yachao Zhang, Yong Xie, Fangyong Wang, Yuan Xie, Yanyun Qu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出RF-CLIP方法,通过模拟人类注意力重聚焦行为,解决CLIP在开放词汇语义分割中的多模态对齐问题,提升密集预测性能。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15720 2025-11-21 cs.AI 57%

Automated Hazard Detection in Construction Sites Using Large Language and Vision-Language Models

利用大语言和视觉-语言模型进行施工工地自动危险检测

Islem Sahraoui

机构 * University of Houston Cullen College of Engineering Department of Civil and Environmental Engineering(德克萨斯大学休斯顿分校库伦工程学院土木与环境工程系)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

AI总结 本研究利用大语言和视觉-语言模型,通过分析文本和图像数据,提高施工工地的安全隐患检测效率。

Comments Master thesis, University of Houton

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2501.01094 2025-11-21 cs.SD cs.AI cs.MM eess.AS 82%

MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions

MMVA: 基于估值和唤醒度的多模态匹配

Suhwan Choi, Kyu Won Kim, Myungjoo Kang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM、eess.AS

AI总结 MMVA通过多模态匹配基于估值和唤醒度,实现跨图像、音乐和音乐字幕的情感内容捕捉,并在估值-唤醒度预测任务中取得最佳性能。

Comments Paper accepted in Artificial Intelligence for Music workshop at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16225 2025-11-21 cs.LG 78%

Real-Time Inference for Distributed Multimodal Systems under Communication Delay Uncertainty

在通信延迟不确定性下的分布式多模态系统实时推理

Victor Croisfelt, João Henrique Inacio de Souza, Shashi Raj Pandey, Beatriz Soret, Petar Popovski

机构 * SNS JU project 6G-GOALS under the EU’s Horizon Europe program(欧盟地平线欧洲计划下的SNS JU项目6G-GOALS)

专题命中 音频语音多模态 :multimodal(title);audio-visual(abstract)

AI总结 本文提出了一种基于自适应时间窗口整合的非阻塞推理方法,以应对通信延迟不确定性,提升分布式多模态系统的实时推理性能。

Comments 6 pages, 3 figures, submitted to IEEE ICC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10108 2025-11-21 cs.LG q-bio.NC 50%

NeuroXVocal: Detection and Explanation of Alzheimer's Disease through Non-invasive Analysis of Picture-prompted Speech

NeuroXVocal:通过非侵入性分析提示语音检测和解释阿尔茨海默病

Nikolaos Ntampakis, Konstantinos Diamantaras, Ioanna Chouvarda, Magda Tsolaki, Vasileios Argyriou, Panagiotis Sarigianndis

机构 * International Hellenic University, Sindos, Greece MetaMind Innovations, Kozani, Greece Aristotle University of Thessaloniki, Thessaloniki, Greece Greek Association of Alzheimer’s Disease \& Related Disorders, Thessaloniki, Greece Kingston University London, London, UK University of Western Macedonia, Kozani, Greece

专题命中 音频语音多模态 :cross-modal(abstract)

AI总结 NeuroXVocal通过非侵入性语音分析实现阿尔茨海默病的检测与解释,结合高精度分类和文献支持的可解释性,提升临床诊断效果。

Journal ref Medical Image Computing and Computer Assisted Intervention - MICCAI 2025. Lecture Notes in Computer Science, vol 15973. Springer, Cham (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 6 篇

2511.16227 2025-11-21 cs.CV 79%

SwiTrack: Tri-State Switch for Cross-Modal Object Tracking

SwiTrack:跨模态目标跟踪的三态开关

Boyue Xu, Ruichao Hou, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) School of Mathematics, Southeast University(东南大学数学学院) Purple Mountain Laboratories(紫金山实验室)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 SwiTrack通过三态开关框架提升跨模态目标跟踪的鲁棒性和精度,实现7.2%和4.3%的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15722 2025-11-21 cs.AI 79%

Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods

多模态大语言模型中的空间推理:任务、基准和方法的调查

Weichen Liu, Qiyao Xue, Haoming Wang, Xiangyu Yin, Boyuan Yang, Wei Gao

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文调查多模态大语言模型中的空间推理问题,从认知角度分类任务和基准,分析评估方法与改进策略,揭示模型与人类推理间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16169 2025-11-21 eess.SP 78%

UT-OSANet: A Multimodal Deep Learning model for Evaluating and Classifying Obstructive Sleep Apnea

UT-OSANet: 一种用于评估和分类阻塞性睡眠呼吸暂停的多模态深度学习模型

Zijian Wang, Xiaoyu Bao, Chenhao Zhao, Jihui Zhang, Sizhi Ai, Yuanqing Li

专题命中 视频多模态 :multimodal(title);cross-modal(abstract)

AI总结 UT-OSANet是一种多模态深度学习模型,用于高精度评估和分类阻塞性睡眠呼吸暂停,通过多种输入模态和灵活训练策略实现事件层面的诊断。

Comments 12 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16524 2025-11-21 cs.CV 74%

BoxingVI: A Multi-Modal Benchmark for Boxing Action Recognition and Localization

BoxingVI:一种多模态基准,用于拳击动作识别与定位

Rahul Kumar, Vipul Baghel, Sudhanshu Singh, Bikash Kumar Badatya, Shivam Yadav, Babji Srinivasan, Ravi Hegde

机构 * Indian Institute of Technology Gandhinagar(印度理工学院甘地纳加尔) Indian Institute of Technology Madras(印度理工学院马德拉斯) Dr. A. P. J. Abdul Kalam Technical University(阿卜杜勒·卡拉姆技术大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 BoxingVI提供了一个多模态数据集,用于拳击动作识别与定位,旨在促进低资源环境下的实时视觉动作识别研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20470 2025-11-21 cs.CV 57%

Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence

Conan:基于多尺度视觉证据的逐步学习以像侦探一样推理

Kun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai, Hao Zhou, Jie Zhou, Fandong Meng, Xu Sun

机构 * State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 Conan通过多阶段渐进冷启动策略和AIR RLVR框架,实现证据基础的多步视频推理,超越基线模型,达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05274 2025-11-21 cs.CV 57%

From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos

从游戏到回放:面向时间精细粒度视频的组合视频检索

Animesh Gupta, Jay Parmar, Ishan Rajendrakumar Dave, Mubarak Shah

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 TF-CoVR提出了一种针对时间精细粒度视频检索的框架,通过预训练视频编码器和对比学习提升检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 1 篇

2511.16150 2025-11-21 cs.CV 86%

Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval

基于推理的嵌入:利用多模态大语言模型推理提升多模态检索

Chunxu Liu, Jiyuan Yang, Ruopeng Gao, Yuhan Zhu, Feng Zhu, Rui Zhao, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Sensetime Research(商汤科技研究院) Beijing Institute of Technology(北京理工大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(title);分类 cs.CV

AI总结 本文提出基于推理的嵌入方法,利用多模态大语言模型的推理能力提升多模态检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 7 篇

2510.22946 2025-11-21 cs.CV 83%

LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation

LightFusion: 一种轻量级、双融合框架用于统一多模态理解和生成

Zeyu Wang, Zilong Chen, Chenhui Gou, Feng Li, Chaorui Deng, Deyao Zhu, Kunchang Li, Weihao Yu, Haoqin Tu, Haoqi Fan, Cihang Xie

机构 * UC Santa Cruz(加州大学圣克ruz分校) Tsinghua University(清华大学) Monash University(墨尔本大学) ByteDance Seed Project(字节跳动种子项目)

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 LightFusion通过双融合机制实现轻量级多模态理解和生成,仅用350亿个标记训练即在多个基准测试中取得优异成绩。

Comments Preprint. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17534 2025-11-21 cs.CV cs.CL cs.MM 82%

Co-Reinforcement Learning for Unified Multimodal Understanding and Generation

协同强化学习用于统一多模态理解和生成

Jingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang, Chao Ma

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 本文提出CoRL框架,通过协同强化学习提升多模态大语言模型在生成与理解任务上的性能。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08987 2025-11-21 cs.AI 79%

Towards Efficient Multimodal Unified Reasoning Model via Model Merging

通过模型合并实现高效的多模态统一推理模型

Qixiang Yin, Huanjin Yao, Jianghao Chen, Jiaxing Huang, Zhicheng Zhao, Fei Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学) Beijing Key Laboratory of Network System and Network Culture(北京网络系统与网络文化重点实验室) Key Laboratory of Interactive Technology and Experience System, Ministry of Culture and Tourism(文化和旅游部交互技术与体验系统重点实验室) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 Tiny-R1V通过两阶段优化和模型合并方法,实现高效多模态推理,提升轻量模型在多种任务中的性能。

Comments Technical report, Code will be available at https://github.com/buptyqx/Tiny-R1V

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16671 2025-11-21 cs.CV cs.AI cs.CL 67%

Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation

生成中思考:在视觉生成过程中穿插文本推理

Ziyu Guo, Renrui Zhang, Hongyu Li, Manyuan Zhang, Xinyan Chen, Sifan Wang, Yan Feng, Peng Pei, Pheng-Ann Heng

机构 * CUHK(香港中文大学) IMIXR MMLab Meituan Project(美团项目)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出TwiG框架,通过在视觉生成过程中穿插文本推理,提升生成内容的上下文感知和语义丰富性。

Comments Project Page: https://think-while-gen.github.io Code: https://github.com/ZiyuGuo99/Thinking-while-Generating

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23584 2025-11-21 cs.CV 57%

VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement

VividFace: 高质量和高效的一步扩散用于视频面部增强

Shulian Zhang, Yong Guo, Long Peng, Ziyang Wang, Ye Chen, Wenbo Li, Xiao Zhang, Yulun Zhang, Jian Chen

机构 * South China University of Technology(华南理工大学) Max Planck Institute for Informatics(马克斯·普朗克研究所(信息学)) University of Science and Technology of China(中国科学技术大学) The Chinese University of Hong Kong(香港中文大学) Nanjing University of Science and Technology(南京理工大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

AI总结 VividFace通过单步扩散框架和联合训练策略,高效提升视频面部增强的高质量与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14023 2025-11-21 cs.CL 57%

Synthetic Data Generation Using Large Language Models: Advances in Text and Code

利用大语言模型生成合成数据:文本与代码领域的进展

Mihai Nadas, Laura Diosan, Andreea Tomescu

机构 * Faculty of Mathematics and Computer Science, Babeş-Bolyai University(巴贝什-博耶亚大学数学与计算机科学系) KlusAI Research Lab(KlusAI研究实验室)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CL

AI总结 本文探讨了利用大语言模型生成合成数据在文本和代码领域的新进展,分析了其在低资源任务和代码应用中的潜力及挑战。

Comments 24 pages, 6 tables, 1 figure, 64 references

Journal ref IEEE Access 13, 134615-134633 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15550 2025-11-21 cs.RO 50%

UltraDP: Generalizable Carotid Ultrasound Scanning with Force-Aware Diffusion Policy

UltraDP: 通用的颈动脉超声扫描与力感知扩散策略

Ruoqu Chen, Xiangjie Yan, Kangchen Lv, Gao Huang, Zheng Li, Xiang Li

机构 * Department of Automation, Tsinghua University(自动化系,清华大学) Department of Surgery, Chow Yuk Ho Technology Centre for Innovative Medicine, Li Ka Shing Institute of Health Science and Multi-scale Medical Robotics Center, The Chinese University of Hong Kong, Hong Kong(外科系,创新医学技术中心,利文斯科健康科学研究所及多尺度医学机器人中心,香港中文大学,香港)

专题命中 多模态生成 :multi-modal(abstract)

AI总结 UltraDP通过力感知扩散策略实现通用颈动脉超声扫描,利用多感官输入和混合控制器提高扫描成功率至95%。

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 9 篇

2511.16221 2025-11-21 cs.CV cs.CL 81%

Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions

大语言模型能读取环境吗?一个多模态基准用于评估多当事人社交互动中的欺骗

Caixin Kang, Yifei Huang, Liangyang Ouyang, Mingfang Zhang, Ruicong Liu, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出多模态互动欺骗评估任务,通过新颖数据集评估多种MLLMs的欺骗检测能力,揭示其在多模态社交线索处理上的不足,并提出SoCoT和DSEM模块提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19390 2025-11-21 eess.SP 78%

Depression diagnosis from patient interviews using multimodal machine learning

基于多模态机器学习的抑郁症诊断

Jana Weber, Marcel Weber, Juan Miguel Lopez Alcaraz

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本研究通过多模态机器学习整合语音、语言和临床信息,提升抑郁症诊断的准确性与临床实用性。

Comments Accepted by Frontiers in Psychiatry, 19 pages, 5 figures, source code under https://github.com/UOLMDA2025/Depression

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15205 2025-11-21 cs.LG cs.AI cs.CL 62%

Long-Short Distance Graph Neural Networks and Improved Curriculum Learning for Emotion Recognition in Conversation

长短期距离图神经网络与改进的课程学习用于对话中的情绪识别

Xinran Li, Xiujuan Xu, Jiaqi Qiao

机构 * School of Software Technology, Dalian University of Technology(大连理工大学软件学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出长短期距离图神经网络与改进课程学习方法,用于对话中情绪识别,通过多模态特征提取和优化训练策略提升模型性能。

Comments Accepted by the 28th European Conference on Artificial Intelligence (ECAI 2025)

Journal ref ECAI 2025, Frontiers in Artificial Intelligence and Applications, Volume 413, pp. 4033-4040, IOS Press, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16440 2025-11-21 cs.CV 57%

StreetView-Waste: A Multi-Task Dataset for Urban Waste Management

StreetView-Waste: 一个用于城市垃圾管理的多任务数据集

Diogo J. Paulo, João Martins, Hugo Proença, João C. Neves

机构 * University of Beira Interior(贝拉内陆大学) IT: Instituto de Telecomunicações(电信研究所) NOVA LINCS

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 StreetView-Waste数据集通过多任务设置,提供垃圾桶检测、跟踪和溢出分割的基准,结合启发式方法和几何先验提升模型性能。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14945 2025-11-21 cs.CV 57%

Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities

无监督发现人类活动中的长期时空周期性工作流

Fan Yang, Quanting Xie, Atsunori Moteki, Shoichi Masui, Shan Jiang, Kanji Uchino, Yonatan Bisk, Graham Neubig

机构 * Fujitsu Research of America, USA(美国富士通研究机构) Fujitsu Limited, Japan(日本富士通有限公司) Carnegie Mellon University, USA(美国卡内基梅隆大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出首个包含长期周期性工作流的基准,通过轻量级基线实现无监督检测和异常检测,显著优于现有方法并具备实际部署优势。

Comments accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16163 2025-11-21 cs.CV 57%

An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs

一张图像等于一万字:针对视觉语言模型的冗余文本诱导攻击

Zhi Luo, Zenghui Yuan, Wenqi Wei, Daizong Liu, Pan Zhou

机构 * Huazhong University of Science and Technology(华中科技大学) Fordham University(福特汉姆大学) Wuhan University(武汉大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出一种针对视觉语言模型的冗余文本诱导攻击,通过两阶段框架生成恶意图像以诱导模型生成冗长文本,提升攻击效果与可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16048 2025-11-21 cs.RO cs.AI cs.HC 57%

Semantic Glitch: Agency and Artistry in an Autonomous Pixel Cloud

语义裂隙:自主像素云中的代理与艺术性

Qing Zhang, Jing Huang, Mingyang Xu, Jun Rekimoto

机构 * The University of Tokyo(东京大学) Tokyo University of the Arts(东京艺术大学) Keio University(庆应大学) SONY CSL Kyoto(索尼 CSL 京都)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 本文提出语义裂隙,通过多模态大语言模型实现自主导航,展示低 fidelity方法在创造不完美但富有艺术性的机器人同伴中的应用。

Comments NeurIPS 2025 Creative AI Track, The Thirty-Ninth Annual Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏