arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2511.15497 2025-12-16 eess.SP 50%

A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems

机械系统中空化强度识别的机器学习综述

Yu Sha, Ningtao Liu, Haofeng Liu, Junqi Tao, Zhenxing Niu, Guojun Huang, Yao Yao, Jiaqi Liang, Moxian Qian, Horst Stoecker, Domagoj Vnucec, Andreas Widl, Kai Zhou

专题命中 安全评测 :safety(abstract)

AI总结 本文综述了机械系统中空化强度识别的机器学习发展,强调传统方法与深度学习的演变,并展望未来在多源数据处理和工业应用中的发展方向。

Comments 43 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11680 2025-12-15 cs.CV 50%

Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing

跨模态上下文感知学习:用于遥感中视觉提示引导的多模态图像理解

Xu Zhang, Jiabin Fang, Zhuoming Ding, Jin Yuan, Xuan Liu, Qianjun Zhang, Zhiyong Li

机构 * College of Computer Science and Electronic Engineering, Hunan University(计算机科学与电子工程学院,湖南大学) School of Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University(机器人学院及机器人视觉感知与控制技术国家工程研究中心,湖南大学) School of Computing and Artificial Intelligence, Southwest Jiaotong University(计算与人工智能学院,西南交通大学)

专题命中 安全评测 :alignment(abstract)

AI总结 CLV-Net通过跨模态上下文感知学习,利用视觉提示引导遥感多模态图像理解,提升目标识别精度和用户意图对齐能力。

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11215 2025-12-15 cs.CV 50%

SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection

SmokeBench: 评估多模态大语言模型用于野火烟雾检测

Tianye Qi, Weihao Li, Nick Barnes

机构 * Australian National University(澳大利亚国立大学)

专题命中 安全评测 :safety(abstract)

AI总结 SmokeBench评估多模态大语言模型在野火烟雾检测中的性能,发现模型在烟雾定位方面存在显著局限,尤其在早期阶段。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10674 2025-12-12 cs.CV 50%

Geo6DPose: Fast Zero-Shot 6D Object Pose Estimation via Geometry-Filtered Feature Matching

Geo6DPose:通过几何过滤特征匹配实现快速零样本6D物体姿态估计

Javier Villena Toro, Mehdi Tarkian

机构 * Linköping University(利乌普斯大学)

专题命中 安全评测 :alignment(abstract)

AI总结 Geo6DPose通过几何过滤特征匹配实现快速零样本6D姿态估计,无需训练或网络访问,具备高效推理与高鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09271 2025-12-11 cs.CV 50%

LongT2IBench: A Benchmark for Evaluating Long Text-to-Image Generation with Graph-structured Annotations

LongT2IBench: 一个用于评估长文本到图像生成的基准,具有图结构注释

Zhichao Yang, Tianjiao Gu, Jianjie Wang, Feiyu Lin, Xiangfei Sheng, Pengfei Chen, Leida Li

专题命中 安全评测 :alignment(abstract)

AI总结 LongT2IBench通过图结构注释和LongT2IExpert提出,用于评估长文本到图像生成的对齐和解释能力。

Comments The paper has been accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09092 2025-12-11 cs.CV 50%

Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters

解释未见的:多模态视觉-语言推理用于地下矿难中的情境感知

Mizanur Rahman Jewel, Mohamed Elmahallawy, Sanjay Madria, Samuel Frimpong

机构 * Missouri University of Science and Technology(密苏里科学与技术大学) Washington State University(华盛顿州立大学)

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出MDSE框架,通过多模态视觉-语言推理提升地下矿难情境感知能力,实现更准确的描述生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17353 2025-12-11 cs.CE 50%

RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding

RoadBench: 一种用于道路损伤理解的视觉-语言基础模型和基准

Xi Xiao, Yunbei Zhang, Janet Wang, Lin Zhao, Yuxiang Wei, Hengjia Li, Yanshu Li, Xinyuan Song, Xiao Wang, Swalpa Kumar Roy, Hao Xu, Tianyang Wang

专题命中 安全评测 :safety(abstract)

AI总结 RoadBench通过整合视觉与文本信息,提出RoadCLIP模型,显著提升道路损伤识别性能,为基础设施监测提供新基准。

Comments Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08860 2025-12-10 cs.CV 50%

Tri-Bench: Stress-Testing VLM Reliability on Spatial Reasoning under Camera Tilt and Object Interference

Tri-Bench:在相机倾斜和物体干扰下测试VLM在空间推理中的可靠性

Amit Bendkhale

机构 * Amit Bendkhale(独立研究者)

专题命中 安全评测 :trustworthy(abstract)

AI总结 Tri-Bench通过测试VLM在相机倾斜和物体干扰下的空间推理可靠性,揭示了模型在几何推理中的不足。

Comments 6 pages, 3 figures. Code and data: https://github.com/Amiton7/Tri-Bench. Accepted to the AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08645 2025-12-10 cs.CV 50%

Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation

图像生成链:迈向可监控和可控的图像生成

Young Kyung Kim, Oded Schlesinger, Yuzhou Zhao, J. Matias Di Martino, Guillermo Sapiro

机构 * Duke University(杜克大学) Princeton University(普林斯顿大学) Universidad Católica del Uruguay(乌拉圭天主教大学) Apple(苹果公司)

专题命中 安全评测 :safety(abstract)

AI总结 CoIG框架通过将图像生成过程分解为可监控的步骤,提升图像生成的可控性和可解释性。

Comments 19 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11110 2025-12-10 cs.CV 50%

Beyond accuracy: quantifying the reliability of Multiple Instance Learning for Whole Slide Image classification

超越准确性:量化多实例学习在全滑片图像分类中的可靠性

Hassan Keshvarikhojasteh, Marc Aubreville, Christof A. Bertram, Josien P. W. Pluim, Mitko Veta

机构 * Department of Biomedical Engineering, Eindhoven University of Technology(埃因霍温理工大学生物医学工程系) Flensburg Artificial Intelligence Research Group (FLAIR), Flensburg University of Applied Sciences(弗劳恩霍夫应用科技大学人工智能研究组) University of Veterinary Medicine Vienna(维也纳兽医大学)

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出三种量化指标评估多实例学习在全滑片图像分类中的可靠性,发现MEAN-POOL-INS模型在可靠性上表现优异,为未来研究提供可靠基准。

Journal ref PloS one. 2025 Dec 5;20(12):e0337261

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03080 2025-12-10 q-bio.TO 50%

Explainable AI for computational pathology identifies model limitations and tissue biomarkers

可解释AI用于计算病理学识别模型限制和组织生物标志物

Jakub R. Kaczmarzyk, Chanwoo Kim, Soham Gadgil, Deepika Savant, Zhen Zhao, Joel H. Saltz, Su-In Lee, Peter K. Koo

专题命中 安全评测 :trustworthy(abstract)

AI总结 HIPPO通过可解释AI方法识别模型限制和组织生物标志物,提升数字病理学中模型的可解释性和临床应用价值。

Comments The following authors contributed equally: Jakub R. Kaczmarzyk, Chanwoo Kim, Soham Gadgil

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07449 2025-12-09 cs.PF 50%

AFarePart: Accuracy-aware Fault-resilient Partitioner for DNN Edge Accelerators

AFarePart: 用于DNN边缘加速器的精度感知故障容错分区器

Mukta Debnath, Krishnendu Guha, Debasri Saha, Amlan Chakrabarti, Susmita Sur-Kolay

专题命中 安全评测 :safety(abstract)

AI总结 本文提出了一种精度感知的DNN分区框架,通过NSGA-II多目标优化提升故障容错性,实验显示在性能开销小的情况下,故障容忍度提升达27.7%。

Comments 6 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13036 2025-12-09 cs.CV 50%

uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data

uCLIP: 无配对数据下多语言视觉语言模型的参数高效扩展

Dahyun Chung, Donghyun Shin, Yujin Sung, Seunggi Moon, Jinwoo Jeon, Byung-Jun Lee

专题命中 安全评测 :alignment(abstract)

AI总结 uCLIP通过无需配对数据的轻量级框架,在低资源语言中实现高效多语言视觉语言对齐。

Comments Our project page can be found at https://dinyudin203.github.io/uCLIP-project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06645 2025-12-09 cs.MA 50%

Analyzing Collision Rates in Large-Scale Mixed Traffic Control via Multi-Agent Reinforcement Learning

通过多智能体强化学习分析大规模混合交通控制中的碰撞率

Muyang Fan

专题命中 安全评测 :safety(abstract)

AI总结 本研究通过多智能体强化学习分析大规模混合交通控制中的碰撞率影响因素,探讨交通密度、信号协调和转向策略对碰撞风险的作用,为提升交通系统安全性和效率提供理论支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06583 2025-12-09 econ.GN q-fin.EC 50%

Tournament-Based Performance Evaluation and Systematic Misallocation: Why Forced Ranking Systems Produce Random Outcomes

基于竞赛的绩效评估与系统性误分配:为什么强制排名系统产生随机结果

Jeremy McEntire

专题命中 安全评测 :alignment(abstract)

AI总结 本文揭示强制排名系统因系统性误分配导致随机结果,指出其无法有效解决委托-代理问题,反而加剧了分配误差。

Comments 31 pages, 6 tables. Agent-based simulation demonstrating structural allocation failures in tournament-based forced distribution evaluation mechanisms. Includes sensitivity analyses across team bias levels, alternative distributions, and cutoff percentages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05559 2025-12-08 q-fin.CP 50%

A Unified AI System For Data Quality Control and DataOps Management in Regulated Environments

统一的AI系统用于受监管环境中的数据质量控制和DataOps管理

Devender Saini, Bhavika Jain, Nitish Ujjwal, Philip Sommer, Dan Romuald Mbanga, Dhagash Mehta

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出了一种统一的AI驱动框架,用于受监管环境中数据质量控制和DataOps管理,通过整合多种方法提升数据管道的可靠性与合规性。

Comments 10 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05511 2025-12-08 cs.CV 50%

Rethinking Infrared Small Target Detection: A Foundation-Driven Efficient Paradigm

重新思考红外小目标检测:一种基础驱动的高效范式

Chuang Yu, Jinmiao Zhao, Yunpeng Liu, Yaokun Li, Xiujun Shu, Yuanhao Feng, Bo Wang, Yimian Dai, Xiangyu Yue

机构 * Key Laboratory of Opto-Electronic Information Processing, Chinese Academy of Sciences(光电信息处理重点实验室,中国科学院) Shenyang Institute of Automation, Chinese Academy of Sciences(沈阳自动化研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Sun Yat-sen University(中山大学) Tencent(腾讯) Nankai University(南开大学) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab)

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出FDEP范式,通过基础模型冻结表示与任务特征融合,提升红外小目标检测精度,构建综合评估指标,实现多数据集SOTA性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05470 2025-12-08 cs.SE 50%

Everything is Context: Agentic File System Abstraction for Context Engineering

一切皆为上下文:面向上下文工程的代理文件系统抽象

Xiwei Xu, Robert Mao, Quan Bai, Xuewu Gu, Yechao Li, Liming Zhu

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出一种基于文件系统的上下文工程抽象,通过统一挂载、元数据和访问控制,实现可验证的上下文管理,支持可问责和以人类为中心的AI协作。

Comments Submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14257 2025-12-08 cs.SE 50%

C2SaferRust: Transforming C Projects into Safer Rust with NeuroSymbolic Techniques

C2SaferRust:利用神经符号技术将C项目转化为更安全的Rust

Vikram Nitin, Rahul Krishna, Luiz Lemos do Valle, Baishakhi Ray

专题命中 安全评测 :safety(abstract)

AI总结 C2SaferRust结合规则系统和LLMs,将C代码转换为更安全的Rust,减少指针和unsafe代码,提升安全性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03635 2025-12-04 cs.LO cs.SE 50%

Formal Analysis of the Sigmoid Function and Formal Proof of the Universal Approximation Theorem

对Sigmoid函数的正式分析及通用逼近定理的正式证明

Dustin Bryant, Jim Woodcock, Simon Foster

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文在Isabelle/HOL中形式化分析了Sigmoid函数并证明了通用逼近定理,填补了形式化证明库的空白,提升了神经网络的可信度。

Comments 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11045 2025-12-04 cs.CV 50%

StableV2V: Stablizing Shape Consistency in Video-to-Video Editing

StableV2V: 视频到视频编辑中的形状一致性稳定

Chang Liu, Rui Li, Kaidong Zhang, Yunwei Lan, Dong Liu

专题命中 安全评测 :alignment(abstract)

AI总结 StableV2V通过分解编辑流程并建立运动与提示的对齐,提升视频到视频编辑的形状一致性与效率。

Comments Project page: https://alonzoleeeooo.github.io/StableV2V, code: https://github.com/AlonzoLeeeooo/StableV2V, model weights: https://huggingface.co/AlonzoLeeeooo/StableV2V, dataset (DAVIS-Edit): https://huggingface.co/datasets/AlonzoLeeeooo/DAVIS-Edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17699 2025-12-03 cs.CV q-bio.QM 50%

Benchmarking Class Activation Map Methods for Explainable Brain Hemorrhage Classification on Hemorica Dataset

在Hemorica数据集上基于类激活图方法的可解释性脑出血分类基准测试

Z. Rafati, M. Hoseyni, J. Khoramdel, A. Nikoofard

机构 * Faculty of Computer Engineering, K. N. Toosi University of Technology(计算机工程学院,K.N.托菲大学) Faculty of Electrical Engineering, K. N. Toosi University of Technology(电气工程学院,K.N.托菲大学) Faculty of Mechanical Engineering, Tarbiat Modares University(机械工程学院,塔里比亚特莫达res大学)

专题命中 安全评测 :alignment(abstract)

AI总结 本研究通过在Hemorica数据集上评估九种CAM算法,发现AblationCAM在像素级Dice和IoU指标上表现最佳,为脑出血分类的可解释性提供了新的基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02410 2025-12-03 cs.MA cs.CR 50%

Decentralized Multi-Agent System with Trust-Aware Communication

去中心化多智能体系统与信任感知通信

Yepeng Ding, Ahmed Twabi, Junwei Yu, Lingfeng Zhang, Tohru Kondo, Hiroyuki Sato

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出了一种基于区块链的去中心化多智能体系统,通过信任感知通信协议解决传统集中式系统中的信任、可扩展性和抗审查问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02329 2025-12-03 cs.SE 50%

Towards autonomous normative multi-agent systems for Human-AI software engineering teams

迈向自主规范的多智能体系统用于人机软件工程团队

Hoa Khanh Dam, Geeta Mahala, Rashina Hoda, Xi Zheng, Cristina Conati

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出自主规范的多智能体系统,通过大型语言模型赋能,实现人机协作的高效软件开发。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22686 2025-12-03 cs.CV 50%

Emergent Extreme-View Geometry in 3D Foundation Models

三维基础模型中的涌现极端视角几何

Yiwen Zhang, Joseph Tung, Ruojin Cai, David Fouhey, Hadar Averbuch-Elor

机构 * Cornell University(康奈尔大学) New York University(纽约大学) Kempner Institute, Harvard University(哈佛大学凯普勒研究所)

专题命中 安全评测 :alignment(abstract)

AI总结 本文研究了三维基础模型在极端视角下的几何理解能力,并提出轻量级对齐方案提升其相对姿态估计性能,同时引入新的未见过的互联网场景基准。

Comments Project page is at https://ext-3dfms.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00928 2025-12-02 cs.MM 50%

Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation

增强多模态大语言模型的模态内理解以实现鲁棒的多模态关键词生成

Jiajun Cao, Qinggang Zhang, Yunbo Tang, Zhishang Xiang, Chang Yang, Jinsong Su

专题命中 安全评测 :alignment(abstract)

AI总结 AimKP通过增强模态内语义学习和跨模态对齐,提升多模态关键词生成的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12077 2025-12-02 cs.CV 50%

Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound

通过视觉学习听觉:现在是让视觉语言模型理解艺术情感的时候了

Dengming Zhang, Weitao You, Jingxiong Li, Weishen Lin, Wenda Shi, Xue Zhao, Heda Zuo, Junxian Wu, Lingyun Sun

机构 * Zhejiang University(浙江大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 安全评测 :alignment(abstract)

AI总结 VAEmotionLLM通过两阶段框架,利用有限音频预训练使视觉语言模型理解艺术中的多模态情感

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23594 2025-12-02 cs.CV 50%

PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection

PRISM-Bench: 一个包含推理错误检测的基于谜题的视觉任务基准

Yusu Qian, Cheng Wan, Chao Jia, Yinfei Yang, Qingyu Zhao, Zhe Gan

专题命中 安全评测 :trustworthy(abstract)

AI总结 PRISM-Bench通过检测推理错误评估多模态模型的视觉推理能力,揭示了流畅生成与忠实推理之间的差距。

Comments This paper's first error detection task's ground truth data contains hallucination introduced by gpt and needs to be withdrawn

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19480 2025-12-02 cs.CV 50%

Learning by Aligning 2D Skeleton Sequences and Multi-Modality Fusion

通过对齐2D骨骼序列和多模态融合进行学习

Quoc-Huy Tran, Muhammad Ahmed, Murad Popattia, M. Hassan Ahmed, Andrey Konin, M. Zeeshan Zia

机构 * Retrocausal, Inc.(Retrocausal公司)

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出通过2D骨骼热图和多模态融合提升时间视频对齐的自监督学习方法,实现更高的准确性和鲁棒性。

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00082 2025-12-02 cs.CV 50%

Exploring Diagnostic Prompting Approach for Multimodal LLM-based Visual Complexity Assessment: A Case Study of Amazon Search Result Pages

探索用于多模态大语言模型视觉复杂性评估的诊断提示方法:亚马逊搜索结果页面案例研究

Divendar Murtadak, Yoon Kim, Trilokya Akula

专题命中 安全评测 :alignment(abstract)

AI总结 本研究通过对比诊断提示与传统提示方法,发现其在提升多模态大语言模型对亚马逊搜索结果页面视觉复杂性评估的可靠性方面有显著改进,但仍有待进一步优化。

Comments 9 pages, 4 figures, 9 tables. Study on diagnostic prompting for multimodal LLM-based visual complexity assessment of Amazon search result pages

详情

展开后加载摘要…

URL PDF HTML 收藏