arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-26 至 2025-11-26 共收录 21 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 21 篇

2511.19863 2025-11-26 cs.CY 88%

International AI Safety Report 2025: Second Key Update: Technical Safeguards and Risk Management

2025国际人工智能安全报告:第二关键更新:技术安全措施与风险管理工作

Yoshua Bengio, Stephen Clare, Carina Prunkl, Maksym Andriushchenko, Ben Bucknall, Philip Fox, Nestor Maslej, Conor McGlynn, Malcolm Murray, Shalaleh Rismani, Stephen Casper, Jessica Newman, Daniel Privitera, Sören Mindermann, Daron Acemoglu, Thomas G. Dietterich, Fredrik Heintz, Geoffrey Hinton, Nick Jennings, Susan Leavy, Teresa Ludermir, Vidushi Marda, Helen Margetts, John McDermid, Jane Munga, Arvind Narayanan, Alondra Nelson, Clara Neppel, Gopal Ramchurn, Stuart Russell, Marietje Schaake, Bernhard Schölkopf, Alavaro Soto, Lee Tiedrich, Gaël Varoquaux, Andrew Yao, Ya-Qin Zhang, Leandro Aguirre, Olubunmi Ajala, Fahad Albalawi, Noora AlMalek, Christian Busch, André Carvalho, Jonathan Collas, Amandeep Gill, Ahmet Hatip, Juha Heikkilä, Chris Johnson, Gill Jolly, Ziv Katzir, Mary Kerema, Hiroaki Kitano, Antonio Krüger, Aoife McLysaght, Oleksii Molchanovskyi, Andrea Monti, Kyoung Mu Lee, Mona Nemer, Nuria Oliver, Raquel Pezoa, Audrey Plonk, José Portillo, Balaraman Ravindran, Hammam Riza, Crystal Rugege, Haroon Sheikh, Denise Wong, Yi Zeng, Liming Zhu

专题命中 安全评测 :safety(title,abstract);AI safety(title,abstract);分类 cs.CY

AI总结 2025国际人工智能安全报告第二关键更新,探讨通用人工智能风险管理工作进展,包括技术安全措施、对抗训练、数据整理及制度框架的建立。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04704 2025-11-26 cs.CV cs.AI 85%

HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model

HoliSafe: 视觉-语言模型的综合安全性评估与建模

Youngwan Lee, Kangsan Kim, Kwanyong Park, Ilcahe Jung, Soojin Jang, Seanie Lee, Yong-Ju Lee, Sung Ju Hwang

机构 * ETRI KAIST AI(韩国科学技术院人工智能研究所) University of Seoul(首尔大学)

专题命中 安全评测 :safety(title,abstract);alignment(abstract);jailbreak(abstract);分类 cs.AI

AI总结 HoliSafe提出了一种模块化框架和视觉守护模块,通过综合安全性数据集提升视觉-语言模型的安全性,展现其在多个基准中的优越表现。

Comments Project page: https://youngwanlee.github.io/holisafe

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20604 2025-11-26 cs.CL cs.AI cs.LG 82%

On Evaluating LLM Alignment by Evaluating LLMs as Judges

通过评估LLM作为裁判来评估LLM对齐

Yixin Liu, Pengfei Liu, Arman Cohan

机构 * Yale University(耶鲁大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出AlignEval基准,通过评估LLM作为裁判的能力来衡量其对齐人类偏好的效果,结果优于现有自动评估基准。

Comments NeurIPS 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19477 2025-11-26 cs.SE 67%

Building Browser Agents: Architecture, Security, and Practical Solutions

构建浏览器代理:架构、安全与实用解决方案

Aram Vardanyan

专题命中 安全评测 :safety(abstract);prompt injection(abstract)

AI总结 本文提出通过专用工具和编程约束构建安全浏览器代理,实现85%的成功率。

Comments 30 pages, 22 figures. Production architecture and benchmark evaluation of browser agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19930 2025-11-26 cs.GT cs.CY cs.LG 62%

Designing Reputation Systems for Manufacturing Data Trading Markets: A Multi-Agent Evaluation with Q-Learning and IRL-Estimated Utilities

为制造数据交易市场设计声誉系统:基于Q学习和IRL估计效用的多智能体评估

Kenta Yamamoto, Teruaki Hayashi

机构 * Department of Systems Innovation, Graduate School of Engineering, The University of Tokyo Tokyo, Japan(系统创新部门,工学研究生院,东京大学东京,日本)

专题命中 安全评测 :alignment(abstract);分类 cs.CY、cs.LG

AI总结 本研究通过多智能体模拟器评估了五种声誉系统,发现PeerTrust在数据价格与质量一致性及防止垄断方面表现最佳,并提出混合声誉机制提升市场稳定性。

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19845 2025-11-26 cs.LG cs.CY stat.ML 62%

SX-GeoTree: Self-eXplaining Geospatial Regression Tree Incorporating the Spatial Similarity of Feature Attributions

SX-GeoTree: 自解释地理回归树结合特征归因的空间相似性

Chaogui Kang, Lijian Luo, Qingfeng Guan, Yu Liu

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY、cs.LG

AI总结 SX-GeoTree通过整合空间相似性与模块度最大化,提升地理回归树的解释稳定性与空间残差均匀性。

Comments 41 pages, 7 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17681 2025-11-26 cs.LG cs.AI 62%

Unlearning as Ablation: Toward a Falsifiable Benchmark for Generative Scientific Discovery

反例作为消去:面向生成科学发现的可证伪基准

Robert Yang

机构 * S6 Research(S6研究)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出'反例作为消去'作为可证伪基准,旨在检验AI在科学发现中的生成能力,通过系统移除目标结果并评估模型能否重新推导,推动AI-科学的基准发展。

Comments 6 pages + appendix. Accepted to NeurIPS 2025 AI4Science Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10161 2025-11-26 cs.CL cs.AI 62%

LaajMeter: A Framework for LaaJ Evaluation

LaajMeter:一种用于LaaJ评估的框架

Samuel Ackerman, Gal Amram, Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich, Raviv Gal, Wesam Ibraheem, Avi Ziv

机构 * IBM Research, Israel(IBM以色列研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 LaaJMeter是一种用于LaaJ评估的模拟框架,通过生成合成数据来系统分析评估度量标准,帮助验证特定任务的LaaJ质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08701 2025-11-26 cs.CV cs.AI cs.LG 62%

SafeFix: Targeted Model Repair via Controlled Image Generation

SafeFix: 通过受控图像生成实现目标模型修复

Ouyang Xu, Baoming Zhang, Ruiyu Mao, Yunhui Guo

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 SafeFix通过生成语义忠实的图像来修复模型,提升模型对罕见案例的鲁棒性,减少系统性错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19480 2025-11-26 cs.LG cs.AI 62%

Exploiting the Experts: Unauthorized Compression in MoE-LLMs

利用专家:MoE-LLMs中的未经授权压缩

Pinaki Prasad Guha Neogi, Ahmad Mohammadshirazi, Dheeraj Kulshrestha, Rajiv Ramnath

机构 * Ohio State University(俄亥俄州立大学) Flairsoft

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了MoE-LLMs在特定任务中的可剪枝性,提出专家归因框架和防御策略以防止未经授权的压缩和微调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20650 2025-11-26 cs.CV cs.AI 57%

MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities

MedROV:面向跨多种医学影像模态的实时开放词汇检测

Tooba Tehreem Sheikh, Jean Lahoud, Rao Muhammad Anwer, Fahad Shahbaz Khan, Salman Khan, Hisham Cholakkal

机构 * Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 MedROV是首个实时开放词汇医学影像检测模型,通过大规模数据集和伪标签策略提升检测性能,实现40 mAP50的提升并达到70 FPS的实时处理速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20526 2025-11-26 cs.AI 57%

Assessing LLMs' Performance: Insights from the Chinese Pharmacist Exam

评估大语言模型的性能:来自中国药师资格考试的洞察

Xinran Wang, Boran Zhu, Shujuan Zhou, Ziwen Long, Dehua Zhou, Shu Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本研究比较了DeepSeek-R1和ChatGPT-4o在药师资格考试中的表现,发现DeepSeek-R1在准确性上显著优于后者,强调了领域特定模型在评估中的重要性及人类监督的必要性。

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19952 2025-11-26 cs.LG 57%

Hierarchical Spatio-Temporal Attention Network with Adaptive Risk-Aware Decision for Forward Collision Warning in Complex Scenarios

具有自适应风险感知决策的分层时空注意力网络用于复杂场景的前方碰撞预警

Haoran Hu, Junren Shi, Shuo Jiang, Kun Cheng, Xia Yang, Changhao Piao

机构 * organization= School of Automation \& School of Industrial Internet, Chongqing University of Posts organization= Platform Technology Development Department, AVATR Technology Co. LTD. , city= Chongqing , postcode= 400000 , country= China organization= School of Vehicle Mobility, Tsinghua University , city= Beijing , postcode= 100084 , country= China

专题命中 安全评测 :safety(abstract);分类 cs.LG

AI总结 本文提出一种结合分层时空注意力网络和动态风险阈值调整算法的前方碰撞预警框架,通过高效模型和自适应机制提升复杂场景下的预警精度与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19657 2025-11-26 cs.LG 57%

Structured Noise Modeling for Enhanced Time-Series Forecasting

结构噪声建模以提升时间序列预测

Sepideh Koohfar

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

AI总结 本文提出结构化噪声建模框架,通过模糊-去噪机制提升时间序列预测的准确性与稳定性,适用于能源、基础设施等时间敏感领域。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19654 2025-11-26 cs.CR cs.AI 57%

Accuracy and Efficiency Trade-Offs in LLM-Based Malware Detection and Explanation: A Comparative Study of Parameter Tuning vs. Full Fine-Tuning

在基于大语言模型的恶意软件检测与解释中的准确性与效率权衡:参数调优与全微调的比较研究

Stephen C. Gravereaux, Sheikh Rabiul Islam

机构 * Department of Cybersecurity University at Albany - State University of New York Albany, NY, USA(网络安全系 美国纽约州立大学阿尔巴尼分校)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 本研究比较了参数调优与全微调在恶意软件检测中的性能,发现LoRA在保持解释质量的同时显著提升了效率。

Comments Accepted in IEEE Big Data 2025

Journal ref IEEE Big Data 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01265 2025-11-26 cs.CL 57%

AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs

AraFinNews: 基于领域适应大语言模型的阿拉伯语金融摘要

Mo El-Haj, Paul Rayson

机构 * School of Computing(计算学院) Lancaster University(兰卡斯特大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 AraFinNews通过领域适应大语言模型提升阿拉伯语金融文本摘要的连贯性和准确性。

Comments 9 pages

Journal ref IEEE BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15097 2025-11-26 cs.CR cs.AI 57%

MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm

MAIF:基于 artifacts 的代理范式强化 AI 可信度与溯源

Vineeth Sai Narajala, Manish Bhatt, Idan Habler, Ronald F. Del Rosario, Ads Dawson

机构 * Security Researcher Cisco(安全研究员 希思科) Researcher OWASP/Project Kuiper Security(研究员 OWASP/项目 Kuiper 安全) Adversarial AI Security Research Cisco(对抗性AI安全研究 希思科)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 MAIF 通过以 artifacts 为中心的代理范式,解决 AI 的可信度、安全性和问责问题,实现数据的主动信任执行和高效处理。

Comments 7 Pages, 2 Figures, 6 Tables, Repo: https://github.com/vineethsai/maifscratch-1, Added additional Author and fixed Citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20550 2025-11-26 cs.LO 50%

Verifying Numerical Methods with Isabelle/HOL

用Isabelle/HOL验证数值方法

Dustin Bryant, Jonathan Julian Huerta y Munive, Simon Foster

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出基于ITrees的Isabelle/HOL框架,用于验证数值方法,通过形式化规范和自动证明方法,实现从形式化到可执行代码的端到端验证流程。

Comments 30 pages, 30 listings, for accompanying formalisation, see https://zenodo.org/records/17679526

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20218 2025-11-26 cs.CV 50%

Text-guided Controllable Diffusion for Realistic Camouflage Images Generation

基于文本引导的可控扩散生成逼真伪装图像

Yuhang Qian, Haiyan Chen, Wentong Li, Ningzhong Liu, Jie Qin

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出CT-CIG方法,通过文本引导和可控扩散生成逼真且逻辑合理的伪装图像,利用VLM和FIRM模块提升伪装图像质量。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09436 2025-11-26 cs.CV 50%

Scaling up self-supervised learning for improved surgical foundation models

提升自监督学习以改进手术基础模型

Tim J. M. Jaspers, Ronald L. P. D. de Jong, Yiping Li, Carolus H. J. Kusters, Franciscus H. A. Bakker, Romy C. van Jaarsveld, Gino M. Kuiper, Richard van Hillegersberg, Jelle P. Ruurda, Willem M. Brinkman, Josien P. W. Pluim, Peter H. N. de With, Marcel Breeuwer, Yasmina Al Khalil, Fons van der Sommen

机构 * Department of Electrical Engineering, Video Coding \& Architectures, Eindhoven University of Technology, Eindhoven, The Netherlands Department of Biomedical Engineering, Medical Image Analysis, Eindhoven University of Technology, Eindhoven, The Netherlands Department of Surgery, University Medical Center Utrecht, Utrecht, The Netherlands Department of Oncological Urology, University Medical Center Utrecht, Utrecht, The Netherlands Department of Urology, Catharina Hospital, Eindhoven, The Netherlands

专题命中 安全评测 :safety(abstract)

AI总结 本研究提出SurgeNetXL,通过大规模预训练提升手术计算机视觉性能,实现多个任务上的显著改进。

Journal ref Medical Image Analysis, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19526 2025-11-26 cs.CV 50%

Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models

感知分类:评估和引导视觉语言模型中的层次场景推理

Jonathan Lee, Xingrui Wang, Jiawei Peng, Luoxin Ye, Zehan Zheng, Tiezheng Zhang, Tao Wang, Wufei Ma, Siyi Chen, Yu-Cheng Chou, Prakhar Kaushik, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出感知分类基准测试,旨在评估和引导视觉语言模型在层次场景推理中的能力,揭示模型在属性驱动推理上的不足,并展示通过上下文示例提升性能的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏