arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2503.08906 2025-12-02 cs.CV cs.AI cs.CL cs.MM 62%

Prompt-OT: An Optimal Transport Regularization Paradigm for Knowledge Preservation in Vision-Language Model Adaptation

Prompt-OT: 一种用于视觉-语言模型适应中知识保留的最优传输正则化范式

Xiwen Chen, Wenhui Zhu, Peijie Qiu, Hao Wang, Huayu Li, Haiyu Wu, Aristeidis Sotiras, Yalin Wang, Abolfazl Razi

机构 * Morgan Stanley(摩根士丹利) Clemson University(克莱姆森大学) Arizona State University(亚利桑那州立大学) Washington University in St. Louis(圣路易斯华盛顿大学) University of Arizona(亚利桑那大学) University of Notre Dame(圣母大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 Prompt-OT通过最优传输正则化提升视觉-语言模型在知识保留和跨领域适应中的性能

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21742 2025-12-01 cs.CL cs.AI 62%

EduMod-LLM: A Modular Approach for Designing Flexible and Transparent Educational Assistants

EduMod-LLM: 一种用于设计灵活且透明教育助手的模块化方法

Meenakshi Mittal, Rishi Khare, Mihran Miroyan, Chancharik Mitra, Narges Norouzi

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 EduMod-LLM通过模块化方法评估教育问答系统各组件性能,揭示失败模式,提升透明度和教学对齐效果。

Comments Proceedings of the AAAI Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19930 2025-11-26 cs.GT cs.CY cs.LG 62%

Designing Reputation Systems for Manufacturing Data Trading Markets: A Multi-Agent Evaluation with Q-Learning and IRL-Estimated Utilities

为制造数据交易市场设计声誉系统:基于Q学习和IRL估计效用的多智能体评估

Kenta Yamamoto, Teruaki Hayashi

机构 * Department of Systems Innovation, Graduate School of Engineering, The University of Tokyo Tokyo, Japan(系统创新部门,工学研究生院,东京大学东京,日本)

专题命中 安全评测 :alignment(abstract);分类 cs.CY、cs.LG

AI总结 本研究通过多智能体模拟器评估了五种声誉系统,发现PeerTrust在数据价格与质量一致性及防止垄断方面表现最佳,并提出混合声誉机制提升市场稳定性。

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19845 2025-11-26 cs.LG cs.CY stat.ML 62%

SX-GeoTree: Self-eXplaining Geospatial Regression Tree Incorporating the Spatial Similarity of Feature Attributions

SX-GeoTree: 自解释地理回归树结合特征归因的空间相似性

Chaogui Kang, Lijian Luo, Qingfeng Guan, Yu Liu

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY、cs.LG

AI总结 SX-GeoTree通过整合空间相似性与模块度最大化,提升地理回归树的解释稳定性与空间残差均匀性。

Comments 41 pages, 7 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17681 2025-11-26 cs.LG cs.AI 62%

Unlearning as Ablation: Toward a Falsifiable Benchmark for Generative Scientific Discovery

反例作为消去:面向生成科学发现的可证伪基准

Robert Yang

机构 * S6 Research(S6研究)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出'反例作为消去'作为可证伪基准,旨在检验AI在科学发现中的生成能力,通过系统移除目标结果并评估模型能否重新推导,推动AI-科学的基准发展。

Comments 6 pages + appendix. Accepted to NeurIPS 2025 AI4Science Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10161 2025-11-26 cs.CL cs.AI 62%

LaajMeter: A Framework for LaaJ Evaluation

LaajMeter:一种用于LaaJ评估的框架

Samuel Ackerman, Gal Amram, Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich, Raviv Gal, Wesam Ibraheem, Avi Ziv

机构 * IBM Research, Israel(IBM以色列研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 LaaJMeter是一种用于LaaJ评估的模拟框架,通过生成合成数据来系统分析评估度量标准,帮助验证特定任务的LaaJ质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08701 2025-11-26 cs.CV cs.AI cs.LG 62%

SafeFix: Targeted Model Repair via Controlled Image Generation

SafeFix: 通过受控图像生成实现目标模型修复

Ouyang Xu, Baoming Zhang, Ruiyu Mao, Yunhui Guo

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 SafeFix通过生成语义忠实的图像来修复模型,提升模型对罕见案例的鲁棒性,减少系统性错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19480 2025-11-26 cs.LG cs.AI 62%

Exploiting the Experts: Unauthorized Compression in MoE-LLMs

利用专家:MoE-LLMs中的未经授权压缩

Pinaki Prasad Guha Neogi, Ahmad Mohammadshirazi, Dheeraj Kulshrestha, Rajiv Ramnath

机构 * Ohio State University(俄亥俄州立大学) Flairsoft

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了MoE-LLMs在特定任务中的可剪枝性,提出专家归因框架和防御策略以防止未经授权的压缩和微调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19324 2025-11-25 cs.IR cs.AI cs.CL 62%

What Drives Cross-lingual Ranking? Retrieval Approaches with Multilingual Language Models

什么推动跨语言排序?具有多语言语言模型的检索方法

Roksana Goworek, Olivia Macmillan-Scott, Eda B. Özyiğit

机构 * The Alan Turing Institute(艾伦·图灵研究所) Queen Mary University of London(伦敦女王学院) University College London(伦敦大学学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了跨语言检索中影响排序的因素,发现专门训练的密集检索模型和对比学习在低资源和跨脚本场景中表现更优,优于传统的翻译和单语言检索方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19257 2025-11-25 cs.CR cs.AI cs.LG 62%

Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation

Medusa: 跨模态可转移的对抗攻击用于多模态医疗检索增强生成

Yingjia Shang, Yi Liu, Huimin Wang, Furong Li, Wenfang Sun, Wu Chengyu, Yefeng Zheng

机构 * Westlake University(西湖大学) Heilongjiang University(黑龙江大学) City University of Hong Kong(香港城市大学) Tencent(腾讯)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

AI总结 Medusa提出了一种针对多模态医疗检索增强生成系统的跨模态可转移对抗攻击方法,通过优化扰动和双循环策略实现高攻击成功率并抵御主流防御措施。

Comments Accepted at KDD 2026 First Cycle (full version). Authors marked with * contributed equally. Yi Liu is the lead author

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18860 2025-11-25 cs.CL cs.AI 62%

Generating Reading Comprehension Exercises with Large Language Models for Educational Applications

利用大语言模型生成阅读理解练习用于教育应用

Xingyu Huang, Fei Jiang, Jianli Xiao

机构 * School of Optical-Electrical and Computer Engineering, University of Shanghai for Science and Technology, Shanghai, China(光学电子与计算机工程学院,上海理工大学,上海,中国) Chongqing Academy of Science and Technology, Chongqing, China(重庆市科学技术院,重庆,中国)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出RCEG框架,利用大语言模型自动生成高质量、个性化的英语阅读理解练习,通过内容候选生成、判别器筛选和综合评估指标提升练习质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09711 2025-11-25 cs.CL cs.AI 62%

PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry

PsychiatryBench: 一个面向心理治疗的多任务基准测试

Aya E. Fouda, Abdelrahamn A. Hassan, Radwa J. Hanafy, Mohammed E. Fouda

机构 * Compumacy for Artificial Intelligence solutions(人工智能解决方案公司) Department of Behavioural Health- Saint Elizabeths Hospital(行为健康部门-圣伊丽莎白医院)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 PsychiatryBench通过权威教科书和病例书构建多任务基准,评估LLMs在心理治疗中的临床一致性与安全性,揭示模型在多轮随访等任务中的不足,推动专门调优和更严谨的评估方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16699 2025-11-24 cs.CL cs.AI 62%

Detecting and Steering LLMs' Empathy in Action

检测和引导大语言模型的同理心行为

Juan P. Cadile

机构 * Department of Philosophy University of Rochester(哲学系罗切斯特大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 研究通过对比提示检测和引导大语言模型的同理心行为,发现不同模型在引导效果和鲁棒性上存在差异。

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16482 2025-11-21 cs.LG cs.AI stat.ML 62%

Correlation-Aware Feature Attribution Based Explainable AI

基于相关性的特征归因可解释人工智能

Poushali Sengupta, Yan Zhang, Frank Eliassen, Sabita Maharjan

机构 * Department of Informatics, University of Oslo, Norway(奥斯陆大学信息学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 ExCIR通过相关性感知的归因方法,提供高效、一致且可扩展的可解释人工智能解决方案。

Comments Accepted, 2026 International Conference on Advances in Artificial Intelligence and Machine Learning (AAIML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00091 2025-11-21 cs.CY cs.AI 62%

A First-Principles Based Risk Assessment Framework and the IEEE P3396 Standard

基于第一性原理的风险评估框架及IEEE P3396标准

Richard J. Tong, Marina Cortês, Jeanine A. DeFalco, Mark Underwood, Janusz Zalewski

机构 * Chair, IEEE Artificial Intelligence Standards Committee (AISC) Institute of Astrophysics Space Sciences, University of Lisbon, Portugal Vice Chair, IEEE Artificial Intelligence Standards Committee University of New Haven Florida Gulf Coast University, United States State Academy of Applied Sciences, Ciechanow, Poland

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文提出基于第一性原理的生成式AI风险评估框架,通过信息分类系统识别不同风险类型并归责于相关方,旨在提升AI治理的严谨性与责任性。

Comments 8 pages with 3 tables. This manuscript is prepared for publication by the Institute of Electrical and Electronics Engineers, Standards Association (IEEE-SA), Sponsor Committee - Artificial Intelligence Standards Committee (C/AISC) as a White Paper of Working Group p3396 at https://standards.ieee.org/ieee/3396/11379/

Journal ref 2025 IEEE Conference on Artificial Intelligence (CAI), pp. 1588-1595, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14010 2025-11-20 cs.CL cs.AI 62%

Knowledge-Grounded Agentic Large Language Models for Multi-Hazard Understanding from Reconnaissance Reports

Chenchen Kuai, Zihao Li, Braden Rosen, Stephanie Paal, Navid Jafari, Jean-Louis Briaud, Yunlong Zhang, Youssef M. A. Hashash, Yang Zhou

机构 * organization= Department One , addressline= Address One , city= City One , postcode= 00000 , state= State One , country= Country One organization= Department Two , addressline= Address Two , city= City Two , postcode= 22222 , state= State Two , country= Country Two organization= Zachry Department of Civil \& Environmental Engineering, Texas A\&M University , addressline= 3136 TAMU , city= College Station , postcode= 77843 , state= TX , country= USA organization= Department of Engineering Technology Industrial Distribution, Texas A\&M University , city= College Station , postcode= 77843 , state= TX , country= USA organization= Department of Civil Environmental Engineering, University of Illinois Urbana-Champaign , city= Urbana , postcode= 61801 , state= IL , country= USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14767 2025-11-20 cs.IR cs.AI cs.CY 62%

An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market

Minh-Thuan Nguyen, Thien Vo-Thanh, Thai-Duy Dinh, Xuan-Quang Phan, Tan-Ha Mai, Lam-Son Lê

机构 * Computer Science Department Vietnamese-German University, Vietnam(越南德意志大学计算机科学系) Business Administration Department FPT University, Vietnam(越南FPT大学商学院) IT Operation Department Mantu Group, Vietnam(越南Mantu集团IT运营部) CSIE Department National Taiwan University, Taiwan(台湾国立台湾大学CSIE系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments Accepted at ACOMPA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17979 2025-11-20 cs.AI cs.CL 62%

Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities

Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng, Yanyan Zhao, Xuda Zhi, Yongbo Huang, Hao He, Wanxiang Che, Ting Liu, Bing Qin

专题命中 安全评测 :harmlessness(abstract);分类 cs.CL、cs.AI

Comments To appear at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13948 2025-11-19 cs.CV cs.CL cs.LG 62%

EchoAgent: Guideline-Centric Reasoning Agent for Echocardiography Measurement and Interpretation

Matin Daghyani, Lyuyang Wang, Nima Hashemi, Bassant Medhat, Baraa Abdelsamad, Eros Rojas Velez, XiaoXiao Li, Michael Y. C. Tsang, Christina Luong, Teresa S. M. Tsang, Purang Abolmaesumi

机构 * University of British Columbia(不列颠哥伦比亚大学) Vancouver General Hospital(温哥华总医院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

Comments 12 pages, Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13809 2025-11-19 cs.LG cs.AI 62%

ScoresActivation: A New Activation Function for Model Agnostic Global Explainability by Design

Emanuel Covaci, Fabian Galis, Radu Balan, Daniela Zaharie, Darian Onchis

机构 * West University of Timișoara, Romania(蒂米șоara西大学) University of Maryland, United States(马里兰大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Paper submitted to ECAI 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13753 2025-11-19 cs.LG cs.AI cs.CR 62%

Robustness of LLM-enabled vehicle trajectory prediction under data security threats

Feilong Wang, Fuqiang Liu

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 20 pages, 2 figures, 11 tables, working paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18646 2025-11-19 cs.AI cs.CL 62%

Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap

Jun Wang, Ninglun Gu, Kailai Zhang, Zijiao Zhang, Yelun Bao, Jin Yang, Xu Yin, Liwei Liu, Yihuan Liu, Pengyong Li, Gary G. Yen, Junchi Yan

机构 * Department of Networks, China Mobile Communications Group Co.,Ltd.(中国移动通信集团有限公司网络部) Xidian University(西安电子科技大学) Oklahoma State University(俄克拉荷马州立大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Preprint. Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13712 2025-11-18 cs.LG cs.AI 62%

From Black Box to Insight: Explainable AI for Extreme Event Preparedness

Kiana Vu, İsmet Selçuk Özer, Phung Lai, Zheng Wu, Thilanka Munasinghe, Jennifer Wei

机构 * Department of Cybersecurity University at Albany, SUNY Albany, NY, USA Dept. of Atmospheric \& Environmental Sciences University at Albany, SUNY Albany, NY, USA Lally School of Management Rensselaer Polytechnic Institute Albany, NY, USA Goddard Space Flight Center NASA Greenbelt, MD, USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04688 2025-11-18 cs.CL cs.LG 62%

Evaluating LLMs' Reasoning Over Ordered Procedural Steps

Adrita Anika, Md Messal Monem Miah

机构 * Amazon(亚马逊公司) Texas A&M University(德克萨斯A&M大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to IJCNLP-AACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11691 2025-11-18 cs.LG cs.AI cs.SD 62%

Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues

Seham Nasr, Zhao Ren, David Johnson

机构 * Center for Cognitive Interaction Technology (CITEC), Bielefeld University, Germany(认知交互技术中心(CITEC),比勒菲尔德大学,德国) Cognitive Systems Lab, University of Bremen, Germany(认知系统实验室,不莱梅大学,德国)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11597 2025-11-18 cs.AI cs.CL 62%

CLINB: A Climate Intelligence Benchmark for Foundational Models

Michelle Chen Huebscher, Katharine Mach, Aleksandar Stanić, Markus Leippold, Ben Gaiarin, Zeke Hausfather, Elisa Rawat, Erich Fischer, Massimiliano Ciaramita, Joeri Rogelj, Christian Buck, Lierni Sestorain Saralegui, Reto Knutti

机构 * University of Miami(迈阿密大学) University of Zurich(苏黎世大学) Stripe(Stripe公司) ETH Zurich(苏黎世联邦理工学院) Imperial College London(伦敦帝国理工学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Questions, system prompt and model judge prompts available here: https://www.kaggle.com/datasets/deepmind/clinb-questions

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11583 2025-11-18 cs.LG cs.AI cs.IR 62%

Parallel and Multi-Stage Knowledge Graph Retrieval for Behaviorally Aligned Financial Asset Recommendations

Fernando Spadea, Oshani Seneviratne

机构 * Rensselaer Polytechnic Institute(伦斯勒理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 3 figures, RAGE-KG 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11066 2025-11-17 cs.CV cs.AI cs.CL 62%

S2D-ALIGN: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report Generation

Jiechao Gao, Chang Liu, Yuangang Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10067 2025-11-17 cs.AI cs.CL 62%

Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning

Yuxuan Zhou, Yubin Wang, Bin Wang, Chen Ning, Xien Liu, Ji Wu, Jianye Hao

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Huawei Noah’s Ark Lab(华为诺亚实验室) College of AI, Tsinghua University(清华大学人工智能学院) Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10002 2025-11-17 cs.CL cs.AI 62%

PustakAI: Curriculum-Aligned and Interactive Textbooks Using Large Language Models

Shivam Sharma, Riya Naik, Tejas Gawas, Heramb Patil, Kunal Korgaonkar

机构 * Department of Computer Science and Information Systems(计算机科学与信息系统系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏