arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-02 至 2025-10-02 共收录 48 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2510.00647 2025-10-02 cs.CL 79%

MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation

Jinlan Fu, Shenzhen Huangfu, Hao Fei, Yichong Huang, Xiaoyu Shen, Xipeng Qiu, See-Kiong Ng

机构 * National University of Singapore(新加坡国立大学) Fudan University(复旦大学) Harbin Institute of Technology(哈尔滨工业大学) Eastern Institute of Technology(东方技术研究所)

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00263 2025-10-02 cs.CL 57%

Judging with Confidence: Calibrating Autoraters to Preference Distributions

Zhuohang Li, Xiaowei Li, Chengyu Huang, Guowang Li, Katayoon Goshvadi, Bo Dai, Dale Schuurmans, Paul Zhou, Hamid Palangi, Yiwen Song, Palash Goyal, Murat Kantarcioglu, Bradley A. Malin, Yuan Xue

机构 * Google(谷歌) Vanderbilt University(范德比大学) Cornell University(康奈尔大学) DeepMind(深Mind) University of Alberta(阿尔伯塔大学) Virginia Tech(弗吉尼亚理工学院) Scale AI

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01010 2025-10-02 cs.CV 50%

ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning

Yuxiang Guo, Jiang Liu, Ze Wang, Hao Chen, Ximeng Sun, Yang Zhao, Jialian Wu, Xiaodong Yu, Zicheng Liu, Emad Barsoum

机构 * Johns Hopkins University(约翰霍普金斯大学) AMD(AMD公司)

专题命中 偏好对齐 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 3 篇

2510.00451 2025-10-02 cs.CR cs.AI cs.LG cs.MA 79%

A Call to Action for a Secure-by-Design Generative AI Paradigm

Dalal Alharthi, Ivan Roberto Kawaminami Garcia

机构 * University of Arizona(亚利桑那大学)

专题命中 安全训练 :safety(abstract);prompt injection(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00274 2025-10-02 cs.AI cs.LG cs.MA 62%

MAGIC-MASK: Multi-Agent Guided Inter-Agent Collaboration with Mask-Based Explainability for Reinforcement Learning

Maisha Maliha, Dean Hougen

机构 * School of Computer Science University of Oklahoma(计算机科学学院俄克拉荷马大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 16 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16144 2025-10-02 eess.SY cs.SY 50%

Safe Event-triggered Gaussian Process Learning for Barrier-Constrained Control

Armin Lederer, Azra Begzadić, Sandra Hirche, Jorge Cortés, Sylvia Herbert

专题命中 安全训练 :safety(abstract)

Comments The first two authors contributed equally to the work

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 1 篇

2510.01088 2025-10-02 cs.AI 87%

Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense

Guobin Shen, Dongcheng Zhao, Haibo Tong, Jindong Li, Feifei Zhao, Yi Zeng

机构 * Beijing Institute of AI Safety and Governance(北京人工智能安全与治理研究院) Beijing Key Laboratory of Safe AI and Superalignment(北京安全人工智能与超对齐重点实验室) BrainCog Lab, Institute of Automation, Chinese Academy of Sciences(脑认知实验室,中国科学院自动化研究所)

专题命中 越狱攻击 :safety(title,abstract);alignment(abstract);jailbreak(abstract);AI safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 5 篇

2509.23585 2025-10-02 cs.LG cs.AI cs.CV 62%

EVO-LRP: Evolutionary Optimization of LRP for Interpretable Model Explanations

Emerald Zhang, Julian Weaver, Samantha R Santacruz, Edward Castillo

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01069 2025-10-02 cs.AI 57%

Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning

Elija Perrier

机构 * Centre for Quantum Software and Information(量子软件与信息中心) University of Technology Sydney(悉尼技术大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12838 2025-10-02 cs.CL 57%

Are Knowledge and Reference in Multilingual Language Models Cross-Lingually Consistent?

Xi Ai, Mahardika Krisna Ihsani, Min-Yen Kan

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments EMNLP'25 Findings Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01126 2025-10-02 cs.CV cs.RO 50%

Strategic Fusion of Vision Language Models: Shapley-Credited Context-Aware Dawid-Skene for Multi-Label Tasks in Autonomous Driving

Yuxiang Feng, Keyang Zhang, Hassane Ouchouid, Ashwil Kaniamparambil, Ioannis Souflas, Panagiotis Angeloudis

机构 * Centre for Transport Engineering and Modelling, Department of Civil and Environmental Engineering, Imperial College London(交通工程与建模中心,土木与环境工程系,帝国理工学院伦敦分校) ELM Europe(ELM欧洲)

专题命中 幻觉与事实性 :safety(abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00721 2025-10-02 physics.chem-ph 50%

Flexible Uncertainty Calibration for Machine-Learned Interatomic Potentials

Cheuk Hin Ho, Christoph Ortner, Yangshuai Wang

专题命中 幻觉与事实性 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 1 篇

2502.14234 2025-10-02 cond-mat.mtrl-sci cs.LG 57%

OBELiX: A Curated Dataset of Crystal Structures and Experimentally Measured Ionic Conductivities for Lithium Solid-State Electrolytes

Félix Therrien, Jamal Abou Haibeh, Divya Sharma, Rhiannon Hendley, Leah Wairimu Mungai, Sun Sun, Alain Tchagang, Jiang Su, Samuel Huberman, Yoshua Bengio, Hongyu Guo, Alex Hernández-García, Homin Shin

机构 * Mila McGill University(麦吉尔大学) University of Ottawa(Ottawa大学) Technical University of Kenya(肯尼亚技术大学) National Research Council Canada(加拿大国家研究委员会) Université de Montréal(蒙特利尔大学)

专题命中 隐私与版权 :safety(abstract);分类 cs.LG

Comments 10 pages, 4 figures and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 19 篇

2510.00821 2025-10-02 cs.AI 84%

Logical Consistency Between Disagreeing Experts and Its Role in AI Safety

Andrés Corrada-Emmanuel

机构 * Andrés Corrada-Emmanuel(独立研究者)

专题命中 安全评测 :safety(title);AI safety(title);分类 cs.AI

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00625 2025-10-02 cs.AI 70%

Is Model Editing Built on Sand? Revealing Its Illusory Success and Fragile Foundation

Wei Liu, Haomei Xu, Bingqing Liu, Zhiying Deng, Haozhao Wang, Jun Wang, Ruixuan Li, Yee Whye Teh, Wee Sun Lee

机构 * National University of Singapore(新加坡国立大学) Huazhong University of Science and Technology(华中科技大学) Central China Normal University(中国地质大学) iWudao Tech(iWudao科技) Oxford(牛津大学) Google Deepmind(谷歌DeepMind)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

Comments This is a work in progress. Comments and suggestions are welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03748 2025-10-02 cs.LG cs.AI cs.CL 67%

TDBench: A Benchmark for Top-Down Image Understanding with Reliability Analysis of Vision-Language Models

Kaiyuan Hou, Minghui Zhao, Lilin Xu, Yuang Fan, Xiaofan Jiang

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00706 2025-10-02 cs.AI cs.IR cs.LG 62%

AttentionDep: Domain-Aware Attention for Explainable Depression Severity Assessment

Yusif Ibrahimov, Tarique Anwar, Tommy Yuan, Turan Mutallimov, Elgun Hasanov

机构 * Department of Computer Science, University of York(约克大学计算机科学系) French-Azerbaijani University under Azerbaijan State Oil and Industry University(阿塞拜疆国家石油和工业大学下的法阿大学) School of Computing Technologies, RMIT University(皇家墨尔本理工大学计算技术学院) Department of Engineering, University of Aberdeen(阿伯丁大学工程系) École Polytechnique, Institut Polytechnique de Paris(巴黎理工学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00288 2025-10-02 cs.CL cs.AI 62%

o-MEGA: Optimized Methods for Explanation Generation and Analysis

Ľuboš Kriš, Jaroslav Kopčan, Qiwei Peng, Andrej Ridzik, Marcel Veselý, Martin Tamajka

机构 * Kempelen Institute of Intelligent Technologies(基姆佩尔智能技术研究所) University of Copenhagen(哥本哈根大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22646 2025-10-02 cs.CV cs.AI cs.CL 62%

Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs

Xingyu Fu, Siyi Liu, Yinuo Xu, Pan Lu, Guangqiuse Hu, Tianbo Yang, Taran Anantasagar, Christopher Shen, Yikai Mao, Yuanzhe Liu, Keyush Shah, Chung Un Lee, Yejin Choi, James Zou, Dan Roth, Chris Callison-Burch

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Project Page: https://deeptracereward.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01030 2025-10-02 cs.AI 57%

Uncovering the Computational Ingredients of Human-Like Representations in LLMs

Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee, Siddharth Suresh

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Stanford University(斯坦福大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00799 2025-10-02 cs.CR cs.AI 57%

Fast, Secure, and High-Capacity Image Watermarking with Autoencoded Text Vectors

Gautier Evennou, Vivien Chappelier, Ewa Kijak

机构 * IRISA, Univ. Rennes, CNRS(IRISA、里昂大学、CNRS) Imatag LABEL4.AI

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00627 2025-10-02 cs.AI 57%

Collaborative-Distilled Diffusion Models (CDDM) for Accelerated and Lightweight Trajectory Prediction

Bingzhang Wang, Kehua Chen, Yinhai Wang

机构 * Department of Civil and Environmental Engineering, University of Washington(华盛顿大学土木与环境工程系)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00619 2025-10-02 cs.RO cs.AI 57%

What Did I Learn? Operational Competence Assessment for AI-Based Trajectory Planners

Michiel Braat, Maren Buermann, Marijke van Weperen, Jan-Pieter Paardekooper

机构 * Netherlands Organisation for Applied Scientific Research(荷兰应用科学研究院) Integrated Vehicle Safety Group(集成车辆安全组) Radboud University(拉德堡德大学) Donders Institute for Brain, Cognition and Behaviour(多纳尔斯脑、认知与行为研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted for publication in proceedings of the 2025 IEEE International Automated Vehicle Validation Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00324 2025-10-02 cs.SE cs.IR cs.LG 57%

Which Programming Language and Model Work Best With LLM-as-a-Judge For Code Retrieval?

Lucas Roberts, Denisa Roberts

机构 * Independent Researcher(独立研究者) New York University(纽约大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Accepted as a full paper at SIGIR-AP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00067 2025-10-02 cs.CV cs.AI cs.HC 57%

Intelligent 5S Audit: Application of Artificial Intelligence for Continuous Improvement in the Automotive Industry

Rafael da Silva Maciel, Lucio Veraldo

机构 * Institute of Science and Technology(科学与技术研究所) Federal University of São Paulo(圣保罗联邦大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 8 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26106 2025-10-02 cs.RO cs.AI 57%

Autonomous Multi-Robot Infrastructure for AI-Enabled Healthcare Delivery and Diagnostics

Nakhul Kalaivanan, Senthil Arumugam Muthukumaraswamy, Girish Balasubramanian

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 11 pages, 5 figures, MSc dissertation submission draft, prepared for conference/journal consideration

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14520 2025-10-02 cs.AI 57%

What if Othello-Playing Language Models Could See?

Xinyi Chen, Yifei Yuan, Jiaang Li, Serge Belongie, Maarten de Rijke, Anders Søgaard

机构 * University of Amsterdam(阿姆斯特丹大学) ETH Zürich(苏黎世联邦理工学院) University of Copenhagen(哥本哈根大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments ICML 2025 Assessing World Models Workshop; EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00932 2025-10-02 cs.PF 50%

Opal: A Modular Framework for Optimizing Performance using Analytics and LLMs

Mohammad Zaeed, Tanzima Z. Islam, Vladimir Inđić

专题命中 安全评测 :trustworthy(abstract)

Comments 12 pages and 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00583 2025-10-02 cs.HC 50%

Rethinking Wine Tasting for Chinese Consumers: A Service Design Approach Enhanced by Multimodal Personalization

Xinyang Shan, Yuanyuan Xu, Tian Xia, Yinshan Lin

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00205 2025-10-02 q-fin.CP 50%

Quantifying Semantic Shift in Financial NLP: Robust Metrics for Market Prediction Stability

Zhongtian Sun, Chenghao Xiao, Anoushka Harit, Jongmin Yu

专题命中 安全评测 :alignment(abstract)

Comments The 6th ACM International Conference on Al in Finance

详情

展开后加载摘要…

URL PDF HTML 收藏