arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-06 至 2025-11-06 共收录 28 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 2 篇

2509.06733 2025-11-06 cs.AI cs.CL cs.IR 73%

Reinforcement Learning Foundations for Deep Research Systems: A Survey

Wenjun Li, Zhi Chen, Jingru Lin, Hannan Cao, Wei Han, Sheng Liang, Zhi Zhang, Kuicai Dong, Dexun Li, Chen Zhang, Yong Liu

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments 39 pages, second version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19248 2025-11-06 cs.LG 70%

Inference-Time Reward Hacking in Large Language Models

Hadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju, Flavio du Pin Calmon

机构 * Harvard University(哈佛大学)

专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2025 (Spotlight Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 4 篇

2511.03247 2025-11-06 cs.CR cs.LG 81%

Death by a Thousand Prompts: Open Model Vulnerability Analysis

Amy Chang, Nicholas Conley, Harish Santhanalakshmi Ganesan, Adam Swanda

机构 * Cisco AI Threat Research & Security(思科AI威胁研究与安全)

专题命中 安全训练 :alignment(abstract);safety(abstract);jailbreak(abstract);prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02875 2025-11-06 cs.CY cs.AI 62%

Academics and Generative AI: Empirical and Epistemic Indicators of Policy-Practice Voids

R. Yamamoto Ravenor

机构 * Tokyo Women’s Medical University(东京女子医科大学)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.CY

Comments 14 pages, 2 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03106 2025-11-06 cs.AI 57%

Large language models require a new form of oversight: capability-based monitoring

Katherine C. Kellogg, Bingyang Ye, Yifan Hu, Guergana K. Savova, Byron Wallace, Danielle S. Bitterman

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02997 2025-11-06 cs.AI 57%

Evaluating Control Protocols for Untrusted AI Agents

Jon Kutasov, Chloe Loughridge, Yuqi Sun, Henry Sleight, Buck Shlegeris, Tyler Tracy, Joe Benton

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2511.03271 2025-11-06 cs.CR cs.CL 79%

Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs

Yize Liu, Yunyun Hou, Aina Sui

专题命中 越狱攻击 :jailbreak(title);red teaming(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03434 2025-11-06 cs.HC cs.AI cs.MA cs.NI cs.SI 57%

Inter-Agent Trust Models: A Comparative Study of Brief, Claim, Proof, Stake, Reputation and Constraint in Agentic Web Protocol Design-A2A, AP2, ERC-8004, and Beyond

Botao 'Amber' Hu, Helena Rong

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

Comments Submitted to AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 1 篇

2503.11519 2025-11-06 cs.CV cs.CL 57%

Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models

Hao Cheng, Erjia Xiao, Yichi Wang, Lingfeng Zhang, Qiang Zhang, Jiahang Cao, Kaidi Xu, Mengshu Sun, Xiaoshuai Hao, Jindong Gu, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Oxford(牛津大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) The Hong Kong University of Science and Technology(香港科学与技术大学) Beijing University of Technology(北京工业大学) Tsinghua University(清华大学) City University of Hong Kong(香港城市大学)

专题命中 提示注入 :prompt injection(abstract);分类 cs.CL

Comments This paper is accepted by IJCAI2025 Workshop on Deepfake Detection, Localization, and Interpretability as Best Student Paper

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 3 篇

2510.12839 2025-11-06 cs.CL cs.AI cs.CE cs.CY 67%

FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs

Yingjia Wan, Haochen Tan, Xiao Zhu, Xinyu Zhou, Zhiwei Li, Qingsong Lv, Changxuan Sun, Jiaqi Zeng, Yi Xu, Jianqiao Lu, Yinhong Liu, Zhijiang Guo

机构 * UCLA(美国大学洛杉矶分校) CUHK(香港中文大学) HKUST (GZ)(香港科技大学(广州)) HKUST(香港科技大学) Tsinghua University(清华大学) ECNU(华东师范大学) NVIDIA(英伟达) UCL(伦敦大学学院) HKU(香港大学) University of Cambridge(剑桥大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments EMNLP 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23845 2025-11-06 cs.CL 57%

Read Your Own Mind: Reasoning Helps Surface Self-Confidence Signals in LLMs

Jakub Podolak, Rajeev Verma

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

Comments Presented at UncertaiNLP Workshop at EMNLP 2025 https://aclanthology.org/2025.uncertainlp-main.21.pdf

Journal ref UncertaiNLP Workshop at Empirical Methods in Natural Language Processing 2025 (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03166 2025-11-06 cs.CL 57%

Measuring Aleatoric and Epistemic Uncertainty in LLMs: Empirical Evaluation on ID and OOD QA Tasks

Kevin Wang, Subre Abdoul Moktar, Jia Li, Kangshuo Li, Feng Chen

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments Accepted by UDM-KDD'24

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 7 篇

2211.01446 2025-11-06 cs.LG 74%

Trustworthy Representation Learning via Information Funnels and Bottlenecks

João Machado de Freitas, Bernhard C. Geiger

机构 * Christian Doppler Laboratory for Dependable Intelligent Systems in Harsh Environments(可信智能系统在恶劣环境中的克里斯蒂安·多普勒实验室) Graz University of Technology(格拉茨技术大学) Know Center Research GmbH(Know Center研究有限责任公司)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Published in Machine Learning (Springer), vol. 114, no. 12, Article 267, 2025

Journal ref Mach Learn 114, 267 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03641 2025-11-06 cs.CR cs.AI cs.CL cs.CY 67%

Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology

Thomas Souverain

机构 * Department of AI Ethics, CEA Paris-Saclay(人工智能伦理系,CEA巴黎萨克雷)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 17 pages, 2 Tables and 2 Pictures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21861 2025-11-06 cs.LG cs.AI cs.CL 67%

The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems

Bentley DeVilling

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 18 pages, 2 figures. Category: cs.LG. Code and data: https://github.com/Course-Correct-Labs/mirror-loop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03051 2025-11-06 cs.AI cs.IR 57%

No-Human in the Loop: Agentic Evaluation at Scale for Recommendation

Tao Zhang, Kehui Yao, Luyi Ma, Jiao Chen, Reza Yousefi Maragheh, Kai Zhao, Jianpeng Xu, Evren Korpeoglu, Sushant Kumar, Kannan Achan

机构 * Walmart Global Tech(沃尔玛全球技术)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 4 page, NeurIPS 2025 Workshop: Evaluating the Evolving LLM Lifecycle

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20462 2025-11-06 cs.AI 57%

TAMO: Fine-Grained Root Cause Analysis via Tool-Assisted LLM Agent with Multi-Modality Observation Data in Cloud-Native Systems

Xiao Zhang, Qi Wang, Mingyi Li, Yuan Yuan, Mengbai Xiao, Fuzhen Zhuang, Dongxiao Yu

机构 * School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院) Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07955 2025-11-06 cs.HC 50%

Implementation Considerations for Automated AI Grading of Student Work

Zewei Tian, Alex Liu, Lief Esbenshade, Shawon Sarkar, Zachary Zhang, Kevin He, Min Sun

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02996 2025-11-06 cs.CV 50%

SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics

Ailar Mahdizadeh, Puria Azadi Moghadam, Xiangteng He, Shahriar Mirabbasi, Panos Nasiopoulos, Leonid Sigal

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能向量研究所)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 2 篇

2511.01885 2025-11-06 cs.AI cs.LG q-bio.NC 81%

Mirror-Neuron Patterns in AI Alignment

Robyn Wyrick

机构 * Department of Computer Science University of Bath, United Kingdom(计算机科学系 英国巴斯大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 51 pages, Masters thesis. 10 tables, 7 figures, project data & code here: https://github.com/robynwyrick/mirror-neuron-frog-and-toad

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17228 2025-11-06 cs.CY cs.AI 62%

Survey on AI Ethics: A Socio-technical Perspective

Dave Mbiazi, Meghana Bhange, Maryam Babaei, Ivaxi Sheth, Patrik Kenfack, Samira Ebrahimi Kahou

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments Updated to the peer-reviewed version accepted and published in Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)

Journal ref Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 7 篇

2511.03280 2025-11-06 cs.LG stat.AP 79%

A Probabilistic Approach to Pose Synchronization for Multi-Reference Alignment with Applications to MIMO Wireless Communication Systems

Rob Romijnders, Gabriele Cesa, Christos Louizos, Kumar Pratik, Arash Behboodi

机构 * University of Amsterdam(阿姆斯特丹大学) QUvA-Lab(QUvA实验室) Qualcomm AI Research(高通人工智能研究)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments To appear in NeurIPS workshop: AI and ML for Next-Generation Wireless Communications (AI4NextG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11857 2025-11-06 cs.CL 79%

Post Persona Alignment for Multi-Session Dialogue Generation

Yi-Pei Chen, Noriki Nishida, Hideki Nakayama, Yuji Matsumoto

机构 * RIKEN AIP(日本理化学研究所Advanced Institute for Physical Research) The University of Tokyo(东京大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03559 2025-11-06 cs.CL cs.AI 62%

AILA--First Experiments with Localist Language Models

Joachim Diederich

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03159 2025-11-06 cs.LG cs.AI 62%

CoTox: Chain-of-Thought-Based Molecular Toxicity Reasoning and Prediction

Jueon Park, Yein Park, Minju Song, Soyon Park, Donghyeon Lee, Seungheun Baek, Jaewoo Kang

机构 * Department of Computer Science and Engineering(计算机科学与工程系) AIGEN Sciences(AIGEN公司)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to IEEE BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05588 2025-11-06 cs.RO cs.AI cs.LG 62%

Deep Learning Warm Starts for Trajectory Optimization on the International Space Station

Somrita Banerjee, Abhishek Cauligi, Marco Pavone

机构 * Apple(苹果公司) Johns Hopkins University(约翰霍普金斯大学) Stanford University(斯坦福大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to 2025 International Conference on Space Robotics (iSpaRo). Presented at RSS 2025 Workshop on Space Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03410 2025-11-06 cs.CL 57%

Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG

Longpeng Qiu, Ting Li, Shuai Mao, Nan Yang, Xiaohui Yan

机构 * University of Chinese Academy of Sciences(中国科学院大学) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15737 2025-11-06 cs.LG cs.CV 57%

Probability Density from Latent Diffusion Models for Out-of-Distribution Detection

Joonas Järve, Karl Kaspar Haavel, Meelis Kull

机构 * Institute of Computer Science, University of Tartu, Estonia(计算机科学研究所,塔尔图大学,爱沙尼亚)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments ECAI 2025

Journal ref Frontiers in Artificial Intelligence and Applications 413 (ECAI 2025) 5027 - 5034

详情

展开后加载摘要…

URL PDF HTML 收藏