arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-20 至 2025-11-20 共收录 34 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2507.20964 2025-11-20 cs.AI cs.CC cs.GT cs.LG cs.MA 84%

Core Safety Values for Provably Corrigible Agents

Aran Nayebi

机构 * Aran Nayebi(独立研究者)

专题命中 偏好对齐 :safety(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG

Comments 14 pages. To appear in AAAI 2026 Machine Ethics Workshop (W37) Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17184 2025-11-20 cs.CL 79%

Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models

Xudong Han, Junjie Yang, Tianyang Wang, Ziqian Bi, Xinyuan Song, Junfeng Hao, Junhao Song

机构 * Department of Informatics, University of Sussex(信息学院,苏塞克斯大学) Pingtan Research Institute, Xiamen University(平潭研究院,厦门大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Department of Computer Science, Purdue University(计算机科学系,普渡大学) Department of Computer Science, Emory University(计算机科学系,埃默里大学) AI Agent Lab, Vokram Group(AI代理实验室,Vokram集团) Department of Computing, Imperial College London(计算系,帝国理工学院伦敦分校)

专题命中 偏好对齐 :alignment(title);safety(abstract);分类 cs.CL

Comments 24 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03527 2025-11-20 cs.CL cs.LG 62%

FANAL -- Financial Activity News Alerting Language Modeling Framework

Urjitkumar Patel, Fang-Chun Yeh, Chinmay Gondhalekar, Hari Nalluri

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted for the IEEE International Workshop on Large Language Models for Finance, 2024. This is a preprint version

Journal ref 2024 IEEE International Conference on Big Data, Dec 15-18, 2024, Electronic ISBN: 979-8-3503-6248-0, Electronic ISSN: 2573-2978

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15573 2025-11-20 cs.CY 57%

Two-Faced Social Agents: Context Collapse in Role-Conditioned Large Language Models

Vikram K Suresh

专题命中 偏好对齐 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15038 2025-11-20 cs.SD cs.AI eess.AS 57%

Aligning Generative Music AI with Human Preferences: Methods and Challenges

Dorien Herremans, Abhinaba Roy

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

Comments Accepted at the AAAI-2026 Senior Member Track

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2511.15208 2025-11-20 cs.LG 57%

Reasoning in Diffusion Large Language Models is Concentrated in Dynamic Confusion Zones

Ranfei Chen, Ming Chen, Kaifei Wang

机构 * Institute of Computing Technology Chinese Academy of Sciences Beijing, China(计算技术研究所中国科学院北京)

专题命中 安全训练 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15358 2025-11-20 cs.RO 50%

Platform-Agnostic Reinforcement Learning Framework for Safe Exploration of Cluttered Environments with Graph Attention

Gabriele Calzolari, Vidya Sumathy, Christoforos Kanellakis, George Nikolakopoulos

机构 * Robotics and AI Group, Department of Computer Science, Electrical and Space Engineering, Luleå University of Technology(机器人与人工智能组,计算机科学、电气与空间工程系,吕勒奥技术大学)

专题命中 安全训练 :safety(abstract)

Comments 8 pages, 6 figures, submitted to the 2026 IEEE International Conference on Robotics & Automation

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 1 篇

2511.15316 2025-11-20 cs.CV 50%

What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNs

Zhihan Ren, Lijun He, Jiaxi Liang, Xinzhu Fu, Haixia Bi, Fan Li

机构 * Xi’an Jiaotong University(西安交通大学)

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 3 篇

2503.15511 2025-11-20 cs.HC cs.CY cs.LG 62%

The Trust Calibration Maturity Model for Characterizing and Communicating Trustworthiness of AI Systems

Scott T Steinmetz, Asmeret Naugle, Paul Schutte, Matt Sweitzer, Alex Washburne, Lisa Linville, Daniel Krofcheck, Michal Kucer, Samuel Myren

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CY、cs.LG

Comments 19 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15005 2025-11-20 cs.CL cs.AI 62%

Mathematical Analysis of Hallucination Dynamics in Large Language Models: Uncertainty Quantification, Advanced Decoding, and Principled Mitigation

Moses Kiprono

机构 * Catholic University of America(美国天主教大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments 10 pages, theoretical/mathematical LLM research, no figures, intended for peer-reviewed journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01840 2025-11-20 cs.LG 57%

Optimizing In-Context Learning for Efficient Full Conformal Prediction

Weicao Deng, Sangwoo Park, Min Li, Osvaldo Simeone

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 2 篇

2511.14964 2025-11-20 cs.CY cs.AI cs.HC 73%

How Should the Law Treat Future AI Systems? Fictional Legal Personhood versus Legal Identity

Heather J. Alexander, Jonathan A. Simon, Frédéric Pinard

专题命中 隐私与版权 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments 69 pages. Forthcoming in Case Western Journal of Law, Technology & the Internet (publication offer date september 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15614 2025-11-20 cs.RO cs.AI 57%

Optimus-Q: Utilizing Federated Learning in Adaptive Robots for Intelligent Nuclear Power Plant Operations through Quantum Cryptography

Optimus-Q:利用联邦学习在自适应机器人中进行智能核电站操作的量子密码学应用

Sai Puppala, Ismail Hossain, Jahangir Alam, Sajedul Talukder

机构 * University of Texas at El Paso(德克萨斯理工大学) Southern Illinois University Carbondale(南方伊利诺伊大学卡本代尔分校)

专题命中 隐私与版权 :safety(abstract);分类 cs.AI

AI总结 Optimus-Q机器人结合联邦学习和量子密码学,实现智能核电站的污染监测与安全提升。

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 12 篇

2509.01418 2025-11-20 cs.CL 79%

On the Alignment of Large Language Models with Global Human Opinion

Yang Liu, Masahiro Kaneko, Chenhui Chu

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments 28 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04759 2025-11-20 cs.AI 79%

Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented Reasoning

Tianhui Cai, Yifan Liu, Zewei Zhou, Haoxuan Ma, Seth Z. Zhao, Zhiwen Wu, Xu Han, Zhiyu Huang, Jiaqi Ma

专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15206 2025-11-20 cs.CR cs.IT math.IT 78%

Trustworthy GenAI over 6G: Integrated Applications and Security Frameworks

Bui Duc Son, Trinh Van Chien, Dong In Kim

专题命中 安全评测 :trustworthy(title,abstract)

Comments 8 pages, 5 figures. Submitted for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14805 2025-11-20 cs.SE cs.AI 70%

Towards Continuous Assurance with Formal Verification and Assurance Cases

Dhaminda B. Abeywickrama, Michael Fisher, Frederic Wheeler, Louise Dennis

机构 * Department of Computer Science, The University of Manchester(曼彻斯特大学计算机科学系) Regulatory Support Directorate, Amentum(Amentum监管支持部门)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14010 2025-11-20 cs.CL cs.AI 62%

Knowledge-Grounded Agentic Large Language Models for Multi-Hazard Understanding from Reconnaissance Reports

Chenchen Kuai, Zihao Li, Braden Rosen, Stephanie Paal, Navid Jafari, Jean-Louis Briaud, Yunlong Zhang, Youssef M. A. Hashash, Yang Zhou

机构 * organization= Department One , addressline= Address One , city= City One , postcode= 00000 , state= State One , country= Country One organization= Department Two , addressline= Address Two , city= City Two , postcode= 22222 , state= State Two , country= Country Two organization= Zachry Department of Civil \& Environmental Engineering, Texas A\&M University , addressline= 3136 TAMU , city= College Station , postcode= 77843 , state= TX , country= USA organization= Department of Engineering Technology Industrial Distribution, Texas A\&M University , city= College Station , postcode= 77843 , state= TX , country= USA organization= Department of Civil Environmental Engineering, University of Illinois Urbana-Champaign , city= Urbana , postcode= 61801 , state= IL , country= USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14767 2025-11-20 cs.IR cs.AI cs.CY 62%

An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market

Minh-Thuan Nguyen, Thien Vo-Thanh, Thai-Duy Dinh, Xuan-Quang Phan, Tan-Ha Mai, Lam-Son Lê

机构 * Computer Science Department Vietnamese-German University, Vietnam(越南德意志大学计算机科学系) Business Administration Department FPT University, Vietnam(越南FPT大学商学院) IT Operation Department Mantu Group, Vietnam(越南Mantu集团IT运营部) CSIE Department National Taiwan University, Taiwan(台湾国立台湾大学CSIE系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments Accepted at ACOMPA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17979 2025-11-20 cs.AI cs.CL 62%

Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities

Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng, Yanyan Zhao, Xuda Zhi, Yongbo Huang, Hao He, Wanxiang Che, Ting Liu, Bing Qin

专题命中 安全评测 :harmlessness(abstract);分类 cs.CL、cs.AI

Comments To appear at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15203 2025-11-20 cs.CR cs.AI 57%

Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks

Zimo Ji, Xunguang Wang, Zongjie Li, Pingchuan Ma, Yudong Gao, Daoyuan Wu, Xincheng Yan, Tian Tian, Shuai Wang

机构 * The Hong Kong University of Science and Technology(香港科技大学) Zhejiang University of Technology(浙江工业大学) Lingnan University(岭南大学) School of Cyber Science and Engineering, Southeast University(东南大学计算机科学与工程学院) ZTE Corporation(中兴通讯有限公司)

专题命中 安全评测 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14903 2025-11-20 cs.LG cs.SE 57%

It's LIT! Reliability-Optimized LLMs with Inspectable Tools

Ruixin Zhang, Jon Donnelly, Zhicheng Guo, Ghazal Khalighinejad, Haiyang Huang, Alina Jade Barnett, Cynthia Rudin

机构 * Department of Computer Science(计算机科学系) Duke University(杜克大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on Multi-Turn Interactions in Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14439 2025-11-20 cs.CL 57%

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

Jinru Ding, Lu Lu, Chao Ding, Mouxiao Bian, Jiayuan Chen, Wenrao Pang, Ruiyao Chen, Xinwei Peng, Renjie Lu, Sijie Ren, Guanxu Zhu, Xiaoqin Wu, Zhiqiang Liu, Rongzhao Zhang, Luyi Jiang, Bing Han, Yunqiu Wang, Jie Xu

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12527 2025-11-20 cs.LG stat.ML 57%

Selective Risk Certification for LLM Outputs via Information-Lift Statistics: PAC-Bayes, Robustness, and Skeleton Design

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

机构 * Department of Computer Science Iowa State University(计算机科学系爱荷华州立大学) Department of Civil, Construction and Environmental Engineering Iowa State University(土木、建设与环境工程系爱荷华州立大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15308 2025-11-20 cs.CV 50%

Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Yilin Di, Joao F. Henriques, Daniel Cremers

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) Technical University of Munich(慕尼黑技术大学) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

专题命中 安全评测 :alignment(abstract)

Comments This paper builds upon and extends our earlier conference paper Text2Loc presented at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 1 篇

2508.20230 2025-11-20 cs.LG 57%

Coresets from Trajectories: Selecting Data via Correlation of Loss Differences

Manish Nagaraj, Deepak Ravikumar, Kaushik Roy

机构 * Purdue University(普渡大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Journal ref Transactions on Machine Learning Research 2025, issn=2835-8856

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 8 篇

2502.05934 2025-11-20 cs.AI cs.CC cs.GT cs.LG cs.MA 84%

Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis

Aran Nayebi

机构 * Aran Nayebi(独立研究者)

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG

Comments 21 pages, 1 figure, 1 table. To appear in AAAI 2026 Special Track on AI Alignment (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02505 2025-11-20 cs.CV cs.AI 57%

ESA: Energy-Based Shot Assembly Optimization for Automatic Video Editing

Yaosen Chen, Wei Wang, Tianheng Zheng, Xuming Wen, Han Yang, Yanru Zhang

机构 * Sobey Media Intelligence Laboratory(索贝媒体智能实验室) University of Electronic Science and Technology of China(电子科学与技术大学) SiChuan University(四川大学) Qinghai Normal University(青海师范大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00275 2025-11-20 cs.CV cs.AI 57%

AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding

Md Asaduzzaman Jabin, Hanqi Jiang, Yiwei Li, Patrick Kaggwa, Eugene Douglass, Juliet N. Sekandi, Tianming Liu

机构 * University of Georgia(佐治亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: 7th International Workshop on Large Scale Holistic Video Understanding: Toward Video Foundation Models

Journal ref Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15118 2025-11-20 cs.CV 50%

Unbiased Semantic Decoding with Vision Foundation Models for Few-shot Segmentation

Jin Wang, Bingfeng Zhang, Jian Pang, Weifeng Liu, Baodi Liu, Honglong Chen

机构 * School of Control Science and Engineering, China University of Petroleum (East China)(控制科学与工程学院,中国石油大学(华东))

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏