arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-23 至 2025-10-23 共收录 40 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2405.11647 2025-10-23 cs.AI cs.LG 73%

Hummer: Towards Limited Competitive Preference Dataset

Li Jiang, Yusen Wu, Junwu Xiong, Jingqing Ruan, Yichuan Ding, Qingpei Guo, Zujie Wen, Jun Zhou, Xiaotie Deng

机构 * Mila, McGill University(蒙特利尔大学Mila实验室) Peking University(北京大学) Ant Group(蚂蚁集团)

专题命中 偏好对齐 :alignment(abstract);jailbreak(abstract);分类 cs.AI、cs.LG

Journal ref COLM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19050 2025-10-23 cs.AI cs.LG 62%

Rectifying Shortcut Behaviors in Preference-based Reward Learning

Wenqian Ye, Guangtao Zheng, Aidong Zhang

机构 * University of Virginia(弗吉尼亚大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15061 2025-10-23 cs.LG cs.CL 62%

Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models

Samuel Paech, Allen Roush, Judah Goldfeder, Ravid Shwartz-Ziv

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.LG

Comments 11 pages + appendices, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18895 2025-10-23 cs.SE cs.AI cs.HC 57%

CosmoCore Affective Dream-Replay Reinforcement Learning for Code Generation

Santhosh Kumar Ravindran

机构 * Microsoft Corporation(微软公司)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 1 篇

2507.16814 2025-10-23 cs.LG cs.CV 57%

Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning

Junhao Shen, Haiteng Zhao, Yuzhe Gu, Songyang Gao, Kuikun Liu, Haian Huang, Jianfei Gao, Dahua Lin, Wenwei Zhang, Kai Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab)

专题命中 安全训练 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 红队测试 1 篇

2510.19303 2025-10-23 cs.CR cs.AI cs.LG cs.MA cs.SE 62%

Collaborative penetration testing suite for emerging generative AI algorithms

Petar Radanliev

专题命中 红队测试 :red teaming(abstract);分类 cs.AI、cs.LG

Journal ref Appl Intell 55, 1030 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 3 篇

2510.19476 2025-10-23 cs.LG cs.AI 81%

A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring

Julian Schulz

机构 * Meridian Research, Cambridge(梅迪安研究,剑桥)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18918 2025-10-23 cs.CL cs.AI 62%

Misinformation Detection using Large Language Models with Explainability

Jainee Patel, Chintan Bhatt, Himani Trivedi, Thanh Thi Nguyen

机构 * Department of Computer Engineering, LDRP Institute of Technology and Research(计算机工程系,LDRP技术与研究学院) University of Wollongong(沃林根大学) Monash University(莫纳什大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication in the Proceedings of the 8th International Conference on Algorithms, Computing and Artificial Intelligence (ACAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19310 2025-10-23 cs.CL 57%

JointCQ: Improving Factual Hallucination Detection with Joint Claim and Query Generation

Fan Xu, Huixuan Zhang, Zhenliang Zhang, Jiahao Wang, Xiaojun Wan

机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学) Trustworthy Technology and Engineering Laboratory, Huawei(可信技术与工程实验室,华为)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 3 篇

2506.10946 2025-10-23 cs.LG cs.AI cs.CL 67%

GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models

Peizhi Niu, Evelyn Ma, Huiting Zhou, Duo Zhou, Huan Zhang, S. Rasoul Etesami, Olgica Milenkovic

机构 * Department of Electrical & Computer Engineering(电气与计算机工程系) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 隐私与版权 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08727 2025-10-23 cs.LG cs.AI cs.CL cs.IT math.IT 67%

Memorization-Compression Cycles Improve Generalization

Fangyuan Yu

机构 * Thoughtworks

专题命中 隐私与版权 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 12 pages, 6 figures, NeurIPS2025 NEGEL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19036 2025-10-23 cs.CL 57%

From Memorization to Generalization: Fine-Tuning Large Language Models for Biomedical Term-to-Identifier Normalization

Suswitha Pericharla, Daniel B. Hier, Tayo Obafemi-Ajayi

专题命中 隐私与版权 :alignment(abstract);分类 cs.CL

Comments Submitted for publication to BMC BioData Mining

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 15 篇

2510.19384 2025-10-23 cs.LG 79%

Learning Noise-Resilient and Transferable Graph-Text Alignment via Dynamic Quality Assessment

Yuhang Liu, Minglai Shao, Zengyi Wo, Yunlong Chu, Bing Hao, Shengzhong Liu, Ruijie Wang, Jianxin Li

机构 * School of New Media and Communication, Tianjin University(新媒体与传播学院,天津大学) Baidu(百度) Shanghai Jiao Tong University(上海交通大学) School of Computer Science and Engineering, Beihang University(计算机科学与工程学院,北航)

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17452 2025-10-23 cs.CV cs.AI 79%

Training-Free Label Space Alignment for Universal Domain Adaptation

Dujin Lee, Sojung An, Jungmyung Wi, Kuniaki Saito, Donghyun Kim

机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments 22 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19152 2025-10-23 cs.LG 70%

Subliminal Corruption: Mechanisms, Thresholds, and Interpretability

Reya Vir, Sarvesh Bhatnagar

机构 * Department of Computer Science, Columbia University, New York, NY, USA(哥伦比亚大学计算机科学系) Department of Computer Science(计算机科学系) Engineering, University of Michigan, MI, USA(密歇根大学工程学院)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19799 2025-10-23 cs.CY cs.AI cs.HC cs.LG cs.SE econ.GN q-fin.EC 67%

Integrating Transparent Models, LLMs, and Practitioner-in-the-Loop: A Case of Nonprofit Program Evaluation

Ji Ma, Albert Casella

机构 * LBJ School of Public Affairs, The University of Texas at Austin(德克萨斯大学奥斯汀分校劳伦斯·伯克曼公共事务学院) Gradel Institute of Charity, University of Oxford(牛津大学格拉德利慈善研究所) Michael & Susan Dell Foundation(迈克尔与苏珊·德尔基金会)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19955 2025-10-23 cs.LG cs.AI cs.CL 67%

MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research

Hui Chen, Miao Xiong, Yujie Lu, Wei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, Bryan Hooi

机构 * National University of Singapore(新加坡国立大学) University of California, Santa Barbara(加州大学圣巴巴拉分校) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 49 pages, 9 figures. Accepted by NeurIPS 2025 D&B Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19032 2025-10-23 cs.CL cs.CY cs.HC 62%

When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation

Abeer Badawi, Elahe Rahimi, Md Tahmid Rahman Laskar, Sheri Grach, Lindsay Bertrand, Lames Danok, Jimmy Huang, Frank Rudzicz, Elham Dolatabadi

机构 * York University(约克大学) Vector Institute(向量研究所) Dalhousie University(达尔豪斯大学) IWK Health Hospital(IWK健康医院) King’s College London(伦敦国王学院)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14623 2025-10-23 cs.LG cs.AI 62%

LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching

Zhuo Cao, Xuan Zhao, Lena Krieger, Hanno Scharr, Ira Assent

机构 * IAS-8, Forschungszentrum Jülich, Germany(茹里希研究所,德国) Munich Center for Machine Learning (MCML), LMU Munich, Germany(慕尼黑机器学习中心(MCML),慕尼黑大学,德国) Aarhus University, Denmark(奥胡斯大学,丹麦)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted as a poster presentation at NeurIPS 2025. Camera-ready version. 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19616 2025-10-23 cs.CL 57%

PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models

Farhan Farsi, Shayan Bali, Fatemeh Valeh, Parsa Ghofrani, Alireza Pakniat, Kian Kashfipour, Amir H. Payberah

机构 * Amirkabir University of Technology(阿米尔卡比尔技术大学) King’s College London(伦敦国王学院) Politecnico di Milano(米兰理工学院) KTH Royal Institute of Technology(皇家理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19599 2025-10-23 cs.CV cs.AI 57%

XBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography

Haozhe Luo, Shelley Zixin Shu, Ziyu Zhou, Sebastian Otalora, Mauricio Reyes

机构 * ARTORG Center for Biomedical Engineering Research, University of Bern, Switzerland(ARTORG生物医学工程研究中心,伯尔尼大学,瑞士) Shanghai Jiao Tong University, China(上海交通大学,中国) Kaiko.AI, Switzerland(Kaiko.AI,瑞士) Dept. of Radiation Oncology, Inselspital, Bern University Hospital(放射肿瘤科,因斯普尔茨医院,伯尔尼大学医院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19419 2025-10-23 cs.CL 57%

BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models

Yuan Gao, Suchir Salhan, Andrew Caines, Paula Buttery, Weiwei Sun

机构 * ALTA Institute(ALTA研究院) Department of Computer Science & Technology, University of Cambridge(计算机科学与技术系,剑桥大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted Paper at the BabyLM Workshop 2025 @ EMNLP (Presentation in Suzhou, China)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19212 2025-10-23 stat.ME cs.AI 57%

No Intelligence Without Statistics: The Invisible Backbone of Artificial Intelligence

Ernest Fokoué

机构 * School of Mathematics and Statistics(数学与统计学学院) Rochester Institute of Technology(罗切斯特理工学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 37 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19002 2025-10-23 cs.GT cs.LG econ.TH math.OC 57%

Impartial Selection with Predictions

Javier Cembrano, Felix Fischer, Max Klimm

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20228 2025-10-23 cs.CR cs.CL 57%

Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID

Xia Han, Qi Li, Jianbing Ni, Mohammad Zulkernine

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) School of Computing(计算机学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted by TrustCom2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07376 2025-10-23 eess.SY cs.RO cs.SY 50%

AttentionSwarm: Reinforcement Learning with Attention Control Barier Function for Crazyflie Drones in Dynamic Environments

Grik Tadevosyan, Valerii Serpiva, Aleksey Fedoseev, Roohan Ahmed Khan, Demetros Aschu, Faryal Batool, Nickolay Efanov, Artem Mikhaylov, Dzmitry Tsetserukou

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Intelligent Space Robotics Laboratory(智能空间机器人实验室) Skolkovo Institute of Science and Technology(斯克尔科维科学与技术学院) AI Center(人工智能中心) Moscow Institute of Physics and Technology(莫斯科物理技术学院)

专题命中 安全评测 :safety(abstract)

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09814 2025-10-23 cs.GR cs.CV cs.SD eess.AS 50%

Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis

Zeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao, Chuan Lin, Baoquan Chen, Libin Liu

机构 * School of Electronics Engineering and Computer Science, Peking University(电子工程与计算机科学学院,北京大学) School of Computer Science, Peking University(计算机学院,北京大学) Renmin University of China(中国人民大学) Shandong University(山东大学) Peking University(北京大学) State Key Lab of General AI(通用人工智能国家重点实验室)

专题命中 安全评测 :alignment(abstract)

Comments SIGGRAPH 2024 (Journal Track); Project page: https://pku-mocca.github.io/Semantic-Gesticulator-Page

Journal ref ACM Transactions on Graphics (TOG) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 6 篇

2510.19008 2025-10-23 cs.HC cs.AI cs.LG cs.MA 73%

Plural Voices, Single Agent: Towards Inclusive AI in Multi-User Domestic Spaces

Joydeep Chandra, Satyam Kumar Navneet

机构 * BNRIST, Tsinghua University(北京理工大学、清华大学) Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19327 2025-10-23 cs.MA cs.AI 70%

SORA-ATMAS: Adaptive Trust Management and Multi-LLM Aligned Governance for Future Smart Cities

Usama Antuley, Shahbaz Siddiqui, Sufian Hameed, Waqas Arif, Subhan Shah, Syed Attique Shah

机构 * organization= Department of Computer Science, National University of Computer \& Emerging Sciences , addressline= St-4 Sector 17-D On National Highway , city= Karachi , postcode= 75160 , state= , country= Pakistan organization= Balochistan University of Information Technology, Engineering organization= Department of Computer Science, Birmingham City University , addressline= STEAMhouse, Belmont Row , city= Birmingham , postcode= B4 7RQ , country= United Kingdom

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15831 2025-10-23 cs.CL cs.AI cs.CY 67%

Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs

Vishnu Hari, Kalpana Panda, Srikant Panda, Amit Agarwal, Hitesh Laxmichand Patel

机构 * Birla Institute of Technology and Science (BITS)(巴拉·技术与科学学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏