arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-18 至 2025-09-18 共收录 38 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2405.13541 2025-09-18 cs.CL cs.AI cs.LG 78%

Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts

Yuu Jinnai, Ukyo Honda

机构 * CyberAgent / Tokyo, Japan(CyberAgent)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP Findings, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15495 2025-09-18 cs.SE 67%

SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion

Dongjun Yu, Xiao Yan, Zhenrui Li, Jipeng Xiao, Haochuan He, Yongda Yu, Hao Zhang, Guoping Rong, Xiaobo Huang

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01658 2025-09-18 cs.LG cs.AI cs.IR 62%

CoPL: Collaborative Preference Learning for Personalizing LLMs

Youngbin Choi, Seunghyuk Cho, Minjong Lee, MoonJeong Park, Yesong Ko, Jungseul Ok, Dongwoo Kim

机构 * Graduate School of Artificial Intelligence, POSTECH(POSTECH人工智能研究生院) Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

Comments 19pages, 13 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13869 2025-09-18 cs.CL 57%

Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs

Yang Liu, Chenhui Chu

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments 38 pages, 31 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06652 2025-09-18 cs.CL 57%

IntrEx: A Dataset for Modeling Engagement in Educational Conversations

Xingwei Tan, Mahathi Parvatham, Chiara Gambi, Gabriele Pergola

机构 * Department of Computer Science, University of Warwick, UK(沃里克大学计算机科学系) School of Computer Science, University of Sheffield, UK(谢菲尔德大学计算机科学学院) Department of Psychology, University of Warwick, UK(沃里克大学心理学系)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

Comments EMNLP 2025 Findings camera-ready, 9+7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 3 篇

2509.13164 2025-09-18 cs.RO cs.SY eess.SY 78%

TeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving

Jiawei Wang, Haowei Sun, Xintao Yan, Shuo Feng, Jun Gao, Henry X. Liu

机构 * University of Michigan(密歇根大学) SaferDrive AI The University of Hong Kong(香港大学) Tsinghua University(清华大学) NVIDIA(英伟达)

专题命中 安全训练 :safety(title,abstract)

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13760 2025-09-18 cs.CV 67%

Iterative Prompt Refinement for Safer Text-to-Image Generation

Jinwoo Jeon, JunHyeok Oh, Hayeong Lee, Byung-Jun Lee

机构 * Korea University(韩国大学)

专题命中 安全训练 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14129 2025-09-18 cs.LG cs.CY 62%

Breaking the Cycle of Incarceration With Targeted Mental Health Outreach: A Case Study in Machine Learning for Public Policy

Kit T. Rodolfa, Erika Salomon, Jin Yao, Steve Yoder, Robert Sullivan, Kevin McGuire, Allie Dickinson, Rob MacDougall, Brian Seidler, Christina Sung, Claire Herdeman, Rayid Ghani

机构 * Machine Learning Department(机器学习部门) Heinz College of Public Policy(公共政策学院) Carnegie Mellon University(卡内基梅隆大学) Center for Data Science and Public Policy(数据科学与公共政策中心) University of Chicago(芝加哥大学)

专题命中 安全训练 :safety(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2503.09334 2025-09-18 cs.CR cs.AI 83%

CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning

Adel ElZemity, Budi Arief, Shujun Li

机构 * University of Kent(肯特大学)

专题命中 越狱攻击 :safety(title,abstract);prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12824 2025-09-18 cs.IR 50%

DiffHash: Text-Guided Targeted Attack via Diffusion Models against Deep Hashing Image Retrieval

Zechao Liu, Zheng Zhou, Xiangkun Chen, Tao Liang, Dapeng Lang

专题命中 越狱攻击 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 2 篇

2505.09974 2025-09-18 cs.CR cs.AI 89%

Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data

Adel ElZemity, Budi Arief, Shujun Li

机构 * University of Kent(肯特大学)

专题命中 红队测试 :safety(title,abstract);alignment(abstract);red teaming(abstract);prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13597 2025-09-18 cs.CR cs.AI 57%

Agentic JWT: A Secure Delegation Protocol for Autonomous AI Agents

Abhishek Goswami

专题命中 红队测试 :prompt injection(abstract);分类 cs.AI

Comments 17 pages, 6 figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 4 篇

2509.13919 2025-09-18 cs.CV 78%

Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration

Yuanchen Wu, Ke Yan, Shouhong Ding, Ziyin Zhou, Xiaoqiang Li

机构 * School of Computer Engineering(计算机工程学院) Key Laboratory of Multimedia Trusted Perception(多媒体可信感知关键实验室) Efficient Computing, Xiamen University(高效计算,厦门大学) Tencent Youtu Lab(腾讯优图实验室)

专题命中 幻觉与事实性 :alignment(title,abstract)

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13702 2025-09-18 cs.CL cs.AI 62%

DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models

Xiao Zheng

机构 * School of Computing and Technology(计算机学院) China University of Petroleum(中国石油大学) Qingdao(青岛)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13334 2025-09-18 cs.AI cs.LG 62%

FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness

Anand Swaroop, Akshat Nallani, Saksham Uboweja, Adiliia Uzdenova, Michael Nguyen, Kevin Zhu, Sunishchal Dev, Ashwinee Panda, Vasu Sharma, Maheep Chaudhary

机构 * Algoverse AI Research(Algoverse AI研究院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09206 2025-09-18 cs.RO cs.SY eess.SY 50%

Occupancy-aware Trajectory Planning for Autonomous Valet Parking in Uncertain Dynamic Environments

Farhad Nawaz, Faizan M. Tariq, Sangjae Bae, David Isele, Avinash Singh, Nadia Figueroa, Nikolai Matni, Jovin D'sa

机构 * Honda Research Institute (HRI)(本田研究院) University of Pennsylvania(宾夕法尼亚大学)

专题命中 幻觉与事实性 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 17 篇

2509.13339 2025-09-18 cs.AI 89%

Position: AI Safety Must Embrace an Antifragile Perspective

Ming Jin, Hyunin Lee

专题命中 安全评测 :safety(title,abstract);AI safety(title,abstract);alignment(abstract);分类 cs.AI

Journal ref Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18325 2025-09-18 cs.AI cs.LG 84%

Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary

Licheng Pan, Yongqi Tong, Xin Zhang, Xiaolu Zhang, Jun Zhou, Zhixuan Chu

机构 * The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院) Language and Machine Intelligence Department, Ant Group(蚂蚁集团语言与机器智能部门)

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08380 2025-09-18 cs.AI cs.LG 81%

Co-Investigator AI: The Rise of Agentic AI for Smarter, Trustworthy AML Compliance Narratives

Prathamesh Vasudeo Naik, Naresh Kumar Dintakurthi, Zhanghao Hu, Yue Wang, Robby Qiu

专题命中 安全评测 :trustworthy(title);alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13569 2025-09-18 cs.CL 79%

Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12

John Mendonça, Lining Zhang, Rahul Mallidi, Alon Lavie, Isabel Trancoso, Luis Fernando D'Haro, João Sedoc

机构 * INESC-ID Instituto Superior Técnico - University of Lisbon(理工学院 - 里斯本大学) Department of Technology, Operations, and Statistics, New York University(技术、运营与统计系,纽约大学) Speech Technology and Machine Learning Group - Universidad Politécnica de Madrid(语音技术与机器学习小组 - 马德里理工大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments DSTC12 Track 1 Overview Paper. https://chateval.org/dstc12

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19546 2025-09-18 cs.CL cs.AI 79%

Language Models Identify Ambiguities and Exploit Loopholes

Jio Choi, Mohit Bansal, Elias Stengel-Eskin

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 camera-ready; Code: https://github.com/esteng/ambiguous-loophole-exploitation

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15871 2025-09-18 cs.CY cs.AI cs.CL 75%

A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcare

Manar Aljohani, Jun Hou, Sindhura Kommu, Xuan Wang

机构 * Virginia Tech(维吉尼亚理工大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17514 2025-09-18 cs.AI 74%

TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment

Athanasios Davvetas, Xenia Ziouvelou, Ypatia Dami, Alexios Kaponis, Konstantina Giouvanopoulou, Michael Papademas

专题命中 安全评测 :trustworthy(title);分类 cs.AI

Comments 9 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23804 2025-09-18 cs.CL cs.AI cs.LG 67%

Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequencies

Terrance Liu, Shuyi Wang, Daniel Preotiuc-Pietro, Yash Chandarana, Chirag Gupta

机构 * Carnegie Mellon University(卡内基梅隆大学) Bloomberg(彭博)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10970 2025-09-18 cs.LG cs.AI 62%

The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models

Joshua Au Yeung, Jacopo Dalmasso, Luca Foschini, Richard JB Dobson, Zeljko Kraljevic

机构 * King’s College Hospital(国王学院医院) Nuraxi AI Dev and Doc: AI for Healthcare(Nuraxi AI 人工智能医疗) Sage Bionetworks(Sage 生物网络) University College London(伦敦大学学院) King’s College London(国王学院伦敦)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13680 2025-09-18 cs.SE cs.AI 57%

Prompt Stability in Code LLMs: Measuring Sensitivity across Emotion- and Personality-Driven Variations

Wei Ma, Yixiao Yang, Jingquan Ge, Xiaofei Xie, Lingxiao Jiang

机构 * Singapore Management University(新加坡国立大学) Capital Normal University(首都师范大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13369 2025-09-18 eess.SY cs.CY cs.HC cs.SY 57%

Right-to-Override for Critical Urban Control Systems: A Deliberative Audit Method for Buildings, Power, and Transport

Rashid Mushkani

专题命中 安全评测 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08364 2025-09-18 cs.AI 57%

Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation

Enci Zhang, Xingang Yan, Wei Lin, Tianxiang Zhang, Qianchun Lu

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 14 pages, 3 figs

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17671 2025-09-18 cs.MA cs.AI 57%

ComfyGPT: A Self-Optimizing Multi-Agent System for Comprehensive ComfyUI Workflow Generation

Oucheng Huang, Yuhang Ma, Zeng Zhao, Mingrui Wu, Jiayi Ji, Rongsheng Zhang, Zhipeng Hu, Xiaoshuai Sun, Rongrong Ji

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15022 2025-09-18 cs.CL 57%

Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics

Amalie Brogaard Pauli, Isabelle Augenstein, Ira Assent

机构 * Department of Computer Science, Aarhus University, Denmark(计算机科学系,奥胡斯大学) Department of Computer Science, University of Copenhagen, Denmark(计算机科学系,哥本哈根大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted at EMNLP Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏