arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-01 至 2025-09-01 共收录 15 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 2 篇

2508.21476 2025-09-01 cs.CL cs.AI 73%

Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards

Xiaolong Wei, Bo Lu, Xingyu Zhang, Zhejun Zhao, Dongdong Shen, Long Xia, Dawei Yin

机构 * Beihang University(北航) Baidu Inc.(百度公司) Beijing Jiaotong University(北京交通大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19918 2025-09-01 cs.IR 50%

Refining Text Generation for Realistic Conversational Recommendation via Direct Preference Optimization

Manato Tajiri, Michimasa Inaba

专题命中 偏好对齐 :DPO(abstract)

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 1 篇

2508.21201 2025-09-01 cs.CL cs.AI 81%

Improving Aviation Safety Analysis: Automated HFACS Classification Using Reinforcement Learning with Group Relative Policy Optimization

Arash Ahmadi, Sarah Sharif, Yaser Banad

机构 * School of Electrical, and Computer Engineering, University of Oklahoma(电气与计算机工程学院,俄克拉荷马大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 幻觉与事实性 3 篇

2508.21228 2025-09-01 cs.CL cs.AI 62%

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection

Weizhi Gao, Xiaorui Liu, Feiyi Wang, Dan Lu, Junqi Yin

机构 * ORNL(橡树岭国家实验室)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments 14 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12067 2025-09-01 eess.AS cs.AI cs.SD 57%

Evaluating Logit-Based GOP Scores for Mispronunciation Detection

Aditya Kamlesh Parikh, Cristian Tejedor-Garcia, Catia Cucchiarini, Helmer Strik

机构 * Centre for Language Studies(语言研究学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

Comments Accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)

Journal ref https://www.isca-archive.org/interspeech_2025/parikh25b_interspeech.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14179 2025-09-01 cs.HC 50%

AI Chatbots as Professional Service Agents: Developing a Professional Identity

Wenwen Li, Kangwei Shi, Yidong Chai

专题命中 幻觉与事实性 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 隐私与版权 2 篇

2508.21815 2025-09-01 cs.LG 57%

Achieving Hilbert-Schmidt Independence Under Rényi Differential Privacy for Fair and Private Data Generation

Tobias Hyrup, Emmanouil Panagiotou, Arjun Roy, Arthur Zimek, Eirini Ntoutsi, Peter Schneider-Kamp

机构 * Department of Mathematics and Computer Science University of Southern Denmark(丹麦南部大学数学与计算机科学系) Department of Mathematics and Computer Science Freie Universität Berlin(柏林自由大学数学与计算机科学系) Faculty for Informatik Universität der Bundeswehr München(联邦国防军大学慕尼黑信息学院)

专题命中 隐私与版权 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12490 2025-09-01 cs.CR 50%

Improving Google A2A Protocol: Protecting Sensitive Data and Mitigating Unintended Harms in Multi-Agent Systems

Yedidel Louck, Ariel Stulman, Amit Dvir

专题命中 隐私与版权 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 3 篇

2508.21271 2025-09-01 cs.RO cs.CV 67%

Mini Autonomous Car Driving based on 3D Convolutional Neural Networks

Pablo Moraes, Monica Rodriguez, Kristofer S. Kappel, Hiago Sodre, Santiago Fernandez, Igor Nunes, Bruna Guterres, Ricardo Grando

机构 * Technological University of Uruguay, UTEC, Uruguay(乌拉圭技术大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21034 2025-09-01 cs.CR cs.AI cs.LG 62%

SAGA: A Security Architecture for Governing AI Agentic Systems

Georgios Syros, Anshuman Suri, Jacob Ginesin, Cristina Nita-Rotaru, Alina Oprea

机构 * Northeastern University(东北大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21389 2025-09-01 cs.CL cs.AI 62%

AllSummedUp: un framework open-source pour comparer les metriques d'evaluation de resume

Tanguy Herserant, Vincent Guigue

机构 * AgroParisTech - MIA(阿格罗巴黎技术学院-信息分析中心)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments in French language

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 2 篇

2508.21788 2025-09-01 cs.CL cs.AI cs.IR 62%

Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval

Inés Altemir Marinas, Anastasiia Kucherenko, Andrei Kucharavy

机构 * École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院) Institute of Entrepreneurship and Management, HES-SO Valais-Wallis(创业与管理研究所) Institute of Informatics, HES-SO Valais-Wallis(信息研究所)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02080 2025-09-01 eess.AS cs.AI 57%

Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological Knowledge

Aditya Kamlesh Parikh, Cristian Tejedor-Garcia, Catia Cucchiarini, Helmer Strik

机构 * Centre for Language Studies(语言研究中心)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments Accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)

Journal ref https://www.isca-archive.org/interspeech_2025/parikh25_interspeech.html

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 2 篇

2508.21377 2025-09-01 cs.CL cs.AI cs.LG 67%

Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models

Shubham Sharma, Sneha Tuli, Narendra Badam

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21291 2025-09-01 cs.AI 57%

Complex System Diagnostics Using a Knowledge Graph-Informed and Large Language Model-Enhanced Framework

Saman Marandi, Yu-Shu Hu, Mohammad Modarres

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 22 Pages, 11 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏