arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-13 至 2025-08-13 共收录 36 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2508.08509 2025-08-13 cs.CL cs.AI 86%

Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression

Jadie Adams, Brian Hu, Emily Veenhuis, David Joy, Bharadwaj Ravichandran, Aaron Bray, Anthony Hoogs, Arslan Basharat

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);harmlessness(abstract);分类 cs.CL、cs.AI

Comments AIES '25: Proceedings of the 2025 AAAI/ACM Conference on AI, Ethics, and Society

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08466 2025-08-13 cs.CL 83%

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints

Daren Yao, Jinsong Yuan, Ruike Chen

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08550 2025-08-13 cs.SD cs.CL 79%

Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization

Chaoqun Cui, Liangbin Huang, Shijing Wang, Zhe Tong, Zhaolong Huang, Xiao Zeng, Xiaofeng Liu

机构 * Alibaba Digital Media and Entertainment Group(阿里巴巴数字媒体与娱乐集团) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing Jiaotong University(北京交通大学交通数据挖掘与具身智能重点实验室)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments This paper is accepted by ACL2025 (Main)

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025: 4524-4546

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05496 2025-08-13 cs.AI 70%

InfiAlign: A Scalable and Sample-Efficient Framework for Aligning LLMs to Enhance Reasoning Capabilities

Shuo Cai, Su Lu, Qi Zhou, Kejing Yang, Zhijie Sang, Congkai Xie, Hongxia Yang

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2501.06208 2025-08-13 cs.CL 86%

Enhancing AI Safety Through the Fusion of Low Rank Adapters

Satya Swaroop Gudipudi, Sreeram Vipparla, Harpreet Singh, Shashwat Goel, Ponnurangam Kumaraguru

机构 * IIIT Hyderabad(海德拉巴国家理工学院) NSUT Delhi(德里NSUT)

专题命中 安全训练 :safety(title,abstract);AI safety(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08926 2025-08-13 cs.AI 79%

Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language Models

Wei Cai, Jian Zhao, Yuchu Jiang, Tianle Zhang, Xuelong Li

机构 * Peking University(北京大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) Northwestern Polytechnical University(西北工业大学) Southeast University(东南大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 1 篇

2412.02141 2025-08-13 cs.CV cs.CL 57%

WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image

Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Jipeng Zhang, Xiangjian He, Song Wu, Xiaohan Xing, Sen Yang, Xiyue Wang, Linlin Shen

机构 * Shenzhen University(深圳大学) University of Nottingham Ningbo China(诺丁汉大学宁波分校) City University of Hong Kong(香港城市大学) Stanford University(斯坦福大学) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

Comments ICCV 2025, 38 pages, 22 figures, 35 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 2 篇

2508.09085 2025-08-13 cs.NI cs.AI cs.LG 62%

Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring

Zihan Fang, Zheng Lin, Senkang Hu, Yihang Tao, Yiqin Deng, Xianhao Chen, Yuguang Fang

机构 * Hong Kong JC STEM Lab of Smart City and Department of Computer Science, City University of Hong Kong(香港JC STEM实验室及城市大学计算机科学系)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06795 2025-08-13 cs.CL cs.CV 57%

From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models

Yuying Shang, Xinyi Zeng, Yutao Zhu, Xiao Yang, Zhengwei Fang, Jingyuan Zhang, Jiawei Chen, Zinan Liu, Yu Tian

机构 * University of Chinese Academy of Sciences(中国科学院大学) Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University(计算机科学与技术系,人工智能研究院,清华大学) Gaoling School of Artificial Intelligence, Renmin University of China(人工智能学院,中国人民大学) Kuaishou Technology Inc.(快手科技有限公司) Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University(多信息处理重点实验室,华东师范大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 10 篇

2501.13983 2025-08-13 cs.CL cs.AI 81%

AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models

Yang Fan

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments There are serious academic problems in this paper, such as data falsification and plagiarism in the method of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02078 2025-08-13 cs.CV cs.AI cs.LG 81%

From Lab to Field: Real-World Evaluation of an AI-Driven Smart Video Solution to Enhance Community Safety

Shanle Yao, Babak Rahimi Ardabili, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Christopher Neff, Lauren Bourque, Hamed Tabkhi

机构 * Department of Electrical and Computer Engineering, University of North Carolina at Charlotte(电气与计算机工程系,北卡罗来纳大学夏洛特分校) Department of Public Policy, University of North Carolina at Charlotte(公共政策系,北卡罗来纳大学夏洛特分校)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08273 2025-08-13 cs.CL cs.LG 81%

TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning

Kristian Miok, Blaz Škrlj, Daniela Zaharie, Marko Robnik Šikonja

机构 * Faculty of Computer and Information Science, University of Ljubljana, Slovenia(卢布尔雅那大学计算机与信息科学学院) ICAM - Advanced Environmental Research Institute, West University of Timisoara, Romania(蒂米șoara西大学先进环境研究所) Department of Computer Science, West University of Timisoara, Romania(蒂米șoa拉西大学计算机科学系)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08777 2025-08-13 cs.IR cs.AI cs.LG 62%

Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge

Francesco Fabbri, Gustavo Penha, Edoardo D'Amico, Alice Wang, Marco De Nadai, Jackie Doremus, Paul Gigioli, Andreas Damianou, Oskar Stal, Mounia Lalmas

机构 * Spotify

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted at RecSys '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08629 2025-08-13 cs.CY cs.AI 62%

Securing Educational LLMs: A Generalised Taxonomy of Attacks on LLMs and DREAD Risk Assessment

Farzana Zahid, Anjalika Sewwandi, Lee Brandon, Vimal Kumar, Roopak Sinha

机构 * University of Waikato(怀卡托大学) Deakin University(迪金大学)

专题命中 安全评测 :jailbreak(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08277 2025-08-13 cs.CL cs.LG 62%

Objective Metrics for Evaluating Large Language Models Using External Data Sources

Haoze Du, Richard Li, Edward Gehringer

机构 * Department of Computer Science(计算机科学系) North Carolina State University(北卡罗来纳州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments This version of the paper is lightly revised from the EDM 2025 proceedings for the sake of clarity

Journal ref EDM 2025 Palermo, Italy, July, 2025, pp. 489-495. International Educational Data Mining Society (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06555 2025-08-13 cs.CV cs.CY cs.MA 57%

StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback

Hongbo Ma, Fei Shen, Hongbin Xu, Xiaoce Wang, Gang Xu, Jinkai Zheng, Liangqiong Qu, Ming Li

专题命中 安全评测 :alignment(abstract);分类 cs.CY

Comments 24pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10066 2025-08-13 cs.MM cs.CV 50%

LayLens: Improving Deepfake Understanding through Simplified Explanations

Abhijeet Narang, Parul Gupta, Liuyijia Su, Abhinav Dhall

机构 * Monash University(墨尔本大学)

专题命中 安全评测 :trustworthy(abstract)

Comments Accepted to ACM ICMI 2025 Demos

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19316 2025-08-13 cs.MA 50%

Making Teams and Influencing Agents: Efficiently Coordinating Decision Trees for Interpretable Multi-Agent Reinforcement Learning

Rex Chen, Stephanie Milani, Zhicheng Zhang, Norman Sadeh, Fei Fang

专题命中 安全评测 :safety(abstract)

Comments 17 pages; 2 tables; 12 figures; accepted version, published at the 8th AAAI/ACM Conference on AI, Ethics and Society (AIES '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11002 2025-08-13 cs.SD cs.MM eess.AS 50%

Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation

Yan Rong, Shan Yang, Chenxing Li, Dong Yu, Li Liu

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 8 篇

2508.05387 2025-08-13 cs.LG cs.AI 76%

Echo: Decoupling Inference and Training for Large-Scale RL Alignment on Heterogeneous Swarms

Jie Xiao, Changyuan Fan, Qingnan Ren, Alfred Long, Yuchen Zhang, Rymon Yu, Eric Yang, Lynn Ai, Shaoduo Gan

机构 * Peking University(北京大学)

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07623 2025-08-13 cs.CL 70%

Optimizing Class-Level Probability Reweighting Coefficients for Equitable Prompting Accuracy

Ruixi Lin, Yang You

机构 * Department of Computer Science(计算机科学系) National University of Singapore(新加坡国立大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08804 2025-08-13 cs.LG cs.AI 62%

TechOps: Technical Documentation Templates for the AI Act

Laura Lucaj, Alex Loosley, Hakan Jonsson, Urs Gasser, Patrick van der Smagt

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08544 2025-08-13 cs.CY cs.AI 62%

AI Agents and the Law

Mark O. Riedl, Deven R. Desai

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 2025 AAAI Conference on AI, Ethics, and Society

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05269 2025-08-13 cs.LG cs.AI q-bio.QM 62%

Chemist-aligned retrosynthesis by ensembling diverse inductive bias models

Krzysztof Maziarz, Guoqing Liu, Hubert Misztela, Austin Tripp, Junren Li, Aleksei Kornev, Piotr Gaiński, Holger Hoefling, Mike Fortunato, Rishi Gupta, Marwin Segler

机构 * Microsoft Research AI for Science(微软研究院人工智能与科学研究中心) Novartis Biomedical Research(诺华生物医学研究) University of Cambridge(剑桥大学) Jagiellonian University(雅盖隆大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08333 2025-08-13 cs.CY cs.AI 62%

Normative Moral Pluralism for AI: A Framework for Deliberation in Complex Moral Contexts

David-Doron Yaacov

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Conference version: AIES 2025 (non-archival track), 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09019 2025-08-13 cs.AI 57%

Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs

Shivam Dubey

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08262 2025-08-13 cs.CL 57%

Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models

Alaa Alhamzeh, Mays Al Rebdawi

机构 * First Author Affiliation(第一作者机构) Second Author Affiliation(第二作者机构)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments 8 pages, 4 figures, Passau uni, Master thesis in NLP

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 9 篇

2508.08504 2025-08-13 cs.CY cs.AI cs.LG 82%

When the Domain Expert Has No Time and the LLM Developer Has No Clinical Expertise: Real-World Lessons from LLM Co-Design in a Safety-Net Hospital

Avni Kothari, Patrick Vossler, Jean Digitale, Mohammad Forouzannia, Elise Rosenberg, Michele Lee, Jennee Bryant, Melanie Molina, James Marks, Lucas Zier, Jean Feng

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12833 2025-08-13 cs.CV cs.AI cs.LG 62%

SPIE: Semantic and Structural Post-Training of Image Editing Diffusion Models with AI feedback

Elior Benarous, Yilun Du, Heng Yang

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08678 2025-08-13 cs.CY 57%

Exploring Large Language Model Agents for Piloting Social Experiments

Jinghua Piao, Yuwei Yan, Nian Li, Jun Zhang, Yong Li

专题命中 其他安全 :alignment(abstract);分类 cs.CY

Comments Accepted by COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏