arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-12 至 2025-09-12 共收录 26 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2509.09055 2025-09-12 cs.CL cs.AI cs.LG 92%

Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M

Piyush Pant

机构 * Saarland University(萨尔兰大学)

专题命中 偏好对齐 :DPO(title,abstract);safety(title,abstract);alignment(abstract);RLHF(abstract)

Comments 17 pages, 3 figures. Code and dataset available at https://github.com/PiyushWithPant/Improving-LLM-Safety-and-Helpfulness-using-SFT-and-DPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20655 2025-09-12 cs.CV cs.CL 83%

Improving Alignment in LVLMs with Debiased Self-Judgment

Sihan Yang, Chenhang Cui, Zihao Zhao, Yiyang Zhou, Weilong Yan, Ying Wei, Huaxiu Yao

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) UNC-Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08302 2025-09-12 cs.CL cs.AI 62%

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

Jiahui Li, Lin Li, Tai-wei Chang, Kun Kuang, Long Chen, Jun Zhou, Cheng Yang

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09121 2025-09-12 cs.CL 57%

Compass-v3: Scaling Domain-Specific LLMs for Multilingual E-Commerce in Southeast Asia

Sophia Maria

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07370 2025-09-12 cs.CL 57%

PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions

Yixuan Tang, Yi Yang, Ahmed Abbasi

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Notre Dame(诺丁汉大学)

专题命中 偏好对齐 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2405.03486 2025-09-12 cs.CR cs.CV cs.SI 78%

UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Yiting Qu, Xinyue Shen, Yixin Wu, Michael Backes, Savvas Zannettou, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA赫尔姆霍茨信息安全中心) TU Delft(代尔夫特理工大学)

专题命中 安全训练 :safety(title,abstract)

Comments To Appear in the ACM Conference on Computer and Communications Security (CCS), October 13, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17813 2025-09-12 cs.RO cs.LG 57%

Safe Multi-Agent Navigation guided by Goal-Conditioned Safe Reinforcement Learning

Meng Feng, Viraj Parimi, Brian Williams

机构 * Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology(计算机科学与人工智能实验室,麻省理工学院)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Due to the limitation "The abstract field cannot be longer than 1,920 characters", the abstract here is shorter than that in the PDF file

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 1 篇

2502.04227 2025-09-12 cs.CR 50%

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks

Andreas Happe, Jürgen Cito

专题命中 越狱攻击 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2506.15850 2025-09-12 cs.LG cs.AI 62%

Uncertainty Estimation by Human Perception versus Neural Models

Pedro Mendes, Paolo Romano, David Garlan

机构 * Software and Societal Systems Department, Carnegie Mellon University(卡内基梅隆大学软件与社会系统部门) INESC-ID and Instituto Superior Técnico, Universidade de Lisboa(里斯本大学INESC-ID和理工学院)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 1 篇

2505.07084 2025-09-12 cs.RO 50%

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models

Shucheng Huang, Freda Shi, Chen Sun, Jiaming Zhong, Minghao Ning, Yufeng Yang, Yukun Lu, Hong Wang, Amir Khajepour

机构 * MVSLab, Department of Mechanical and Mechatronics Engineering, University of Waterloo(滑铁卢大学机械与机电工程系MVSLab) CompLING Lab, David R. Cheriton School of Computer Science, University of Waterloo(滑铁卢大学大卫·R·切里顿计算机科学学院CompLING Lab) Department of Data and Systems Engineering, University of Hong Kong(香港大学数据与系统工程系) Department of Mechanical Engineering, University of New Brunswick(新不伦瑞克大学机械工程系) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动性学院)

专题命中 隐私与版权 :safety(abstract)

Comments This work has been accepted to IEEE Transactions on Vehicular Technology. Please refer to the copyright notice for additional information

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 11 篇

2509.09303 2025-09-12 cs.CL 79%

From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models

Grazia Sveva Ascione, Nicolò Tamagnone

机构 * Department of Industrial Engineering, Polytechnic University of Turin(工业工程系,都灵理工学院) Venice School of Management, Ca Foscari University of Venice(威尼斯管理学院,威尼斯福斯卡里大学)

专题命中 安全评测 :trustworthy(title);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08839 2025-09-12 cs.CY 79%

Evaluating the Clinical Safety of LLMs in Response to High-Risk Mental Health Disclosures

Siddharth Shah, Amit Gupta, Aarav Mann, Alexandre Vaz, Benjamin E. Caldwell, Robert Scholz, Peter Awad, Rocky Allemandi, Doug Faust, Harshita Banka, Tony Rousmaniere

专题命中 安全评测 :safety(title,abstract);分类 cs.CY

Comments Previously posted as a preprint on Research Square (DOI: 10.21203/rs.3.rs-7364128/v1), under a CC BY 4.0 License

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08997 2025-09-12 cs.HC 78%

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models

Yaman Yu, Yiren Liu, Jacky Zhang, Yun Huang, Yang Wang

专题命中 安全评测 :safety(title,abstract)

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08852 2025-09-12 cs.CY cs.AI cs.LG 75%

Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned

Kajetan Schweighofer, Barbara Brune, Lukas Gruber, Simon Schmid, Alexander Aufreiter, Andreas Gruber, Thomas Doms, Sebastian Eder, Florian Mayer, Xaver-Paul Stadlbauer, Christoph Schwald, Werner Zellinger, Bernhard Nessler, Sepp Hochreiter

机构 * TRUSTIFAI GMBH(TRUSTIFAI公司) TÜV AUSTRIA HOLDING AG(TÜV奥地利控股公司) Software Competence Center Hagenberg(哈根贝格软件竞争力中心) Johannes Kepler University Linz - Institute for Machine Learning(林茨约瑟夫·约瑟夫斯大学-机器学习研究所)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 63 pages, 27 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08912 2025-09-12 cs.CY cs.HC 74%

Towards Trustworthy AI: Characterizing User-Reported Risks across LLMs "In the Wild"

Lingyao Li, Renkai Ma, Zhaoqian Xue, Junjie Xiong

专题命中 安全评测 :trustworthy(title);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09397 2025-09-12 cs.CV 67%

Decoupling Clinical and Class-Agnostic Features for Reliable Few-Shot Adaptation under Shift

Umaima Rahman, Raza Imam, Mohammad Yaqub, Dwarikanath Mahapatra

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德人工智能大学) Khalifa University(卡利法大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09593 2025-09-12 cs.CL cs.AI 62%

Fluent but Unfeeling: The Emotional Blind Spots of Language Models

Bangzhao Shu, Isha Joshi, Melissa Karnaze, Anh C. Pham, Ishita Kakkar, Sindhu Kothe, Arpine Hovasapian, Mai ElSherief

机构 * Northeastern University(东北大学) UC San Diego(加州大学圣地亚哥分校) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Camera-ready version for ICWSM 2026. First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09680 2025-09-12 cs.CV cs.CL 57%

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Rongyao Fang, Aldrich Yu, Chengqi Duan, Linjiang Huang, Shuai Bai, Yuxuan Cai, Kun Wang, Si Liu, Xihui Liu, Hongsheng Li

机构 * CUHK(香港中文大学) HKU(香港大学) BUAA(北京航空航天大学) Alibaba(阿里巴巴)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Project page: https://flux-reason-6m.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03700 2025-09-12 cs.HC cs.AI 57%

MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning

Liujian Tang, Shaokang Dong, Yijia Huang, Minqi Xiang, Hongtao Ruan, Bin Wang, Shuo Li, Zhiheng Xi, Zhihui Cao, Hailiang Pang, Heng Kong, He Yang, Mingxu Chai, Zhilin Gao, Xingyu Liu, Yingnan Fu, Jiaming Liu, Xuanjing Huang, Yu-Gang Jiang, Tao Gui, Qi Zhang, Kang Wang, Yunke Zhang, Yuran Wang

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01523 2025-09-12 cs.CL 57%

CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering

Zongxi Li, Yang Li, Haoran Xie, S. Joe Qin

机构 * School of Data Science, Lingnan University(数据科学学院) School of Science and Technology, Hong Kong Metropolitan University(科技学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09311 2025-09-12 cs.CV 50%

Image Recognition with Vision and Language Embeddings of VLMs

Illia Volkov, Nikita Kisel, Klara Janouskova, Jiri Matas

机构 * Visual Recognition Group, Faculty of Electrical Engineering, Czech Technical University in Prague(视觉识别组,电气工程学院,布拉格捷克技术大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 2 篇

2509.08910 2025-09-12 cs.CV cs.AI 83%

PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability

Tung Vu, Lam Nguyen, Quynh Dao

机构 * Posts and Telecommunications Institute of Technology(邮电技术研究所) Hanoi Architectural University(河内建筑大学)

专题命中 AI治理与伦理 :safety(title,abstract);alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08829 2025-09-12 cs.CY cs.AI cs.IR 62%

PerFairX: Is There a Balance Between Fairness and Personality in Large Language Model Recommendations?

Chandan Kumar Sah

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 10 pages, 5 figures. Accepted to the Workshop on Multimodal Continual Learning (MCL) at ICCV 2025. @2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), ICCV's 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 3 篇

2509.09629 2025-09-12 cs.CL 79%

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems

Minghang Zhu, Zhengliang Shi, Zhiwei Xu, Shiguang Wu, Lingjie Wang, Pengjie Ren, Zhaochun Ren, Zhumin Chen

机构 * Shandong University(山东大学) Leiden University(莱顿大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09314 2025-09-12 cs.AI cs.HC 57%

Measuring Implicit Spatial Coordination in Teams: Effects on Collective Intelligence and Performance

Thuy Ngoc Nguyen, Anita Williams Woolley, Cleotilde Gonzalez

机构 * University of Dayton(代顿大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09064 2025-09-12 cs.CV 50%

Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models

Qiuhui Chen, Xuancheng Yao, Huping Ye, Yi Hong

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)

专题命中 其他安全 :alignment(abstract)

Comments Accepted by IEEE Journal of Biomedical and Health Informatics (JBHI)

详情

展开后加载摘要…

URL PDF HTML 收藏