arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-24 至 2025-07-24 共收录 32 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2503.02832 2025-07-24 cs.CL cs.AI cs.LG 87%

AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation

Songming Zhang, Xue Zhang, Tong Zhang, Bojie Hu, Yufeng Chen, Jinan Xu

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation, (Beijing Jiaotong University), Ministry of Education(大数据与人工智能交通运输联合实验室,(北京交通大学)教育部) School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China(计算机科学与技术学院,北京交通大学,北京,中国) Tencent Inc, China(腾讯公司,中国)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2025 Main Conference, code available at: https://github.com/songmzhang/AlignDistil

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16951 2025-07-24 cs.CL 86%

Harnessing RLHF for Robust Unanswerability Recognition and Trustworthy Response Generation in LLMs

Shuyuan Lin, Lei Duan, Philip Hughes, Yuxuan Sheng

机构 * Sichuan University of Science and Engineering(四川理工学院) Zagazig University(扎加齐亚大学)

专题命中 偏好对齐 :RLHF(title,abstract);trustworthy(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05236 2025-07-24 cs.SD cs.AI cs.LG eess.AS 81%

Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance

Shehzeen Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Subhankar Ghosh, Mikyas T. Desta, Roy Fejgin, Rafael Valle, Jason Li

机构 * NVIDIA Corporation(英伟达公司)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Journal ref ICML Workshop on Machine Learning for Audio, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13574 2025-07-24 cs.IR cs.AI 57%

Enhancing Sequential Recommender with Large Language Models for Joint Video and Comment Recommendation

Bowen Zheng, Zihan Lin, Enze Liu, Chen Yang, Enyang Bai, Cheng Ling, Wayne Xin Zhao, Ji-Rong Wen

机构 * Renmin University of China(中国人民大学) Kuaishou Technology Co., Ltd.(快手科技有限公司)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

Comments Accepted by RecSys2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2507.17515 2025-07-24 cs.CV cs.CL 57%

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models

Songshuo Lu, Hua Wang, Zhi Chen, Yaohua Tang

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05018 2025-07-24 cs.LG cs.MA 57%

Joint Pedestrian and Vehicle Traffic Optimization in Urban Environments using Reinforcement Learning

Bibek Poudel, Xuan Wang, Weizi Li, Lei Zhu, Kevin Heaslip

机构 * Min H. Kao Department of Electrical Engineering and Computer Science at University of Tennessee, Knoxville, TN, USA(田纳西大学电气工程与计算机科学系Min H. Kao部门) Department of Electrical and Computer Engineering at George Mason University(乔治·马歇尔大学电气与计算机工程系) Department of Industrial and Systems Engineering at University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校工业与系统工程系) Department of Civil and Environmental Engineering at University of Tennessee, Knoxville, TN, USA(田纳西大学土木与环境工程系)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 幻觉与事实性 4 篇

2507.17477 2025-07-24 cs.AI 85%

An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models

Haoran Sun, Zekun Zhang, Shaoning Zeng

专题命中 幻觉与事实性 :alignment(title,abstract);safety(abstract);harmlessness(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17193 2025-07-24 physics.app-ph cs.LG 70%

Spintronic Bayesian Hardware Driven by Stochastic Magnetic Domain Wall Dynamics

Tianyi Wang, Bingqian Dai, Kin Wong, Yaochen Li, Yang Cheng, Qingyuan Shu, Haoran He, Puyang Huang, Hanshen Huang, Kang L. Wang

专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17722 2025-07-24 cs.CV 50%

BetterCheck: Towards Safeguarding VLMs for Automotive Perception Systems

Malsha Ashani Mahawatta Dona, Beatriz Cabrero-Daniel, Yinan Yu, Christian Berger

机构 * University of Gothenburg(哥德堡大学) Chalmers University of Technology(查尔姆斯理工大学)

专题命中 幻觉与事实性 :safety(abstract)

Comments Accepted in The IEEE International Conference on Intelligent Transportation Systems (ITSC)2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09184 2025-07-24 cs.CV 50%

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models

Qiyan Zhao, Xiaofeng Zhang, Yiheng Li, Yun Xing, Xiaosong Yuan, Feilong Tang, Sinan Fan, Xuhang Chen, Xuyao Zhang, Dahan Wang

机构 * FKLPRIU, Xiamen University of Technology, China(福克斯理工国际大学,厦门理工学院,中国) Shanghai Jiao Tong University, China(上海交通大学,中国) Nanyang Technological University, Singapore(南洋理工大学,新加坡) Jilin University, China(吉林大学,中国) Monash University, Australia(墨尔本大学,澳大利亚) Zhejiang University, China(浙江大学,中国) Huizhou University, China(惠州大学,中国) Chinese Academy of Sciences, China(中国科学院,中国)

专题命中 幻觉与事实性 :alignment(abstract)

Comments Accepted in ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 安全评测 8 篇

2507.17616 2025-07-24 cs.CV cs.AI cs.LG 81%

Vision Transformer attention alignment with human visual perception in aesthetic object evaluation

Miguel Carrasco, César González-Martín, José Aranda, Luis Oliveros

机构 * Escuela de Informática y Telecomunicaciones, Universidad Diego Portáles(戴维·波特莱斯大学信息与电信学院) Department of Specific Didactics, University of Cordoba(科尔多瓦大学特定教学系) Facultad de Ingeniería y Ciencias, Universidad Adolfo Ibáñez(阿道弗·伊巴涅斯大学工程与科学学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 25 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12404 2025-07-24 cs.LG cs.AI 81%

EXGnet: a single-lead explainable-AI guided multiresolution network with train-only quantitative features for trustworthy ECG arrhythmia classification

Tushar Talukder Showrav, Soyabul Islam Lincoln, Md. Kamrul Hasan

机构 * Dept. of Electrical and Electronic Engineering(电子与电气工程系) Bangladesh University of Engineering and Technology(孟加拉工程与技术大学) Dept. of Electronics & Communication Engineering(电子与通信工程系) Khulna University of Engineering and Technology(库尔纳工程与技术大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14900 2025-07-24 cs.CL 79%

From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment

Chongxuan Huang, Yongshi Ye, Biao Fu, Qifeng Su, Xiaodong Shi

机构 * School of Informatics, Xiamen University(厦门大学信息学院) Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能学院) Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室(厦门大学),文化和旅游部)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments ACL main 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17010 2025-07-24 cs.CR cs.AI cs.LG 76%

Towards Trustworthy AI: Secure Deepfake Detection using CNNs and Zero-Knowledge Proofs

H M Mohaimanul Islam, Huynh Q. N. Vo, Aditya Rane

机构 * School of Industrial Engineering and Management(工业工程与管理学院)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments Submitted for peer-review in TrustXR - 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17216 2025-07-24 cs.CL cs.AI 62%

The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models

Giuseppe Russo, Debora Nozza, Paul Röttger, Dirk Hovy

机构 * EPFL(苏黎世联邦理工学院) Bocconi University(博科尼大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17147 2025-07-24 cs.CL 57%

CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards

Cheng Liu, Yifei Lu, Fanghua Ye, Jian Li, Xingyu Chen, Feiliang Ren, Zhaopeng Tu, Xiaolong Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18489 2025-07-24 cs.AI cs.ET cs.RO cs.SE 57%

LLM as a code generator in Agile Model Driven Development

Ahmed R. Sadik, Sebastian Brulin, Markus Olhofer

机构 * Honda Research Institute Europe(本田欧洲研究机构)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17290 2025-07-24 cs.IR 50%

Exploring the Potential of LLMs for Serendipity Evaluation in Recommender Systems

Li Kang, Yuhan Zhao, Li Chen

专题命中 安全评测 :alignment(abstract)

Comments RecSys2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. AI治理与伦理 4 篇

2504.15181 2025-07-24 cs.CY cs.AI 81%

Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures

Lily Stelling, Mick Yang, Rokas Gipiškis, Leon Staufer, Ze Shen Chin, Siméon Campos, Ariel Gil, Michael Chen

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

Comments 166 pages, the Oxford Martin AI Governance Initiative

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17253 2025-07-24 cs.RO 71%

Optimizing Delivery Logistics: Enhancing Speed and Safety with Drone Technology

Maharshi Shastri, Ujjval Shrivastav

机构 * Department of Computer Engineering(计算机工程系) L.R. Tiwari College of Engineering(L.R.蒂瓦里工程学院)

专题命中 AI治理与伦理 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17368 2025-07-24 cs.LG 57%

ViRN: Variational Inference and Distribution Trilateration for Long-Tailed Continual Representation Learning

Hao Dai, Chong Tang, Jagmohan Chauhan

机构 * Department of Computer Science, UCL Centre for Artificial Intelligence, University College London, London, UK(计算机科学系,UCL人工智能中心,伦敦大学学院,伦敦,英国) University of Southampton, Southampton, UK(南安普顿大学,南安普顿,英国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13294 2025-07-24 cs.CY cs.HC 57%

The "Who", "What", and "How" of Responsible AI Governance: A Systematic Review and Meta-Analysis of (Actor, Stage)-Specific Tools

Blaine Kuehnert, Rachel M. Kim, Jodi Forlizzi, Hoda Heidari

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments Accepted to ACM Conference on Fairness, Accountability, and Transparency 2025. 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 其他安全 10 篇

2507.17118 2025-07-24 cs.AI 83%

HySafe-AI: Hybrid Safety Architectural Analysis Framework for AI Systems: A Case Study

Mandar Pitale, Jelena Frtunikj, Abhinaw Priyadershi, Vasu Singh, Maria Spence

机构 * Nvidia Corporation(英伟达公司) Nvidia GmbH(英伟达德国公司)

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16450 2025-07-24 cs.LG eess.SP 74%

RIS-aided Latent Space Alignment for Semantic Channel Equalization

Tomás Hüttebräucker, Mario Edoardo Pandolfo, Simone Fiorellino, Emilio Calvanese Strinati, Paolo Di Lorenzo

机构 * CEA Leti, University Grenoble Alpes(CEA Leti,格勒诺布尔大学) DIAG Department, Sapienza University of Rome(罗马萨皮恩扎大学DIAG系) Consorzio Nazionale Interuniversitario per le Telecomunicazioni (CNIT)(全国大学电信联合体) DIET Department, Sapienza University of Rome(罗马萨皮恩扎大学DIET系)

专题命中 其他安全 :alignment(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17080 2025-07-24 cs.IR cs.AI cs.CV 57%

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings

Ramin Giahi, Kehui Yao, Sriram Kollipara, Kai Zhao, Vahid Mirjalili, Jianpeng Xu, Topojoy Biswas, Evren Korpeoglu, Kannan Achan

机构 * Walmart Global Tech(沃尔玛全球技术)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted at RecSys 2025; DOI:https://doi.org/10.1145/3705328.3748064

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17055 2025-07-24 cs.RO cs.LG 57%

Shared Control of Holonomic Wheelchairs through Reinforcement Learning

Jannis Bähler, Diego Paez-Granados, Jorge Peña-Queralta

机构 * Swiss Paraplegic Research, SPF(瑞士瘫痪研究机构) SCAI Lab, D-HEST, Swiss Federal School of Technology in Zurich - ETH Zurich(SCAI实验室,瑞士联邦理工学院-苏黎世-ETH Zurich) Centre for Artificial Ingelligece, Zurich University of Applied Sciences - ZHAW. Switzerland(人工智能中心,瑞士应用科学大学-ZHAW)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16947 2025-07-24 cs.CL 57%

AI-based Clinical Decision Support for Primary Care: A Real-World Study

Robert Korom, Sarah Kiptinness, Najib Adan, Kassim Said, Catherine Ithuli, Oliver Rotich, Boniface Kimani, Irene King'ori, Stellah Kamau, Elizabeth Atemba, Muna Aden, Preston Bowman, Michael Sharman, Rebecca Soskin Hicks, Rebecca Distler, Johannes Heidecke, Rahul K. Arora, Karan Singhal

机构 * Penda Health(Penda健康) Nairobi County(内罗毕县) OpenAI

专题命中 其他安全 :safety(abstract);分类 cs.CL

Comments Blog: https://openai.com/index/ai-clinical-copilot-penda-health/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12122 2025-07-24 cs.RO cs.AI 57%

ICCO: Learning an Instruction-conditioned Coordinator for Language-guided Task-aligned Multi-robot Control

Yoshiki Yano, Kazuki Shibata, Maarten Kokshoorn, Takamitsu Matsubara

机构 * Division of Information Science, Graduate School of Science and Technology, Nara Institute of Science and Technology (NAIST)(信息科学系,科学技术研究生学校,科学与技术国立研究所) Department of Cognitive Robotics, Faculty of Mechanical Engineering, Delft University of Technology(认知机器人系,机械工程学院,代尔夫特理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 8 pages, 9 figures, to be published in the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17456 2025-07-24 cs.CV 50%

Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection

Francesco Tonini, Lorenzo Vaquero, Alessandro Conti, Cigdem Beyan, Elisa Ricci

机构 * University of Trento(特伦托大学) University of Verona(威尼斯大学) Department of Computer Science(计算机科学系)

专题命中 其他安全 :alignment(abstract)

Comments Accepted to ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10054 2025-07-24 cs.SE 50%

Explicit Vulnerability Generation with LLMs: An Investigation Beyond Adversarial Attacks

Emir Bosnak, Sahand Moslemi, Mayasah Lami, Anil Koyuncu

专题命中 其他安全 :safety(abstract)

Comments Accepted to ICSME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏