arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-01 至 2025-08-01 共收录 29 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2507.22789 2025-08-01 cs.LG cs.AI 81%

G-Core: A Simple, Scalable and Balanced RLHF Trainer

Junyu Wu, Weiming Chang, Xiaotao Liu, Guanyou He, Haoqiang Hong, Boqi Liu, Hongtao Tian, Tao Yang, Yunsheng Shi, Feng Lin, Ting Yao

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

Comments I haven't received company approval yet, and I uploaded it by mistake

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23391 2025-08-01 cs.LG cs.RO 57%

Policy Learning from Large Vision-Language Model Feedback without Reward Modeling

Tung M. Luu, Donghoon Lee, Younghwan Lee, Chang D. Yoo

机构 * School of Electrical Engineering, KAIST(韩国科学技术院电子工程学院)

专题命中 偏好对齐 :safety(abstract);分类 cs.LG

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23478 2025-08-01 cs.CV 50%

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding

Ting Huang, Zeyu Zhang, Hao Tang

专题命中 偏好对齐 :RLHF(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 3 篇

2410.09486 2025-08-01 cs.LG cs.RO 79%

ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning

Yarden As, Bhavya Sukhija, Lenart Treven, Carmelo Sferrazza, Stelian Coros, Andreas Krause

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23226 2025-08-01 cs.CV 78%

Toward Safe, Trustworthy and Realistic Augmented Reality User Experience

Yanming Xiu

机构 * Department of Electrical and Computer Engineering, Duke University(电子工程系,杜克大学)

专题命中 安全训练 :trustworthy(title);safety(abstract)

Comments 2 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18666 2025-08-01 cs.AI cs.CL 62%

AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents

Haoyu Wang, Christopher M. Poskitt, Jun Sun

机构 * Singapore Management University(新加坡管理大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted by the 48th IEEE/ACM International Conference on Software Engineering (ICSE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 幻觉与事实性 3 篇

2507.22915 2025-08-01 cs.CL cs.AI 62%

Theoretical Foundations and Mitigation of Hallucination in Large Language Models

Esmail Gumaan

机构 * Department of Computer Science, University of Sana'a(桑加大学计算机科学系)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23736 2025-08-01 stat.ML cs.LG 57%

DICOM De-Identification via Hybrid AI and Rule-Based Framework for Scalable, Uncertainty-Aware Redaction

Kyle Naddeo, Nikolas Koutsoubis, Rahul Krish, Ghulam Rasool, Nidhal Bouaynaya, Tony OSullivan, Raj Krish

机构 * Rowan University(罗文大学) Moffitt Cancer Center(莫菲特癌症中心) University of South Florida(佛罗里达州立大学) Impact Business Information Solutions, Inc(Impact商务信息解决方案公司)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments 15 pages, 6 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20293 2025-08-01 cs.RO 50%

Decentralized Uncertainty-Aware Multi-Agent Collision Avoidance with Model Predictive Path Integral

Stepan Dergachev, Konstantin Yakovlev

机构 * FRC CSC RAS(俄罗斯科学院应用数学与控制论研究所) HSE University(俄罗斯高等经济大学) AIRI(人工智能研究所)

专题命中 幻觉与事实性 :safety(abstract)

Comments This is a pre-print of the paper accepted to IROS2025. The manuscript includes 8 pages, 4 figures, and 1 table. A supplementary video is available at https://youtu.be/_D4zDYJ4KCk Updated version: added link to source code in the abstract; updated experimental results description in Section VI.A; updated author affiliation and funding information; minor typo corrections

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 隐私与版权 1 篇

2507.18518 2025-08-01 cs.IR 78%

Transform Before You Query: A Privacy-Preserving Approach for Vector Retrieval with Embedding Space Alignment

Ruiqi He, Zekun Fei, Jiaqi Li, Xinyuan Zhu, Biao Yi, Siyi Lv, Weijie Liu, Zheli Liu

专题命中 隐私与版权 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 9 篇

2507.10817 2025-08-01 stat.AP 78%

Is Your Model Risk ALARP? Evaluating Prospective Safety-Critical Applications of Complex Models

Domenic Di Francesco, Alan Forrest, Fiona McGarry, Nicholas Hall, Adam Sobey

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23535 2025-08-01 cs.LG cs.AI cs.CY 67%

Transparent AI: The Case for Interpretability and Explainability

Dhanesh Ramachandram, Himanshu Joshi, Judy Zhu, Dhari Gandhi, Lucas Hartman, Ananya Raval

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22958 2025-08-01 cs.CV cs.AI cs.LG 62%

CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam

Ruslan Khrulev

机构 * Moscow State University(莫斯科国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 15 pages, 3 figures, 10 tables. Code is available at: https://github.com/Karifannaa/Auto-check-EGE-math

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07695 2025-08-01 cs.CL cs.AI 62%

KeyKnowledgeRAG (K^2RAG): An Enhanced RAG method for improved LLM question-answering capabilities

Hruday Markondapatnaikuni, Basem Suleiman, Abdelkarim Erradi, Shijing Chen

机构 * University of Sydney(悉尼大学) University of New South Wales(新南威尔士大学) Qatar University(卡塔尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23000 2025-08-01 cs.LG cs.CV 57%

Planning for Cooler Cities: A Multimodal AI Framework for Predicting and Mitigating Urban Heat Stress through Urban Landscape Transformation

Shengao Yi, Xiaojiang Li, Wei Tu, Tianhong Zhao

机构 * Department of City and Regional Planning, University of Pennsylvania(城市与区域规划系,宾夕法尼亚大学) Guangdong Key Laboratory for Urban Informatics, Guangdong-Hong Kong-Macao Joint Laboratory for Smart Cities(广东城市信息关键实验室、粤港澳大湾区智慧城市联合实验室) Shenzhen Key Laboratory of Spatial Information Smart Sensing(深圳空间信息智能感知关键实验室) Department of Urban Informatics, School of Architecture and Urban Planning, Shenzhen University(城市信息系,深圳大学) College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17186 2025-08-01 cs.CL 57%

FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain

Lingfeng Zeng, Fangqi Lou, Zixuan Wang, Jiajie Xu, Jinyi Niu, Mengping Li, Yifan Dong, Qi Qi, Wei Zhang, Ziwei Yang, Jun Han, Ruilun Feng, Ruiqi Hu, Lejie Zhang, Zhengbo Feng, Yicheng Ren, Xin Guo, Zhaowei Liu, Dongpo Cheng, Weige Cai, Liwen Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12284 2025-08-01 cs.LG stat.ML 57%

GrokAlign: Geometric Characterisation and Acceleration of Grokking

Thomas Walker, Ahmed Imtiaz Humayun, Randall Balestriero, Richard Baraniuk

机构 * Department of Electrical and Computer Engineering, Rice University(电气与计算机工程系, Rice大学) Google Research(谷歌研究院) Department of Computer Science, Brown University(计算机科学系, Brown大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 23 pages, 11 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02123 2025-08-01 cs.CV cs.LG 57%

FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks

Mahadev Prasad Panda, Matteo Tiezzi, Martina Vilas, Gemma Roig, Bjoern M. Eskofier, Dario Zanca

机构 * Department AIBE, FAU Erlangen-Nürnberg(FAU埃朗根-纽伦堡大学AIBE部门) PAVIS, Istituto Italiano di Tecnologia (IIT), Genova, Italy(意大利技术研究院(IIT)帕维斯部门,热那亚,意大利) Ernst Strüngmann Institute for Neuroscience, Frankfurt, Germany(神经科学埃朗根-纽伦堡研究所,法兰克福,德国) Goethe-Universität Frankfurt am Main(法兰克福大学) Institute of AI for Health, Helmholtz Zentrum München, Munich, Germany(健康人工智能研究所,海德堡中心慕尼黑,德国)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Accepted in the International Journal of Computer Vision (Springer Nature)

Journal ref Int J Comput Vis (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23489 2025-08-01 eess.SY cs.SY 50%

Distributionally Robust Cascading Risk Quantification in Multi-Agent Rendezvous: Effects of Time Delay and Network Connectivity

Vivek Pandey, Nader Motee

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 2 篇

2507.23373 2025-08-01 cs.CV 78%

Multi-Prompt Progressive Alignment for Multi-Source Unsupervised Domain Adaptation

Haoran Chen, Zexiao Wang, Haidong Cao, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学)

专题命中 AI治理与伦理 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10073 2025-08-01 cs.CL cs.AI 66%

Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires

Simon Münker

机构 * Tier University(Tier大学)

专题命中 AI治理与伦理 :alignment(abstract,journal_ref);分类 cs.CL、cs.AI

Comments 15pages, 1 figure, 2 tables

Journal ref Proceedings of 0th Symposium on Moral and Legal AI Alignment of the IACAP/AISB Conference, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 8 篇

2507.22920 2025-08-01 cs.CL cs.AI 62%

Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey

Jindong Li, Yali Fu, Jiahong Liu, Linxiao Cao, Wei Ji, Menglin Yang, Irwin King, Ming-Hsuan Yang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22902 2025-08-01 cs.HC cs.AI cs.CL cs.MA 62%

Toward the Autonomous AI Doctor: Quantitative Benchmarking of an Autonomous Agentic AI Versus Board-Certified Clinicians in a Real World Setting

Hashim Hayat, Maksim Kudrautsau, Evgeniy Makarov, Vlad Melnichenko, Tim Tsykunou, Piotr Varaksin, Matt Pavelle, Adam Z. Oskowitz

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22633 2025-08-01 cs.LG cs.AI 62%

H2Tune: Federated Foundation Model Fine-Tuning with Hybrid Heterogeneity

Wei Guo, Siyuan Lu, Yiqi Tong, Zhaojun Hu, Fuzhen Zhuang, Xiao Zhang, Tao Fan, Jin Dong

机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) School of Computer Science and Technology, Heilongjiang University(黑龙江大学计算机科学与技术学院) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) School of Statistics, Renmin University of China(中国人民大学统计学院) Zhongguancun Laboratory, China(中关村实验室) School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院) WeBank Co., Ltd, Shenzhen, China(WeBank有限公司,深圳,中国) Beijing Academy of Blockchain and Edge Computing(北京区块链与边缘计算研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22879 2025-08-01 cs.IR cs.CL 57%

RecGPT Technical Report

Chao Yi, Dian Chen, Gaoyang Guo, Jiakai Tang, Jian Wu, Jing Yu, Mao Zhang, Sunhao Dai, Wen Chen, Wenjun Yang, Yuning Jiang, Zhujin Gao, Bo Zheng, Chi Li, Dimin Wang, Dixuan Wang, Fan Li, Fan Zhang, Haibin Chen, Haozhuang Liu, Jialin Zhu, Jiamang Wang, Jiawei Wu, Jin Cui, Ju Huang, Kai Zhang, Kan Liu, Lang Tian, Liang Rao, Longbin Li, Lulu Zhao, Na He, Peiyang Wang, Qiqi Huang, Tao Luo, Wenbo Su, Xiaoxiao He, Xin Tong, Xu Chen, Xunke Xi, Yang Li, Yaxuan Wu, Yeqiu Yang, Yi Hu, Yinnan Song, Yuchen Li, Yujie Luo, Yujin Yuan, Yuliang Yan, Zhengyang Wang, Zhibo Xiao, Zhixin Ma, Zile Zhou, Ziqi Zhang

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07797 2025-08-01 cs.LG 57%

A Theoretical Framework for Explaining Reinforcement Learning with Shapley Values

Daniel Beechey, Thomas M. S. Smith, Özgür Şimşek

机构 * University of Bath(巴斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23058 2025-08-01 cs.CV cs.AI 57%

Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation

Alexandru Buburuzan

机构 * Department of Computer Science(计算机科学系)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments A dissertation submitted to The University of Manchester for the degree of Bachelor of Science in Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20947 2025-08-01 cs.CV cs.MM 50%

Hierarchical Sub-action Tree for Continuous Sign Language Recognition

Dejie Yang, Zhu Xu, Xinjie Gao, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University, Beijing, China(王轩计算机技术研究所,北京大学,北京,中国) State Key Laboratory of General Artificial Intelligence, Peking Universitys, Beijing, China(通用人工智能国家重点实验室,北京大学,北京,中国)

专题命中 其他安全 :alignment(abstract)

Journal ref ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15135 2025-08-01 cs.RO 50%

Controllable Traffic Simulation through LLM-Guided Hierarchical Reasoning and Refinement

Zhiyuan Liu, Leheng Li, Yuning Wang, Haotian Lin, Hao Cheng, Zhizhe Liu, Lei He, Jianqiang Wang

机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动学院) Artificial Intelligence Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)人工智能研究部) School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院) Department of Computer Sciences, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系)

专题命中 其他安全 :safety(abstract)

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏