arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-15 至 2025-09-15 共收录 24 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 1 篇

2509.10147 2025-09-15 cs.AI 57%

Virtual Agent Economies

Nenad Tomasev, Matija Franklin, Joel Z. Leibo, Julian Jacobs, William A. Cunningham, Iason Gabriel, Simon Osindero

机构 * Google DeepMind(谷歌DeepMind) University of Toronto(多伦多大学)

专题命中 偏好对齐 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2509.09852 2025-09-15 cs.CL 57%

Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization

Chuyuan Li, Austin Xu, Shafiq Joty, Giuseppe Carenini

机构 * Department of Computer Science, University of British Columbia(英属哥伦比亚大学计算机科学系) Salesforce AI Research(Salesforce人工智能研究)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09953 2025-09-15 cs.RO 50%

Detection of Anomalous Behavior in Robot Systems Based on Machine Learning

Mahfuzul I. Nissan, Sharmin Aktar

机构 * Department of Computer Science University of New Orleans(计算机科学系 新奥尔良大学) Department of Computer Science St Mary's University(计算机科学系 斯坦玛丽大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 提示注入 1 篇

2509.09912 2025-09-15 cs.CY cs.CR 79%

When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review

Changjia Zhu, Junjie Xiong, Renkai Ma, Zhicong Lu, Yao Liu, Lingyao Li

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 4 篇

2502.07974 2025-09-15 cs.SE cs.AI cs.LG 81%

From Hazard Identification to Controller Design: Proactive and LLM-Supported Safety Engineering for ML-Powered Systems

Yining Hong, Christopher S. Timperley, Christian Kästner

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.AI、cs.LG

Comments Accepted for publication at the International Conference on AI Engineering (CAIN) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10048 2025-09-15 cs.LG 79%

Uncertainty-Aware Tabular Prediction: Evaluating VBLL-Enhanced TabPFN in Safety-Critical Medical Data

Madhushan Ramalingam

机构 * Engineering University of Moratuwa Sri Lanka(穆塔瓦大学工程学院)

专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10208 2025-09-15 cs.CL cs.AI 62%

SI-FACT: Mitigating Knowledge Conflict via Self-Improving Faithfulness-Aware Contrastive Tuning

Shengqiang Fu

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09989 2025-09-15 cs.CR 50%

rCamInspector: Building Reliability and Trust on IoT (Spy) Camera Detection using XAI

Priyanka Rushikesh Chaudhary, Manan Gupta, Jabez Christopher, Putrevu Venkata Sai Charan, Rajib Ranjan Maiti

专题命中 幻觉与事实性 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 6 篇

2509.10059 2025-09-15 cs.CV cs.AI 57%

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration

Yue Zhou, Litong Feng, Mengcheng Lan, Xue Yang, Qingyun Li, Yiping Ke, Xue Jiang, Wayne Zhang

机构 * Department of Physics, J.K. Institute of Science(J.K.科学研究院物理系) World Scientific University(世界科学大学) University of Intelligent Studies(智能研究大学) East China Normal University(华东师范大学) Nanyang Technological University(南洋理工大学) SenseTime Research(商汤科技研究院) Shanghai Jiao Tong University(上海交通大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 17 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09799 2025-09-15 cs.LG cs.HC 57%

Distinguishing Startle from Surprise Events Based on Physiological Signals

Mansi Sharma, Alexandre Duchevet, Florian Daiber, Jean-Paul Imbert, Maurice Rekrut

机构 * German Research Center for Artificial Intelligence (DFKI), Cognitive Assistants Lab, Saarland Informatics Campus(德国人工智能研究中心(DFKI)、认知助理实验室、萨尔兰州信息技术校区) Fédération ENAC ISAE-SUPAERO ONERA, Université de Toulouse(ENAC ISAE-SUPAERO ONERA联合会、图卢兹大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16708 2025-09-15 cs.CL 57%

Atomic Fact Decomposition Helps Attributed Question Answering

Zhichao Yan, Jiapu Wang, Jiaoyan Chen, Xiaoli Li, Ru Li, Jeff Z. Pan

机构 * School of Computer and Information Technology, Shanxi University(山西大学计算机与信息学院) Beijing Key Laboratory of Multimedia and Intelligent Software Technology, Beijing Institute of Artificial Intelligence, School of Information Science and Technology, Beijing University of Technology(北京人工智能研究院多媒体与智能软件技术重点实验室,北京理工大学信息科学技术学院) University of Manchester(曼彻斯特大学) Singapore University of Technology and Design(新加坡科技设计大学) ILCC, School of Informatics, University of Edinburgh(爱丁堡大学信息学院ILCC)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10284 2025-09-15 cs.MA cs.RO 50%

A Holistic Architecture for Monitoring and Optimization of Robust Multi-Agent Path Finding Plan Execution

David Zahrádka, Denisa Mužíková, David Woller, Miroslav Kulich, Jiří Švancara, Roman Barták

机构 * Faculty of Electrical Engineering, Czech Technical University in Prague, Prague, Czech Republic(布拉格捷克技术大学电气工程学院) Czech Institute of Informatics, Robotics and Cybernetics, Czech Technical University in Prague, Prague, Czech Republic(捷克信息学、机器人学与自动化研究所) Faculty of Mathematics and Physics, Charles University, Prague, Czech Republic(布拉格查理大学数学与物理学院)

专题命中 安全评测 :safety(abstract)

Comments 23 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09922 2025-09-15 econ.GN q-fin.EC 50%

Robo-Advisors Beyond Automation: Principles and Roadmap for AI-Driven Financial Planning

Runhuan Feng, Hong Li, Ming Liu

专题命中 安全评测 :trustworthy(abstract)

Comments 37 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02992 2025-09-15 cs.IR cs.CV 50%

Efficient and Effective Adaptation of Multimodal Foundation Models in Sequential Recommendation

Junchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Kaiwen Zheng, Yongxin Ni, Joemon M. Jose

机构 * School of Computing Science, University of Glasgow(格拉斯哥大学计算机科学学院) Amazon(亚马逊) Telefonica Scientific Research(Telefonica科学研究院) University of Science and Technology of China(中国科学技术大学)

专题命中 安全评测 :alignment(abstract)

Comments Accepted by IEEE Transactions on Knowledge and Data Engineering (TKDE)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 3 篇

2509.08858 2025-09-15 cs.CY cs.LG 81%

Decentralising LLM Alignment: A Case for Context, Pluralism, and Participation

Oriane Peter, Kate Devlin

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CY、cs.LG

Comments Accepted at AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10289 2025-09-15 cs.CY cs.AI 62%

We Need a New Ethics for a World of AI Agents

Iason Gabriel, Geoff Keeling, Arianna Manzini, James Evans

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 6 pages, no figures

Journal ref Nature, 644 (8075), 2025, 38-40

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02773 2025-09-15 cs.CY cs.AI econ.GN q-fin.EC 62%

Web3 x AI Agents: Landscape, Integrations, and Foundational Challenges

Yiming Shen, Jiashuo Zhang, Zhenzhe Shao, Wenxuan Luo, Yanlin Wang, Ting Chen, Zibin Zheng, Jiachi Chen

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 7 篇

2509.10344 2025-09-15 cs.CV cs.AI cs.LG 81%

GLAM: Geometry-Guided Local Alignment for Multi-View VLP in Mammography

Yuexi Du, Lihui Chen, Nicha C. Dvornek

机构 * Department of Biomedical Engineering(生物医学工程系) Department of Radiology & Biomedical Imaging(放射科与生物医学成像系) Yale University(耶鲁大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13204 2025-09-15 cs.CL 79%

Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verification

Jikai Wang, Zhenxu Tian, Juntao Li, Qingrong Xia, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Min Zhang

机构 * Soochow University(苏州大学) Key Laboratory of Data Intelligence and Advanced Computing, Soochow University(数据智能与先进计算关键实验室) Huawei Cloud(华为云)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09740 2025-09-15 q-bio.QM cs.AI cs.CL cs.LG 67%

HypoGeneAgent: A Hypothesis Language Agent for Gene-Set Cluster Resolution Selection Using Perturb-seq Datasets

Ying Yuan, Xing-Yue Monica Ge, Aaron Archer Waterman, Tommaso Biancalani, David Richmond, Yogesh Pandit, Avtar Singh, Russell Littman, Jin Liu, Jan-Christian Huetter, Vladimir Ermakov

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07373 2025-09-15 cs.LG cs.AI 62%

Atherosclerosis through Hierarchical Explainable Neural Network Analysis

Irsyad Adam, Steven Swee, Erika Yilin, Ethan Ji, William Speier, Dean Wang, Alex Bui, Wei Wang, Karol Watson, Peipei Ping

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12039 2025-09-15 cs.CR cs.AI cs.CL cs.SE 62%

Can LLM Prompting Serve as a Proxy for Static Analysis in Vulnerability Detection

Ira Ceka, Feitong Qiao, Anik Dey, Aastha Valecha, Gail Kaiser, Baishakhi Ray

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10380 2025-09-15 eess.SY cs.SY 50%

Merging Physics-Based Synthetic Data and Machine Learning for Thermal Monitoring of Lithium-ion Batteries: The Role of Data Fidelity

Yusheng Zheng, Wenxue Liu, Yunhong Che, Ferdinand Grimm, Jingyuan Zhao, Xiaosong Hu, Simona Onori, Remus Teodorescu, Gregory J. Offer

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09732 2025-09-15 cs.CV 50%

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs

Sary Elmansoury, Islam Mesabah, Gerrit Großmann, Peter Neigel, Raj Bhalwankar, Daniel Kondermann, Sebastian J. Vollmer

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏