arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-18 至 2025-08-18 共收录 25 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 1 篇

2503.22402 2025-08-18 cs.DB cs.AI cs.CL 62%

EllieSQL: Cost-Efficient Text-to-SQL with Complexity-Aware Routing

Yizhang Zhu, Runzhi Jiang, Boyan Li, Nan Tang, Yuyu Luo

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 3 篇

2411.02957 2025-08-18 cs.LG cs.SY eess.SY 79%

Embedding Safety into RL: A New Take on Trust Region Methods

Nikola Milosevic, Johannes Müller, Nico Scherf

机构 * Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所) Center for Scalable Data Analytics and Artificial Intelligence(可扩展数据分析与人工智能中心)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

Comments Accepted at ICML 2025

Journal ref Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09666 2025-08-18 cs.CL 70%

Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillation

Ziyang Ma, Qingyue Yuan, Linhai Zhang, Deyu Zhou

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11504 2025-08-18 cs.LG cs.CY 62%

Predicting and Explaining Traffic Crash Severity Through Crash Feature Selection

Andrea Castellani, Zacharias Papadovasilakis, Giorgos Papoutsoglou, Mary Cole, Brian Bautsch, Tobias Rodemann, Ioannis Tsamardinos, Angela Harden

机构 * Honda Research Institute Europe(霍恩达欧洲研究机构) The Ohio State University(俄亥俄州立大学) Department of Computer Science, University of Crete(克里特大学计算机科学系) American Honda Motor Co., Inc.(美国本田摩托公司)

专题命中 安全训练 :safety(abstract);分类 cs.CY、cs.LG

Comments Preprint. Manuscript under review at "Accident Analysis & Prevention" journal

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 幻觉与事实性 2 篇

2508.11257 2025-08-18 cs.SE cs.AI 57%

Hallucination in LLM-Based Code Generation: An Automotive Case Study

Marc Pavel, Nenad Petrovic, Lukasz Mazur, Vahid Zolfaghari, Fengjunjie Pan, Alois Knoll

机构 * Real-Time Systems Technical University of Munich(实时系统技术大学慕尼黑)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11080 2025-08-18 eess.SY cs.SY 50%

Managing Risks from Large Digital Loads Using Coordinated Grid-Forming Storage Network

Soumya Kundu, Kaustav Chatterjee, Ramij R. Hossain, Sai Pushpak Nandanoori, Veronica Adetola

专题命中 幻觉与事实性 :safety(abstract)

Comments Submitted to IEEE PES T&D Conference and Expo 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 隐私与版权 3 篇

2508.11579 2025-08-18 cs.CY 57%

Intergenerational Support for Deepfake Scams Targeting Older Adults

Karina LaRubbio, Alyssa Lanter, Seihyun Lee, Mahima Ramesh, Diana Freed

专题命中 隐私与版权 :safety(abstract);分类 cs.CY

Comments 3 pages, poster at the Twenty-First Symposium on Usable Privacy and Security (SOUPS) at https://www.usenix.org/conference/soups2025/presentation/larubbio-poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10482 2025-08-18 cs.CL 57%

When Explainability Meets Privacy: An Investigation at the Intersection of Post-hoc Explainability and Differential Privacy in the Context of Natural Language Processing

Mahdi Dhaini, Stephen Meisenbacher, Ege Erdogan, Florian Matthes, Gjergji Kasneci

专题命中 隐私与版权 :trustworthy(abstract);分类 cs.CL

Comments Accepted to AAAI/ACM Conference on AI, Ethics, and Society (AIES 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09186 2025-08-18 cs.CV cs.AI 57%

RL-MoE: An Image-Based Privacy Preserving Approach In Intelligent Transportation System

Abdolazim Rezaei, Mehdi Sookhak, Mahboobeh Haghparast

机构 * Department of Computer Science Texas A\&M University Corpus Christi, USA(计算机科学系德克萨斯A&M大学科罗拉多州科罗拉多市)

专题命中 隐私与版权 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 5 篇

2508.11406 2025-08-18 cs.RO cs.AI 79%

Open, Reproducible and Trustworthy Robot-Based Experiments with Virtual Labs and Digital-Twin-Based Execution Tracing

Benjamin Alt, Mareike Picklum, Sorin Arion, Franklin Kenghagho Kenfack, Michael Beetz

机构 * AICOR Institute for Artificial Intelligence, University of Bremen(人工智能研究所,不莱梅大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 8 pages, 6 figures, submitted to the 1st IROS Workshop on Embodied AI and Robotics for Future Scientific Discovery

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11514 2025-08-18 cs.LG 57%

DiCriTest: Testing Scenario Generation for Decision-Making Agents Considering Diversity and Criticality

Qitong Chu, Yufeng Yue, Danya Yao, Huaxin Pei

机构 * School of Automation, Beijing Institute of Technology(北京理工大学自动化学院) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11317 2025-08-18 cs.CV cs.MM 50%

Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models

Yuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao, Yuqin Dai, Wenhao Yang, Chao Gou, Xiaobo Xia, Tat-Seng Chua

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11196 2025-08-18 cs.CV 50%

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning

Jiajin Guan, Haibo Mei, Bonan Zhang, Dan Liu, Yuanshuang Fu, Yue Zhang

机构 * Research Institute of Electronic Science and Technology, University of Electronic Science and Technology of China(电子科学与技术研究院,电子科技大学) School of Aeronautics and Astronautics, University of Electronic Science and Technology of China(航空航天学院,电子科技大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10947 2025-08-18 cs.CV 50%

MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text

Ronghao Xu, Zhen Huang, Yangbo Wei, Xiaoqian Zhou, Zikang Xu, Ting Liu, Zihang Jiang, S. Kevin Zhou

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 1 篇

2508.11262 2025-08-18 cs.CV cs.AI 57%

Vision-Language Models display a strong gender bias

Aiswarya Konavoor, Raj Abhijit Dandekar, Rajat Dandekar, Sreedath Panat

机构 * Togo AI Labs(Togo人工智能实验室) Vizuara AI Labs(Vizuara人工智能实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 10 篇

2508.11414 2025-08-18 cs.CL 79%

Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions

Shangrui Nie, Florian Mai, David Kaczér, Charles Welch, Zhixue Zhao, Lucie Flek

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 7 pages 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11364 2025-08-18 cs.CL 79%

Feedback Indicators: The Alignment between Llama and a Teacher in Language Learning

Sylvio Rüdian, Yassin Elsir, Marvin Kretschmer, Sabine Cayrou, Niels Pinkwart

机构 * Humboldt-Universität zu Berlin Department of Computer Science(柏林洪堡大学计算机科学系) Humboldt-Universität zu Berlin Language Centre(柏林洪堡大学语言中心) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 11 pages, one table

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11582 2025-08-18 cs.CL cs.AI 62%

Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models

Qiguang Chen, Dengyun Peng, Jinhao Liu, HuiKang Su, Jiannan Guan, Libo Qin, Wanxiang Che

机构 * LARG, Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology(大型语言模型研究组,社会计算与交互机器人研究中心,哈尔滨工业大学) School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10925 2025-08-18 cs.CL cs.AI 62%

gpt-oss-120b & gpt-oss-20b Model Card

OpenAI, :, Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K. Arora, Yu Bai, Bowen Baker, Haiming Bao, Boaz Barak, Ally Bennett, Tyler Bertao, Nivedita Brett, Eugene Brevdo, Greg Brockman, Sebastien Bubeck, Che Chang, Kai Chen, Mark Chen, Enoch Cheung, Aidan Clark, Dan Cook, Marat Dukhan, Casey Dvorak, Kevin Fives, Vlad Fomenko, Timur Garipov, Kristian Georgiev, Mia Glaese, Tarun Gogineni, Adam Goucher, Lukas Gross, Katia Gil Guzman, John Hallman, Jackie Hehir, Johannes Heidecke, Alec Helyar, Haitang Hu, Romain Huet, Jacob Huh, Saachi Jain, Zach Johnson, Chris Koch, Irina Kofman, Dominik Kundel, Jason Kwon, Volodymyr Kyrylov, Elaine Ya Le, Guillaume Leclerc, James Park Lennon, Scott Lessans, Mario Lezcano-Casado, Yuanzhi Li, Zhuohan Li, Ji Lin, Jordan Liss, Lily, Liu, Jiancheng Liu, Kevin Lu, Chris Lu, Zoran Martinovic, Lindsay McCallum, Josh McGrath, Scott McKinney, Aidan McLaughlin, Song Mei, Steve Mostovoy, Tong Mu, Gideon Myles, Alexander Neitz, Alex Nichol, Jakub Pachocki, Alex Paino, Dana Palmie, Ashley Pantuliano, Giambattista Parascandolo, Jongsoo Park, Leher Pathak, Carolina Paz, Ludovic Peran, Dmitry Pimenov, Michelle Pokrass, Elizabeth Proehl, Huida Qiu, Gaby Raila, Filippo Raso, Hongyu Ren, Kimmy Richardson, David Robinson, Bob Rotsted, Hadi Salman, Suvansh Sanjeev, Max Schwarzer, D. Sculley, Harshit Sikchi, Kendal Simon, Karan Singhal, Yang Song, Dane Stuckey, Zhiqing Sun, Philippe Tillet, Sam Toizer, Foivos Tsimpourlas, Nikhil Vyas, Eric Wallace, Xin Wang, Miles Wang, Olivia Watkins, Kevin Weil, Amy Wendling, Kevin Whinnery, Cedric Whitney, Hannah Wong, Lin Yang, Yu Yang, Michihiro Yasunaga, Kristen Ying, Wojciech Zaremba, Wenting Zhan, Cyril Zhang, Brian Zhang, Eddie Zhang, Shengjia Zhao

机构 * OpenAI

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10450 2025-08-18 cs.IR cs.AI cs.CL 62%

TokenRec: Learning to Tokenize ID for LLM-based Generative Recommendation

Haohao Qu, Wenqi Fan, Zihuai Zhao, Qing Li

机构 * Department of Computing, The Hong Kong Polytechnic University(计算机系,香港理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by IEEE TKDE. Codes and data are available at https://github.com/Quhaoh233/TokenRec

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11536 2025-08-18 cs.CL 57%

Language models align with brain regions that represent concepts across modalities

Maria Ryskina, Greta Tuckute, Alexander Fung, Ashley Malkin, Evelina Fedorenko

机构 * Vector Institute for AI(向量人工智能研究所) MIT(麻省理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted to COLM 2025. Code and data can be found at https://github.com/ryskina/concepts-brain-llms

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11404 2025-08-18 cs.RO cs.AI cs.HC 57%

An Exploratory Study on Crack Detection in Concrete through Human-Robot Collaboration

Junyeon Kim, Tianshu Ruan, Cesar Alan Contreras, Manolis Chiou

机构 * Extreme Robotics Lab (ERL) and National Center for Nuclear Robotics (NCNR)(极端机器人实验室(ERL)和核机器人国家中心(NCNR)) University of Birmingham(伯明翰大学) National Center for Nuclear Robotics (NCNR)(核机器人国家中心) Queen Mary University of London(伦敦女王大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10919 2025-08-18 cs.HC cs.AI 57%

Human-AI collaboration or obedient and often clueless AI in instruct, serve, repeat dynamics?

Mohammed Saqr, Kamila Misiejuk, Sonsoles López-Pernas

机构 * University of Eastern Finland(东方芬兰大学) FernUniversität in Hagen(哈根弗伦大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02269 2025-08-18 cs.AI 57%

AirTrafficGen: Configurable Air Traffic Scenario Generation with Large Language Models

Dewi Sid William Gould, George De Ath, Ben Carvell, Nick Pepper

机构 * The Alan Turing Institute(阿尔ัน图灵研究院) NATS University of Exeter(埃克塞特大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 9 pages and appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11153 2025-08-18 cs.CV 50%

LEARN: A Story-Driven Layout-to-Image Generation Framework for STEM Instruction

Maoquan Zhang, Bisser Raytchev, Xiujuan Sun

机构 * Graduate School of Advanced Science and Engineering, Hiroshima University(Hiroshima大学研究生院) Department of Computer Science, Weifang University of Science and Technology(潍坊科技大学计算机科学系)

专题命中 其他安全 :alignment(abstract)

Comments The International Conference on Neural Information Processing (ICONIP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏