arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-23 至 2025-07-23 共收录 32 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2502.18699 2025-07-23 cs.CL cs.LG stat.ME 84%

MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment

Tianze Wang, Dongnan Gui, Yifan Hu, Shuhang Lin, Linjun Zhang

机构 * Department of Statistics, Rutgers University, New Brunswick, United States Department of Computer Science, Rutgers University, New Brunswick, United States College of Management of Technology, EPFL, Switzerland Department of Computer Science, ETH Zurich, Switzerland

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07939 2025-07-23 cs.CL 79%

SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment

Guoxin Zang, Xue Li, Donglin Di, Lanshun Nie, Dechen Zhan, Yang Song, Lei Fan

机构 * Harbin Institute of Technology(哈尔滨工业大学) University of New South Wales(新南威尔士大学)

专题命中 偏好对齐 :alignment(title);DPO(abstract);分类 cs.CL

Comments Accepted by ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16252 2025-07-23 cs.CL cs.AI 62%

Efficient RL for optimizing conversation level outcomes with an LLM-based tutor

Hyunji Nam, Omer Gottesman, Amy Zhang, Dean Foster, Emma Brunskill, Lyle Ungar

机构 * Stanford University(斯坦福大学) Amazon(亚马逊公司) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Pennsylvania(宾夕法尼亚大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15907 2025-07-23 cs.LG cs.AI 62%

Dual Turing Test: A Framework for Detecting and Mitigating Undetectable AI

Alberto Messina

机构 * RAI - Radiotelevisione Italiana, Centre for Research, Technological Innovation and Experimentation (CRITS)(意大利广播电视台,研究中心、技术创新与实验中心(CRITS))

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 3 篇

2507.16322 2025-07-23 cs.AI 57%

Mind the Gap: Evaluating the Representativeness of Quantitative Medical Language Reasoning LLM Benchmarks for African Disease Burdens

Fred Mutisya, Shikoh Gitau, Christine Syovata, Diana Oigara, Ibrahim Matende, Muna Aden, Munira Ali, Ryan Nyotu, Diana Marion, Job Nyangena, Nasubo Ongoma, Keith Mbae, Elizabeth Wamicha, Eric Mibuari, Jean Philbert Nsengemana, Talkmore Chidede

专题命中 安全训练 :alignment(abstract);分类 cs.AI

Comments Preprint. 26 pages, includes appendix and tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00477 2025-07-23 cs.CV 50%

Vision-based Conflict Detection within Crowds based on High-Resolution Human Pose Estimation for Smart and Safe Airport

Karan Kheta, Claire Delgove, Ruolin Liu, Adeola Aderogba, Marc-Olivier Pokam, Muhammed Mehmet Unal, Yang Xing, Weisi Guo

专题命中 安全训练 :safety(abstract)

Comments One of the authors has expressed privacy concerns and made a related request

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06574 2025-07-23 cs.RO 50%

AI Space Cortex: An Experimental System for Future Era Space Exploration

Thomas Touma, Ersin Daş, Erica Tevere, Martin Feather, Ksenia Kolcio, Maurice Prather, Alberto Candela, Ashish Goel, Erik Kramer, Hari Nayar, Lorraine Fesq, Joel W. Burdick

机构 * California Institute of Technology(加州理工学院) Jet Propulsion Laboratory(喷气推进实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2407.09164 2025-07-23 cs.CR cs.AI 79%

ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs

Yuchen Yang, Yiming Li, Hongwei Yao, Bingrun Yang, Yiling He, Tianwei Zhang, Dacheng Tao, Zhan Qin

机构 * State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, Hangzhou(杭州高新区(滨江)区块链与数据安全研究院,杭州) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 越狱攻击 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16576 2025-07-23 cs.CR 50%

From Text to Actionable Intelligence: Automating STIX Entity and Relationship Extraction

Ahmed Lekssays, Husrev Taha Sencar, Ting Yu

专题命中 越狱攻击 :alignment(abstract)

Comments This paper is accepted at RAID 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 3 篇

2507.15906 2025-07-23 cs.LG cs.AI 81%

Towards Reliable, Uncertainty-Aware Alignment

Debangshu Banerjee, Kintan Saha, Aditya Gopalan

机构 * Undergraduate Programme, Indian Institute of Science, India(印度科学研究所)

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15880 2025-07-23 cs.AI 79%

The Recursive Coherence Principle: A Formal Constraint on Scalable Intelligence, Alignment, and Reasoning Architecture

Andy E. Williams

专题命中 幻觉与事实性 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16771 2025-07-23 cs.LG stat.AP stat.ML 57%

A Partitioned Sparse Variational Gaussian Process for Fast, Distributed Spatial Modeling

Michael Grosskopf, Kellin Rumsey, Ayan Biswas, Earl Lawrence

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 2 篇

2507.16731 2025-07-23 cs.DC 67%

Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges

Senyao Li, Haozhao Wang, Wenchao Xu, Rui Zhang, Song Guo, Jingling Yuan, Xian Zhong, Tianwei Zhang, Ruixuan Li

专题命中 隐私与版权 :alignment(abstract);trustworthy(abstract)

Comments 35 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16372 2025-07-23 cs.CR cs.AI 57%

Depth Gives a False Sense of Privacy: LLM Internal States Inversion

Tian Dong, Yan Meng, Shaofeng Li, Guoxing Chen, Zhen Liu, Haojin Zhu

机构 * Shanghai Jiao Tong University(上海交通大学) Southeast University(东南大学)

专题命中 隐私与版权 :safety(abstract);分类 cs.AI

Comments Accepted by USENIX Security 2025. Please cite this paper as "Tian Dong, Yan Meng, Shaofeng Li, Guoxing Chen, Zhen Liu, Haojin Zhu. Depth Gives a False Sense of Privacy: LLM Internal States Inversion. In the 34th USENIX Security Symposium (USENIX Security '25)."

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 7 篇

2507.16033 2025-07-23 cs.HC cs.AI 79%

"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives

Ding Wang, Mark Díaz, Charvi Rastogi, Aida Davani, Vinodkumar Prabhakaran, Pushkar Mishra, Roma Patel, Alicia Parrish, Zoe Ashwood, Michela Paganini, Tian Huey Teh, Verena Rieser, Lora Aroyo

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments Accepted to AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society 2025 (AIES 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10706 2025-07-23 cs.CL cs.AI cs.CY cs.HC cs.RO 75%

SciFi-Benchmark: Leveraging Science Fiction To Improve Robot Behavior

Pierre Sermanet, Anirudha Majumdar, Vikas Sindhwani

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Minor improvements over previous version

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16206 2025-07-23 cs.LG cs.AI 73%

METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark

Xu Yang, Qi Zhang, Shuming Jiang, Yaowen Xu, Zhaofan Zou, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(电信人工智能研究院) Institute of Artificial Intelligence and Robotics(IAIR), Xi’an Jiaotong University(人工智能与机器人研究院) Advanced Technique of Artificial Intelligence(ATAI), Chongqing University of Technology(人工智能先进技术研究院)

专题命中 安全评测 :DPO(abstract);safety(abstract);分类 cs.AI、cs.LG

Comments 9 pages,3 figures ICCV format

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05119 2025-07-23 cs.LG cs.AI cs.AR cs.CV eess.IV 62%

Balancing Robustness and Efficiency in Embedded DNNs Through Activation Function Selection

Jon Gutiérrez-Zaballa, Koldo Basterretxea, Javier Echanobe

机构 * Department of Electronics Technology, University of the Basque Country (UPV/EHU)(电子技术系,巴斯克国家大学(UPV/EHU)) Department of Electricity and Electronics, University of the Basque Country (UPV/EHU)(电力与电子系,巴斯克国家大学(UPV/EHU))

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15874 2025-07-23 cs.AI cs.CL 62%

Why Braking? Scenario Extraction and Reasoning Utilizing LLM

Yin Wu, Daniel Slieter, Vivek Subramanian, Ahmed Abouelazm, Robin Bohn, J. Marius Zöllner

机构 * CARIAD SE(CARIAD公司) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) FZI Research Center for Information Technology(弗劳恩霍夫信息技术研究中心)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15868 2025-07-23 cs.CL cs.AI 62%

Small Edits, Big Consequences: Telling Good from Bad Robustness in Large Language Models

Altynbek Ismailov, Salia Asanova

机构 * Berkeley(伯克利)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16572 2025-07-23 cs.CL 57%

Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models

Mohamad Ballout, Serwan Jassim, Elia Bruni

机构 * Institute of Cognitive Science, University of Osnabrück(认知科学研究所,奥斯纳布吕克大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 4 篇

2402.04247 2025-07-23 cs.CY cs.AI cs.CL cs.LG 77%

Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy

Xiangru Tang, Qiao Jin, Kunlun Zhu, Tongxin Yuan, Yichi Zhang, Wangchunshu Zhou, Meng Qu, Yilun Zhao, Jian Tang, Zhuosheng Zhang, Arman Cohan, Zhiyong Lu, Mark Gerstein

机构 * Yale University(耶鲁大学) National Library of Medicine, National Institutes of Health(国家医学图书馆,国立卫生研究院) Mila-Quebec AI Institute(魁北克AI研究所) Shanghai Jiao Tong University(上海交通大学) OPPO Research Institute(OPPO研究院) Reichman University(里奇曼大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15901 2025-07-23 cs.AI cs.CY cs.MA 62%

Advancing Responsible Innovation in Agentic AI: A study of Ethical Frameworks for Household Automation

Joydeep Chandra, Satyam Kumar Navneet

机构 * Department of CST Tsinghua University Beijing, China(计算机科学与技术系 清华大学 北京中国) Department of CSE Chandigarh University Mohali, India(计算机科学与工程系 印度昌迪加尔大学 摩哈利)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15885 2025-07-23 cs.AI cs.HC cs.LG 62%

ADEPTS: A Capability Framework for Human-Centered Agent Design

Pierluca D'Oro, Caley Drooff, Joy Chen, Joseph Tighe

机构 * FAIR at Meta(Meta 的 FAIR)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14226 2025-07-23 cs.CY 57%

Mapping the Parasocial AI Market: User Trends, Engagement and Risks

Zilan Qian, Mari Izumikawa, Fiona Lodge, Angelo Leone

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 17 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 7 篇

2505.16104 2025-07-23 cs.CL cs.CV cs.LG 81%

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models

Yue Li, Xin Yi, Dongsheng Shi, Gerard de Melo, Xiaoling Wang, Linlin Wang

机构 * East China Normal University(华东师范大学) Hasso Plattner Institute/University of Potsdam(哈索普劳特纳研究所/波茨坦大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.LG

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07525 2025-07-23 cs.CV cs.AI cs.LG 76%

RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment

Difei Gu, Yunhe Gao, Yang Zhou, Mu Zhou, Dimitris Metaxas

机构 * Rutgers University(罗格斯大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

Comments Accepted to MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22116 2025-07-23 cs.CL cs.AI 62%

Multimodal Forecasting of Sparse Intraoperative Hypotension Events Powered by Language Model

Jintao Zhang, Zirui Liu, Mingyue Cheng, Shilong Zhang, Tingyue Pan, Yitong zhou, Qi Liu, Yanhu Xie

机构 * University of Science and Technology of China(中国科学技术大学) The First Affiliated Hospital of University of Science and Technology of China(中国科学技术大学第一附属医院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18527 2025-07-23 cs.IR cs.LG 57%

Probing Ranking LLMs: A Mechanistic Analysis for Information Retrieval

Tanya Chowdhury, Atharva Nijasure, James Allan

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02304 2025-07-23 cs.CL cs.CV 57%

Generative Sign-description Prompts with Multi-positive Contrastive Learning for Sign Language Recognition

Siyu Liang, Yunan Li, Wentian Xin, Huizhou Chen, Xujie Liu, Kang Liu, Qiguang Miao

机构 * Xidian University(西安电子科技大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏