arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-06 至 2025-08-06 共收录 36 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 2 篇

2504.13134 2025-08-06 cs.CL cs.LG stat.ML 84%

Energy-Based Reward Models for Robust Language Model Alignment

Anamika Lochab, Ruqi Zhang

机构 * Department of Computer Science(计算机科学系) Purdue University(普渡大学)

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.LG

Comments Accepted by COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05070 2025-08-06 cs.CL 57%

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation

Tianjiao Li, Mengran Yu, Chenyu Shi, Yanjun Zhao, Xiaojing Liu, Qiang Zhang, Qi Zhang, Xuanjing Huang, Jiayin Wang

机构 * Bilibili Inc.(哔哩哔哩公司) Xi’an Jiaotong University(西安交通大学) School of Computer Science, Fudan University(复旦大学计算机科学学院)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 越狱攻击 3 篇

2508.03054 2025-08-06 cs.AI 79%

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning

Rui Pu, Chaozhuo Li, Rui Ha, Litian Zhang, Lirong Qiu, Xi Zhang

机构 * Key Laboratory of Trustworthy Distributed Computing and Service (MoE)(可信分布式计算与服务重点实验室)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03125 2025-08-06 cs.CR cs.AI cs.MA 57%

Attack the Messages, Not the Agents: A Multi-round Adaptive Stealthy Tampering Framework for LLM-MAS

Bingyu Yan, Ziyi Zhou, Xiaoming Zhang, Chaozhuo Li, Ruilin Zeng, Yirui Qi, Tianbo Wang, Litian Zhang

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03110 2025-08-06 cs.CL 57%

Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation

Zizhong Li, Haopeng Zhang, Jiawei Zhang

机构 * University of California, Davis(加州大学戴维斯分校) University of Hawaii at Mānoa(夏威夷大学马诺阿分校)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 红队测试 1 篇

2503.04856 2025-08-06 cs.CL cs.AI 88%

M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs

Junwoo Ha, Hyunjun Kim, Sangyoon Yu, Haon Park, Ashkan Yousefpour, Yuna Park, Suhyun Kim

机构 * AIM Intelligence(AIM智能研究所) University of Seoul(首尔大学) Korea Advanced Institute of Science and Technology(韩国科学技术院) Seoul National University(首尔国立大学) Yonsei University(延世大学) Korea Institute of Science and Technology(韩国科学技术院) Kyung Hee University(庆熙大学)

专题命中 红队测试 :jailbreak(title,abstract);red teaming(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ACL 2025 (Main Track). Camera-ready version

Journal ref Proc. ACL 2025 (Vol. 1: Long Papers), pp. 16489-16507, Vienna, Austria, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 3 篇

2508.03092 2025-08-06 cs.AI cs.CL 62%

Toward Verifiable Misinformation Detection: A Multi-Tool LLM Agent Framework

Zikun Cui, Tianyi Huang, Chia-En Chiang, Cuiqianhe Du

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12591 2025-08-06 cs.CV cs.CL 57%

CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base

Cong-Duy Nguyen, Xiaobao Wu, Duc Anh Vu, Shuai Zhao, Thong Nguyen, Anh Tuan Luu

机构 * Nanyang Technological University, Singapore(南洋理工大学) National University of Singapore, Singapore(国立新加坡大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03007 2025-08-06 cs.CV 50%

Multi-Granularity Feature Calibration via VFM for Domain Generalized Semantic Segmentation

Xinhui Li, Xiaojie Guo

机构 * Xinhui Li(李新会) Xiaojie Guo(郭小杰)

专题命中 幻觉与事实性 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 1 篇

2508.03101 2025-08-06 cs.NI cs.AI cs.MA 57%

Using the NANDA Index Architecture in Practice: An Enterprise Perspective

Sichao Wang, Ramesh Raskar, Mahesh Lambe, Pradyumna Chari, Rekha Singhal, Shailja Gupta, Rajesh Ranjan, Ken Huang

机构 * Cisco Systems(思科系统) MIT(麻省理工学院) Unify Dynamics Tata Consultancy Services(塔塔咨询公司) CMU(卡内基梅隆大学)

专题命中 隐私与版权 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 16 篇

2409.19177 2025-08-06 cs.LG cs.CL cs.CY 85%

Evidence Is All You Need: Ordering Imaging Studies via Language Model Alignment with the ACR Appropriateness Criteria

Michael S. Yao, Allison Chae, Charles E. Kahn, Walter R. Witschey, James C. Gee, Hersh Sagreiya, Osbert Bastani

专题命中 安全评测 :alignment(title,abstract);trustworthy(abstract);分类 cs.CL、cs.CY、cs.LG

Comments 15 pages main text, 4 figures, 1 table

Journal ref Commun Med 5, 332 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03117 2025-08-06 cs.AI 79%

Toward a Trustworthy Optimization Modeling Agent via Verifiable Synthetic Data Generation

Vinicius Lima, Dzung T. Phan, Jayant Kalagnanam, Dhaval Patel, Nianjun Zhou

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02994 2025-08-06 cs.AI 77%

When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs

Fangyi Yu

机构 * Fangyi Yu

专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00159 2025-08-06 cs.AI cs.CY cs.LG econ.TH math.OC 75%

Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power

Jobst Heitzig, Ram Potham

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01263 2025-08-06 cs.CL cs.AI cs.LO 62%

Bridging LLMs and Symbolic Reasoning in Educational QA Systems: Insights from the XAI Challenge at IJCNN 2025

Long S. T. Nguyen, Khang H. N. Vo, Thu H. A. Nguyen, Tuan C. Bui, Duc Q. Nguyen, Thanh-Tung Tran, Anh D. Nguyen, Minh L. Nguyen, Fabien Baldacci, Thang H. Bui, Emanuel Di Nardo, Angelo Ciaramella, Son H. Le, Ihsan Ullah, Lorenzo Di Rocco, Tho T. Quan

机构 * URA Research Group, Ho Chi Minh City University of Technology (HCMUT), Vietnam Ho Chi Minh City International University (HCMIU), Vietnam University of South-Eastern Norway, Norway Japan Advanced Institute of Science Univ. Bordeaux, CNRS, Bordeaux INP, LaBRI, UMR 5800, F-33400 Talence, France University of Naples Parthenope, Italy VNU Information Technology Institute, Vietnam National University, Vietnam Visual Intelligence Lab, School of Computer Science \& Insight Center for Data Analyitcs, University of Galway, Ireland Sapienza University of Rome, Italy

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments The XAI Challenge @ TRNS-AI Workshop, IJCNN 2025: Explainable AI for Educational Question Answering. Website: https://sites.google.com/view/trns-ai/challenge/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14971 2025-08-06 cs.AI cs.CL cs.SD eess.AS 62%

BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation

Jilong Li, Zhenxi Song, Jiaqi Wang, Meishan Zhang, Honghai Liu, Min Zhang, Zhiguo Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Peng Cheng Laboratory, China(鹏城实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 8 pages (excluding references), accepted by Findings of ACL 2025

Journal ref Findings of the Association for Computational Linguistics: ACL 2025, pages 2762-2778, July 2025, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14961 2025-08-06 cs.LG cs.CR cs.LO 57%

Set-Based Training for Neural Network Verification

Lukas Koller, Tobias Ladner, Matthias Althoff

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments published at Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02947 2025-08-06 cs.RO cs.AI 57%

AeroSafe: Mobile Indoor Air Purification using Aerosol Residence Time Analysis and Robotic Cough Emulator Testbed

M Tanjid Hasan Tonmoy, Rahath Malladi, Kaustubh Singh, Forsad Al Hossain, Rajesh Gupta, Andrés E. Tejada-Martínez, Tauhidur Rahman

机构 * University of California San Diego(加州大学圣地亚哥分校) Plaksha University(Plaksha大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) University of South Florida(佛罗里达州立大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Accepted at IEEE International Conference on Robotics and Automation (ICRA) 2025. Author Accepted Manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02827 2025-08-06 cs.SE cs.AI 57%

Automated Validation of LLM-based Evaluators for Software Engineering Artifacts

Ora Nova Fandina, Eitan Farchi, Shmulik Froimovich, Rami Katan, Alice Podolsky, Orna Raz, Avi Ziv

机构 * IBM Research(IBM研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18905 2025-08-06 cs.CL cs.HC 57%

Large language models provide unsafe answers to patient-posed medical questions

Rachel L. Draelos, Samina Afreen, Barbara Blasko, Tiffany L. Brazile, Natasha Chase, Dimple Patel Desai, Jessica Evert, Heather L. Gardner, Lauren Herrmann, Aswathy Vaikom House, Stephanie Kass, Marianne Kavan, Kirshma Khemani, Amanda Koire, Lauren M. McDonald, Zahraa Rabeeah, Amy Shah

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05347 2025-08-06 cs.CL cs.MA 57%

GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report Evaluation

Zhenxuan Zhang, Kinhei Lee, Peiyuan Jing, Weihang Deng, Huichi Zhou, Zihao Jin, Jiahao Huang, Zhifan Gao, Dominic C Marshall, Yingying Fang, Guang Yang

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20847 2025-08-06 cs.CY 57%

Who Should Run Advanced AI Evaluations -- AISIs?

Merlin Stein, Milan Gandhi, Theresa Kriecherbauer, Amin Oueslati, Robert Trager

专题命中 安全评测 :safety(abstract);分类 cs.CY

Comments Accepted to AIES 2024 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03374 2025-08-06 cs.CV 50%

GRASPing Anatomy to Improve Pathology Segmentation

Keyi Li, Alexander Jaus, Jens Kleesiek, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology, Karlsruhe, Germany(卡尔斯鲁厄理工学院) Helmholtz Information and Data Science School for Health, Germany(海德堡信息与数据科学健康学校) Department of Nuclear Medicine, University of Duisburg-Essen(杜伊斯堡-埃森大学核医学系) German Cancer Consortium (DKTK)- University Hospital Essen, Essen, Germany(德国癌症研究中心(DKTK)-埃森大学医院)

专题命中 安全评测 :alignment(abstract)

Comments Accepted at 16th MICCAI Workshop on Machine Learning in Medical Imaging (MLMI2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02962 2025-08-06 cs.RO 50%

Robot builds a robot's brain: AI generated drone command and control station hosted in the sky

Peter Burke

机构 * Department of Electrical Engineering and Computer Science, University of California, Irvine(电气工程与计算机科学系,加州大学伊文斯顿分校)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22053 2025-08-06 cs.SD cs.MA cs.MM eess.AS 50%

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation

Yan Rong, Jinting Wang, Guangzhi Lei, Shan Yang, Li Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Tencent AI Lab(腾讯人工智能实验室)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20466 2025-08-06 cs.CV 50%

LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs

Woo Yi Yang, Jiarui Wang, Sijing Wu, Huiyu Duan, Yuxin Zhu, Liu Yang, Kang Fu, Guangtao Zhai, Xiongkuo Min

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 2 篇

2508.03292 2025-08-06 cs.CL cs.AI 62%

Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes

Shahed Masoudian, Gustavo Escobedo, Hannah Strauss, Markus Schedl

机构 * Johannes Kepler University (JKU)(约翰内斯·开普勒大学) Linz Institute of Technology (LIT)(林茨技术研究所) University of Innsbruck(因斯布鲁克大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09053 2025-08-06 cs.AI cs.GT cs.LG 62%

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers

Haoran Sun, Yusen Wu, Peng Wang, Wei Chen, Yukun Cheng, Xiaotie Deng, Xu Chu

机构 * CFCS, School of Computer Science, Peking University(计算机科学系,北京大学) School of Business, Jiangnan University(商学院,江南大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments A shorter conference version is published in IJCAI 2025, titled 'Game Theory Meets Large Language Models: A Systematic Survey'

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 8 篇

2508.02823 2025-08-06 cs.HC cs.AI cs.CL cs.SE 62%

NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding Modification

Wenshuo Zhang, Leixian Shen, Shuchang Xu, Jindu Wang, Jian Zhao, Huamin Qu, Linping Yuan

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Waterloo(滑铁卢大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted in UIST 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03651 2025-08-06 cs.HC cs.AI 57%

Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired

Ruei-Che Chang, Rosiana Natalie, Wenqian Xu, Jovan Zheng Feng Yap, Anhong Guo

机构 * University of Michigan(密歇根大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments ACM ASSETS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏