arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-11 至 2025-08-11 共收录 38 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2508.06036 2025-08-11 cs.CV 71%

More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment

Jun Xie, Yingjian Zhu, Feng Chen, Zhenghao Zhang, Xiaohui Fan, Hongzhu Yi, Xinming Wang, Chen Yu, Yue Bi, Zhaoran Zhao, Xiongjun Guan, Zhepeng Wang

机构 * Lenovo Research(联想研究) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Tsinghua University(清华大学) Beijing Jiaotong University(北京交通大学) Shandong University(山东大学)

专题命中 偏好对齐 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06026 2025-08-11 cs.CL cs.AI 62%

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future

Yidong Wang, Xin Wang, Cunxiang Wang, Junfeng Fang, Qiufeng Wang, Jianing Chu, Xuran Meng, Shuxun Yang, Libo Qin, Yue Zhang, Wei Ye, Shikun Zhang

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05434 2025-08-11 cs.LG 57%

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling

Han Qi, Haochen Yang, Qiaosheng Zhang, Zhuoran Yang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Xi’an Jiaotong University(西安交通大学) Peking University(北京大学) Yale University(耶鲁大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05660 2025-08-11 cs.IR cs.AI 57%

Open-Source Agentic Hybrid RAG Framework for Scientific Literature Review

Aditya Nagori, Ricardo Accorsi Casonatto, Ayush Gautam, Abhinav Manikantha Sai Cheruvu, Rishikesan Kamaleswaran

机构 * Department of Surgery, Department of Anesthesiology, Duke University School of Medicine Durham North Carolina United States Faculty of Technology, University of Brasilia Brasilia Federal District Brazil Indian Institute of Technology Goa Goa India Birla Institute of Technology \& Science Pilani Hyderabad India Department of Electrical Computer Engineering, Duke University Pratt School of Engineering Department of Surgery, Department of Anesthesiology, Duke University School of Medicine Durham North Carolina United States Department of Surgery, Department of Anesthesiology, Duke University School of Medicine Faculty of Technology, University of Brasilia Indian Institute of Technology Goa Birla Institute of Technology \& Science Pilani Computer Engineering, Duke University Pratt School of Engineering

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2508.05766 2025-08-11 cs.AI cs.LG cs.SY eess.SY nlin.AO 79%

A Framework for Inherently Safer AGI through Language-Mediated Active Inference

Bo Wen

机构 * IBM T.J. Watson Research Center(IBM T.J. Watson研究院)

专题命中 安全训练 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02539 2025-08-11 cs.LG 70%

VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents

Thong Q. Nguyen, Shubhang Desai, Raja Hasnain Anwar, Firoz Shaik, Vishwas Suryanarayanan, Vishal Chowdhary

机构 * Microsoft(微软公司) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 提示注入 1 篇

2508.06418 2025-08-11 cs.CL 57%

Quantifying Conversation Drift in MCP via Latent Polytope

Haoran Shi, Hongwei Yao, Shuo Shao, Shaopeng Jiao, Ziqi Peng, Zhan Qin, Cong Wang

专题命中 提示注入 :prompt injection(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 2 篇

2508.06017 2025-08-11 cs.SE cs.CL cs.LG 62%

Position: Intelligent Coding Systems Should Write Programs with Justifications

Xiangzhe Xu, Shiwei Feng, Zian Su, Chengpeng Wang, Xiangyu Zhang

机构 * Purdue University(普渡大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.LG

Comments The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06306 2025-08-11 cs.CL cs.AI cs.HC 62%

Humans overrely on overconfident language models, across languages

Neil Rathi, Dan Jurafsky, Kaitlyn Zhou

机构 * Stanford University(斯坦福大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

Comments camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 1 篇

2412.05734 2025-08-11 cs.CR cs.AI cs.LG 73%

LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage

Yuzhou Nie, Zhun Wang, Ye Yu, Xian Wu, Xuandong Zhao, Wenbo Guo, Dawn Song

专题命中 隐私与版权 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 12 篇

2508.06124 2025-08-11 cs.CL 83%

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models

Sayantan Adak, Pratyush Chatterjee, Somnath Banerjee, Rima Hazra, Somak Aditya, Animesh Mukherjee

专题命中 安全评测 :alignment(title,abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05855 2025-08-11 cs.AI cs.RO 79%

Safety of Embodied Navigation: A Survey

Zixia Wang, Jia Hu, Ronghui Mu

机构 * University of Exeter(埃克塞特大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03834 2025-08-11 cs.RO cs.CV 78%

CARE: Enhancing Safety of Visual Navigation through Collision Avoidance via Repulsive Estimation

Joonkyung Kim, Joonyeol Sim, Woojun Kim, Katia Sycara, Changjoo Nam

机构 * Department of Electronic Engineering, Sogang University(电子工程系,首尔大学) Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学)

专题命中 安全评测 :safety(title,abstract)

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00381 2025-08-11 cs.CV cs.AI cs.CE cs.LG 73%

Advancing Welding Defect Detection in Maritime Operations via Adapt-WeldNet and Defect Detection Interpretability Analysis

Kamal Basha S, Athira Nambiar

机构 * Department of Computational Intelligence, Faculty of Engineering and Technology, SRM Institute of Science and Technology(计算智能系,工程与技术学院,SRM科学与技术学院)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13677 2025-08-11 cs.CL cs.AI cs.LG 67%

Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results

Andrea Santilli, Adam Golinski, Michael Kirchhof, Federico Danieli, Arno Blaas, Miao Xiong, Luca Zappella, Sinead Williamson

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14119 2025-08-11 cs.CL cs.AI 62%

Autonomous Structural Memory Manipulation for Large Language Models Using Hierarchical Embedding Augmentation

Derek Yotheringhay, Alistair Kirkland, Humphrey Kirkbride, Josiah Whitesteeple

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11417 2025-08-11 cs.CL cs.AI 62%

Neural Contextual Reinforcement Framework for Logical Structure Language Generation

Marcus Irvin, William Cooper, Edward Hughes, Jessica Morgan, Christopher Hamilton

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06155 2025-08-11 cs.CL 57%

Semantic and Structural Analysis of Implicit Biases in Large Language Models: An Interpretable Approach

Renhan Zhang, Lian Lian, Zhen Qi, Guiran Liu

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05987 2025-08-11 cs.CL 57%

Adversarial Topic-aware Prompt-tuning for Cross-topic Automated Essay Scoring

Chunyun Zhang, Hongyan Zhao, Chaoran Cui, Qilong Song, Zhiqing Lu, Shuai Gong, Kailin Liu

机构 * Shandong University of Finance and Economics(山东财经大学) University of Toronto(多伦多大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08947 2025-08-11 cs.CL 57%

Structured Convergence in Large Language Model Representations via Hierarchical Latent Space Folding

Fenella Harcourt, Naderdel Piero, Gilbert Sutherland, Daphne Holloway, Harriet Bracknell, Julian Ormsby

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05687 2025-08-11 cs.MA cs.AI 57%

Risk Analysis Techniques for Governed LLM-based Multi-Agent Systems

Alistair Reid, Simon O'Callaghan, Liam Carroll, Tiberio Caetano

机构 * Gradient Institute Ltd.(梯度研究所有限公司)

专题命中 安全评测 :red teaming(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06152 2025-08-11 cs.CV 50%

VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation

Kaiyuan Jiang, Ruoxi Sun, Ying Cao, Yuqi Xu, Xinran Zhang, Junyan Guo, ChengSheng Deng

机构 * Peking University(北京大学) LinkSure University of Glasgow(格拉斯哥大学) Boston University(波士顿大学)

专题命中 安全评测 :alignment(abstract)

Comments 17 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 4 篇

2508.05846 2025-08-11 cs.CY cs.AI cs.HC cs.LG cs.RO 82%

Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems

Ahmad Farooq, Kamran Iqbal

机构 * University of Arkansas at Little Rock(阿拉巴马州立大学)

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG

Comments Published in the Proceedings of the 2025 3rd International Conference on Robotics, Control and Vision Engineering (RCVE'25). 6 pages, 3 tables

Journal ref RCVE'25: Proceedings of the 2025 3rd International Conference on Robotics, Control and Vision Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05938 2025-08-11 cs.CL cs.AI cs.CY 75%

Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale

Rafal Kocielnik, Min Kim, Penphob, Boonyarungsrit, Fereshteh Soltani, Deshawn Sambrano, Animashree Anandkumar, R. Michael Alvarez

机构 * California Institute of Technology(加利福尼亚理工学院) Activision Publishing, Inc.(暴雪娱乐公司)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 9 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05913 2025-08-11 cs.HC cs.AI cs.CL 62%

Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction

Stefan Pasch, Min Chul Cha

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06479 2025-08-11 cs.CY 57%

The Problem of Atypicality in LLM-Powered Psychiatry

Bosco Garcia, Eugene Y. S. Chua, Harman Singh Brah

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Preprint of 8/8/2025 -- please cite published version. This article has been published in the Journal of Medical Ethics (2025) following peer review and can also be viewed on the journal's website at 10.1136/jme-2025-110972

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 12 篇

2502.09815 2025-08-11 cs.CL 79%

Statistical Coherence Alignment for Large Language Model Representation Learning Through Tensor Field Convergence

Jonathan Gale, Godfrey Aldington, Harriet Thistlewood, Thomas Tattershall, Basil Wentworth, Vincent Enoasmo

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06300 2025-08-11 cs.HC 78%

Automatic Semantic Alignment of Flow Pattern Representations for Exploration with Large Language Models

Weihan Zhang, Jun Tao

专题命中 其他安全 :alignment(title,abstract)

Comments Accepted by IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03155 2025-08-11 cs.LG cs.AI 62%

Fusing Cross-Domain Knowledge from Multimodal Data to Solve Problems in the Physical World

Yu Zheng

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16658 2025-08-11 cs.CL cs.AI 62%

Contextual Reinforcement in Multimodal Token Compression for Large Language Models

Naderdel Piero, Zacharias Cromwell, Nathaniel Wainwright, Matthias Nethercott

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏