arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-30 至 2025-10-30 共收录 35 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 6 篇

2510.23965 2025-10-30 cs.AI cs.LG stat.ML 84%

The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity

Ali Aouad, Aymane El Gadarri, Vivek F. Farias

机构 * MIT(麻省理工学院)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01183 2025-10-30 cs.LG cs.AI stat.ML 81%

Doubly Robust Alignment for Large Language Models

Erhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu, Francesco Quinzan, Chengchun Shi

机构 * Department of Statistics(统计系) LSE London, UK(伦敦大学学院) Department of Mathematics(数学系) Tsinghua University(清华大学) School of Design LCC, UAL London, UK(伦敦艺术大学设计学院) Department of Engineering Science(工程科学系) University of Oxford(牛津大学)

专题命中 偏好对齐 :alignment(title);RLHF(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03690 2025-10-30 cs.CL 77%

Robust Preference Optimization via Dynamic Target Margins

Jie Sun, Junkang Wu, Jiancan Wu, Zhibo Zhu, Xingyu Lu, Jun Zhou, Lintao Ma, Xiang Wang

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);safety(abstract);分类 cs.CL

Comments 18 pages, 6 figures, accepted to Findings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02745 2025-10-30 cs.AI cs.CL 76%

CURATRON: Complete and Robust Preference Data for Rigorous Alignment of Large Language Models

Son The Nguyen, Niranjan Uma Naresh, Theja Tulabandhula

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Independent Researcher(独立研究者)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01308 2025-10-30 cs.AI cs.CL cs.DB 62%

GradeSQL: Test-Time Inference with Outcome Reward Models for Text-to-SQL Generation from Large Language Models

Mattia Tritto, Giuseppe Farano, Dario Di Palma, Gaetano Rossiello, Fedelucio Narducci, Dharmashankar Subramanian, Tommaso Di Noia

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17220 2025-10-30 cs.CL 61%

RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness

Tianyu Yu, Haoye Zhang, Qiming Li, Qixin Xu, Yuan Yao, Da Chen, Xiaoman Lu, Ganqu Cui, Yunkai Dang, Taiwen He, Xiaocheng Feng, Jun Song, Bo Zheng, Zhiyuan Liu, Tat-Seng Chua, Maosong Sun

机构 * Tsinghua University(清华大学) Shanghai Qi Zhi Institute(上海启智研究院) Harbin Institute of Technology(哈尔滨工业大学) Taobao & Tmall Group of Alibaba(阿里巴巴淘宝与天猫集团) Peng Cheng Laboratory(鹏城实验室) National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL;RLHF(comments)

Comments Project Website: https://github.com/RLHF-V/RLAIF-V

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2510.24820 2025-10-30 cs.CV cs.AI 83%

SafeEditor: Unified MLLM for Efficient Post-hoc T2I Safety Editing

Ruiyang Zhang, Jiahao Luo, Xiaoru Feng, Qiufan Pang, Yaodong Yang, Juntao Dai

机构 * PKU Alignment Team, Peking University(北京大学对齐团队) LLM Safety Centre, Beijing Academy of Artificial Intelligence(北京人工智能研究院大语言模型安全中心)

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25179 2025-10-30 cs.AI 77%

Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University, Australia(计算机学院,麦考瑞大学,澳大利亚)

专题命中 安全训练 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 幻觉与事实性 1 篇

2412.15189 2025-10-30 cs.CL cs.CY 62%

Face the Facts! Evaluating RAG-based Pipelines for Professional Fact-Checking

Daniel Russo, Stefano Menini, Jacopo Staiano, Marco Guerini

机构 * Fondazione Bruno Kessler(布罗诺·凯塞勒基金会) University of Trento(特伦托大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.CY

Comments Code and data at https://github.com/drusso98/face-the-facts - Accepted for publication at INLG 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 安全评测 14 篇

2506.14866 2025-10-30 cs.SE cs.LG 83%

OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents

Thomas Kuntz, Agatha Duzan, Hao Zhao, Francesco Croce, Zico Kolter, Nicolas Flammarion, Maksym Andriushchenko

机构 * EPFL(苏黎世联邦理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :safety(title,abstract);prompt injection(abstract);分类 cs.LG

Comments NeurIPS 2025 Datasets & Benchmarks Track (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06497 2025-10-30 cs.CV cs.AI 79%

Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving

Enming Zhang, Peizhe Gong, Xingyuan Dai, Min Huang, Yisheng Lv, Qinghai Miao

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19169 2025-10-30 cs.CR cs.CL 70%

OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models

Thomas Wang, Haowen Li

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 安全评测 :safety(abstract);prompt injection(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24898 2025-10-30 eess.SY cs.SY 67%

Delay Tolerant Control for Autonomous Driving Using CDOB

Xincheng Cao, Haochong Chen, Levent Guvenc, Bilin Aksun-Guvenc

专题命中 安全评测 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24811 2025-10-30 cs.CL cs.AI cs.LG 67%

ProofSketch: Efficient Verified Reasoning for Large Language Models

Disha Sheshanarayana, Tanishka Magar

机构 * Manipal University Jaipur(马普尔大学斋普尔)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at NeurIPS 2025, ER Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25771 2025-10-30 cs.CL cs.AI 62%

Gaperon: A Peppered English-French Generative Language Model Suite

Nathan Godey, Wissam Antoun, Rian Touchent, Rachel Bawden, Éric de la Clergerie, Benoît Sagot, Djamé Seddah

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25571 2025-10-30 cs.LG cs.DS cs.NA math.NA math.SP math.ST stat.TH 57%

Perturbation Bounds for Low-Rank Inverse Approximations under Noise

Phuc Tran, Nisheeth K. Vishnoi

机构 * Yale University(耶鲁大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25413 2025-10-30 cs.CL 57%

Seeing, Signing, and Saying: A Vision-Language Model-Assisted Pipeline for Sign Language Data Acquisition and Curation from Social Media

Shakib Yazdani, Yasser Hamidullah, Cristina España-Bonet, Josef van Genabith

机构 * German Research Center for Artificial Intelligence (DFKI GmbH)(德国人工智能研究中心(DFKI GmbH)) Saarland Informatics Campus(萨尔兰州信息技术校区) Barcelona Supercomputing Center (BSC-CNS)(巴塞罗那超级计算中心(BSC-CNS))

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted by RANLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25091 2025-10-30 cs.AI 57%

H3M-SSMoEs: Hypergraph-based Multimodal Learning with LLM Reasoning and Style-Structured Mixture of Experts

Peilin Tan, Liang Xie, Churan Zhi, Dian Tu, Chuanqi Shi

机构 * University of California, San Diego(加州大学圣迭戈分校) Wuhan University of Technology(武汉科技大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24951 2025-10-30 cs.LG cs.AR cs.NE 57%

Resource-Efficient and Robust Inference of Deep and Bayesian Neural Networks on Embedded and Analog Computing Platforms

Bernhard Klein

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Ph.D. dissertation, Heidelberg University, October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24749 2025-10-30 cs.SE cs.AI 57%

Beyond Function-Level Search: Repository-Aware Dual-Encoder Code Retrieval with Adversarial Verification

Aofan Liu, Shiyuan Song, Haoxuan Li, Cehao Yang, Yiyan Qi

机构 * International Digital Economy Academy (IDEA)(国际数字经济学院) School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10266 2025-10-30 cs.CV cs.AI 57%

SignMouth: Leveraging Mouthing Cues for Sign Language Translation by Multimodal Contrastive Fusion

Wenfang Wu, Tingting Yuan, Yupeng Li, Daling Wang, Xiaoming Fu

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24999 2025-10-30 cs.CR 50%

SLIP-SEC: Formalizing Secure Protocols for Model IP Protection

Racchit Jain, Satya Lokam, Yehonathan Refael, Adam Hakim, Lev Greenberg, Jay Tenenbaum

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03089 2025-10-30 cs.CV q-bio.NC 50%

Explicitly Modeling Subcortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness

Lucas Piper, Arlindo L. Oliveira, Tiago Marques

机构 * INESC-ID Instituto Superior Técnico(理工学院) Universidade de Lisboa(里斯本大学) Breast Cancer Research Program(乳腺癌研究计划) Champalimaud Foundation(恰帕拉马德基金会) Faculdade de Medicina de Lisboa(里斯本医学院)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. AI治理与伦理 3 篇

2510.24721 2025-10-30 cs.CY cs.AI cs.CL cs.HC 75%

The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims

Matthew Kelly

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 65 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25445 2025-10-30 cs.AI cs.LG 73%

Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions

Mohamad Abou Ali, Fadi Dornaika

机构 * University of the Basque Country(巴斯克大学) Lebanese International University (LIU)(黎巴嫩国际大学) The International University of Beirut(贝鲁特国际大学) IKERBASQUE

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25218 2025-10-30 cs.CY cs.AI 62%

Human Resilience in the AI Era -- What Machines Can't Replace

Shaoshan Liu, Anina Schwarzenbach, Yiyu Shi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 其他安全 9 篇

2510.14205 2025-10-30 cs.CL cs.AI 81%

DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans

Bingsheng Yao, Bo Sun, Yuanzhe Dong, Yuxuan Lu, Dakuo Wang

机构 * Northeastern University(东北大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments In Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19311 2025-10-30 cs.CV cs.AI 79%

DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment

Weizhi Chen, Yupeng Deng, Jin Wei, Jingbo Chen, Jiansheng Chen, Yuman Feng, Zhihao Xi, Diyou Liu, Kai Li, Yu Meng

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院 aerospace information research institute) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) School of Information Network Security, People’s Public Security University of China(中国人民公安大学信息网络安全学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00814 2025-10-30 cs.CL cs.AI cs.CY 67%

Many LLMs Are More Utilitarian Than One

Anita Keshmirian, Razan Baltaji, Babak Hemmatian, Hadi Asghari, Lav R. Varshney

机构 * Forward College(前进学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Nebraska, Lincoln(内布拉斯加大学林肯分校) Technische Universität Berlin(柏林技术大学) Humboldt Institute for Internet and Society(洪堡互联网与社会研究所) Stony Brook University(石溪大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to the Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22149 2025-10-30 cs.AI cs.LG 62%

When Truthful Representations Flip Under Deceptive Instructions?

Xianxuan Long, Yao Fu, Runchao Li, Mu Sheng, Haotian Yu, Xiaotian Han, Pan Li

机构 * Case Western Reserve University(凯斯西储大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏