arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2410.19238 2025-11-17 cs.AI cs.CY 62%

Designing AI-Agents with Personalities: A Psychometric Approach

Muhua Huang, Xijuan Zhang, Christopher Soto, James Evans

机构 * Stanford University(斯坦福大学) University of Chicago Knowledge Lab(芝加哥大学知识实验室) Chicago Center for Computational Social Science(芝加哥计算社会科学中心) York University(约克大学) Colby College(科尔比学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09748 2025-11-14 cs.CL cs.AI 62%

How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation

Muskaan Chopra, Lorenz Sparrenberg, Sarthak Khanna, Rafet Sifa

机构 * Fraunhofer IAIS - Department of Media Engineering(弗劳恩霍夫人工智能研究所-媒体工程部门) University of Bonn - Department of Computer Science(波恩大学-计算机科学系) Lamarr Institute for Machine Learning(拉马尔人工智能与机器学习研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted in IEEE BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09067 2025-11-13 cs.CL cs.AI 62%

MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique

Gailun Zeng, Ziyang Luo, Hongzhan Lin, Yuchen Tian, Kaixin Li, Ziyang Gong, Jianxiong Guo, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist 大学) National University of Singapore(新加坡国立大学) Beijing Normal University(北京师范大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 28 pages, 14 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09363 2025-11-13 cs.AI cs.GT cs.LG 62%

ElicitationGPT: Text Elicitation Mechanisms via Language Models

Yifan Wu, Jason Hartline

机构 * Microsoft Research(微软研究院) Northwestern University(西北大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07982 2025-11-12 cs.CL cs.AI 62%

NOTAM-Evolve: A Knowledge-Guided Self-Evolving Optimization Framework with LLMs for NOTAM Interpretation

Maoqi Liu, Quan Fang, Yuhao Wu, Can Zhao, Yang Yang, Kaiquan Cai

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07659 2025-11-12 cs.CL cs.AI 62%

Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering

Sai Shridhar Balamurali, Lu Cheng

机构 * University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07070 2025-11-11 cs.AI cs.LG 62%

RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services

Fei Zhao, Chonggang Lu, Haofu Qian, Fangcheng Shi, Zijie Meng, Jianzhao Huang, Xu Tang, Zheyong Xie, Zheyu Ye, Zhe Xu, Yao Hu, Shaosheng Cao

机构 * NLP Team, Xiaohongshu Inc.(小红书研究院自然语言处理团队)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06763 2025-11-11 cs.CL cs.AI 62%

Sensitivity of Small Language Models to Fine-tuning Data Contamination

Nicy Scaria, Silvester John Joseph Kennedy, Deepak Subramani

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05527 2025-11-11 cs.CL cs.AI 62%

Bridging Industrial Expertise and XR with LLM-Powered Conversational Agents

Despina Tomkou, George Fatouros, Andreas Andreou, Georgios Makridis, Fotis Liarokapis, Dimitrios Dardanis, Athanasios Kiourtis, John Soldatos, Dimosthenis Kyriazis

机构 * Innov-Acts Ltd.(Innov-Acts有限公司) CYENS Centre of Excellence(CYENS卓越中心) University of Piraeus(比雷埃克斯大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 7 figures

Journal ref 2025 21st International Conference on Distributed Computing in Smart Systems and the Internet of Things (DCOSS-IoT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06470 2025-11-11 cs.AI cs.LG 62%

Brain-Inspired Planning for Better Generalization in Reinforcement Learning

Mingde "Harry" Zhao

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments McGill PhD Thesis (updated on 20251109 for typos and margin adjustments)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06248 2025-11-11 cs.LG cs.AI 62%

Constraint-Informed Active Learning for End-to-End ACOPF Optimization Proxies

Miao Li, Michael Klamkin, Pascal Van Hentenryck, Wenting Li, Russell Bent

机构 * University of Texas at Austin, TX, USA(德克萨斯大学奥斯汀分校)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 8 PAGES

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05516 2025-11-11 cs.CL cs.AI cs.SD eess.AS 62%

Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation

Canxiang Yan, Chunxiang Jin, Dawei Huang, Haibing Yu, Han Peng, Hui Zhan, Jie Gao, Jing Peng, Jingdong Chen, Jun Zhou, Kaimeng Ren, Ming Yang, Mingxue Yang, Qiang Xu, Qin Zhao, Ruijie Xiong, Shaoxiong Lin, Xuezhi Wang, Yi Yuan, Yifei Wu, Yongjie Lyu, Zhengyu He, Zhihao Qiu, Zhiqiang Fang, Ziyuan Huang

机构 * Inclusion AI Ant Group(Inclusion AI Ant集团)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 32 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00710 2025-11-11 cs.AI cs.CL 62%

On Verifiable Legal Reasoning: A Multi-Agent Framework with Formalized Knowledge Representations

Albert Sadowski, Jarosław A. Chudziak

机构 * Warsaw University of Technology(华沙技术大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication at the 34th ACM International Conference on Information and Knowledge Management (CIKM '25)

Journal ref CIKM '25: Proceedings of the 34th ACM International Conference on Information and Knowledge Management (2025) 2535-2545

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17797 2025-11-10 cs.CL cs.AI 62%

Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics

Akshara Prabhakar, Roshan Ram, Zixiang Chen, Silvio Savarese, Frank Wang, Caiming Xiong, Huan Wang, Weiran Yao

机构 * Salesforce AI Research(Salesforce AI研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Technical report; 13 pages plus references and appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04956 2025-11-10 cs.AI cs.CL 62%

ORCHID: Orchestrated Retrieval-Augmented Classification with Human-in-the-Loop Intelligent Decision-Making for High-Risk Property

Maria Mahbub, Vanessa Lama, Sanjay Das, Brian Starks, Christopher Polchek, Saffell Silvers, Lauren Deck, Prasanna Balaprakash, Tirthankar Ghosal

机构 * Oak Ridge National Laboratory(橡树岭国家实验室) Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04703 2025-11-10 cs.CL cs.AI 62%

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Rose Kirk, Fangru Lin, Gabrielle Kaili-May Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yushi Yang, Yilun Zhao, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip H. S. Torr, Cozmin Ududec, Luc Rocher, Adam Mahdi

机构 * University of Oxford(牛津大学) EPFL(苏黎世联邦理工学院) Weizenbaum Institute Berlin(柏林Weizenbaum研究所) Technical University Munich(慕尼黑技术大学) Centre for Digital Governance, Hertie School(赫尔姆霍兹学院数字治理中心) Stanford University(斯坦福大学) UK AI Security Institute(英国人工智能安全研究所) SomosNLP Universdad Politécnica de Madrid(马德里理工大学) Yale University(耶鲁大学) Allen Institute for AI(人工智能研究所) University of Washington(华盛顿大学) Meedan UC Berkeley(加州大学伯克利分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03945 2025-11-07 cs.CL cs.AI 62%

Direct Semantic Communication Between Large Language Models via Vector Translation

Fu-Chun Yang, Jason Eshraghian

机构 * University of California, Santa Cruz(加州大学圣克ruz分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10707 2025-11-04 cs.LG cs.AI 62%

ConTextTab: A Semantics-Aware Tabular In-Context Learner

Marco Spinaci, Marek Polewczyk, Maximilian Schambach, Sam Thelin

机构 * SAP France(SAP法国分公司) SAP SE(SAP德国分公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted as spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14622 2025-11-04 cs.CR cs.AI cs.LG 62%

Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection

Yihao Guo, Haocheng Bian, Liutong Zhou, Ze Wang, Zhaoyi Zhang, Francois Kawala, Milan Dean, Ian Fischer, Yuantao Peng, Noyan Tokgozoglu, Ivan Barrientos, Riyaaz Shaik, Rachel Li, Chandru Venkataraman, Reza Shifteh Far, Moses Pawar, Venkat Sundaranatha, Michael Xu, Frank Chu

机构 * Apple(苹果公司) Cohere(Cohere公司) DeepMind(深度思维公司) Meta MongoDB(MongoDB公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27207 2025-11-03 cs.LG cs.AI 62%

Feature-Function Curvature Analysis: A Geometric Framework for Explaining Differentiable Models

Hamed Najafi, Dongsheng Luo, Jason Liu

机构 * Florida International University(佛罗里达国际大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02927 2025-10-31 eess.SY cs.AI cs.LG cs.SY 62%

Multivariate Physics-Informed Convolutional Autoencoder for Anomaly Detection in Power Distribution Systems with High Penetration of DERs

Mehdi Jabbari Zideh, Sarika Khushalani Solanki

机构 * Lane Department of Computer Science and Electrical Engineering, West Virginia University(计算机科学与电气工程系,西弗吉尼亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref Sustainable Energy, Grids and Networks, Vol. 44, December 2025, 102022

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26402 2025-10-31 cs.AI cs.LG 62%

Autograder+: A Multi-Faceted AI Framework for Rich Pedagogical Feedback in Programming Education

Vikrant Sahu, Gagan Raj Gupta, Raghav Borikar, Nitin Mane

机构 * Indian Institute of Technology(印度理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26037 2025-10-31 cs.CR cs.AI cs.CL 62%

SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning

Kaiwen Zhou, Ahmed Elgohary, A S M Iftekhar, Amin Saied

机构 * Microsoft Responsible AI Research(微软负责任人工智能研究) University of California(加州大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21497 2025-10-31 cs.CV cs.AI cs.CL cs.MA 62%

Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers

Wei Pang, Kevin Qinghong Lin, Xiangru Jian, Xi He, Philip Torr

机构 * University of Waterloo(滑铁卢大学) University of Oxford(牛津大学) Vector Institute(向量研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Project Page: https://github.com/Paper2Poster/Paper2Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25771 2025-10-30 cs.CL cs.AI 62%

Gaperon: A Peppered English-French Generative Language Model Suite

Nathan Godey, Wissam Antoun, Rian Touchent, Rachel Bawden, Éric de la Clergerie, Benoît Sagot, Djamé Seddah

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24488 2025-10-29 cs.CL cs.AI 62%

A word association network methodology for evaluating implicit biases in LLMs compared to humans

Katherine Abramski, Giulio Rossetti, Massimo Stella

机构 * University of Pisa, Department of Computer Science(比萨大学计算机科学系) National Research Council of Italy, Institute of Information Science and Technologies(意大利国家研究委员会信息科学与技术研究所) University of Trento, Department of Psychology and Cognitive Science(特伦托大学心理学与认知科学系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 24 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13499 2025-10-29 cs.CY cs.AI 62%

Reproducible workflow for online AI in digital health

Susobhan Ghosh, Bhanu T. Gullapalli, Daiqi Gao, Asim Gazi, Anna Trella, Ziping Xu, Kelly Zhang, Susan A. Murphy

机构 * Harvard University(哈佛大学) Imperial College London(帝国理工学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24115 2025-10-29 cs.AI cs.LG 62%

HistoLens: An Interactive XAI Toolkit for Verifying and Mitigating Flaws in Vision-Language Models for Histopathology

Sandeep Vissapragada, Vikrant Sahu, Gagan Raj Gupta, Vandita Singh

机构 * Indian Institute of Technology(印度理工学院) All India Institute of Medical Sciences(全印度医学科学研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23854 2025-10-29 cs.CL cs.AI 62%

Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs

Jyotika Singh, Weiyi Sun, Amit Agarwal, Viji Krishnamurthy, Yassine Benajiba, Sujith Ravi, Dan Roth

机构 * Oracle AI

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23471 2025-10-28 stat.ML cs.AI cs.LG 62%

Robust Decision Making with Partially Calibrated Forecasts

Shayan Kiyani, Hamed Hassani, George Pappas, Aaron Roth

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏