arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2510.08930 2025-10-13 cs.HC cs.AI 57%

Co-Authoring the Self: A Human-AI Interface for Interest Reflection in Recommenders

Ruixuan Sun, Junyuan Wang, Sanjali Roy, Joseph A. Konstan

机构 * Grouplens Research, University of Minnesota(Grouplens研究院、明尼苏达大学) Department of Computer Science and Engineering, University of Minnesota(计算机科学与工程系、明尼苏达大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08783 2025-10-13 cs.HC cs.AI 57%

MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces

Reuben A. Luera, Ryan Rossi, Franck Dernoncourt, Samyadeep Basu, Sungchul Kim, Subhojyoti Mukherjee, Puneet Mathur, Ruiyi Zhang, Jihyung Kil, Nedim Lipka, Seunghyun Yoon, Jiuxiang Gu, Zichao Wang, Cindy Xiong Bearfield, Branislav Kveton

机构 * University of California, Berkeley(加州大学伯克利分校) Adobe Research(Adobe研究) Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07881 2025-10-10 cs.CL 57%

CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching

Heyang Liu, Yuhao Wang, Ziyang Cheng, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07829 2025-10-10 cs.HC cs.AI 57%

The Rise of the Knowledge Sculptor: A New Archetype for Knowledge Work in the Age of Generative AI

Cathal Doyle

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 23 pages, 11 figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07748 2025-10-10 cs.AI 57%

Haibu Mathematical-Medical Intelligent Agent:Enhancing Large Language Model Reliability in Medical Tasks via Verifiable Reasoning Chains

Yilun Zhang, Dexing Kong

机构 * Zhejiang Qiushi Institute of Mathematical Medicine(浙江启思数学医学研究院) School of Mathematical Sciences, Zhejiang University(浙江大学数学科学学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07697 2025-10-10 cs.CR cs.AI 57%

Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs

Man Hu, Xinyi Wu, Zuofeng Suo, Jinbo Feng, Linghui Meng, Yanhao Jia, Anh Tuan Luu, Shuai Zhao

机构 * Beijing Electronic Science and Technology Institute, China(北京电子科技研究所) Nanyang Technological University, Singapore(南洋理工大学) Hainan University, China(海南大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07350 2025-10-10 cs.LG 57%

Out-of-Distribution Generalization in Climate-Aware Yield Prediction with Earth Observation Data

Aditya Chakravarty

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Journal ref ICCV 2025 Workshop on Sustainability with Earth observation and AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18709 2025-10-10 cs.CL 57%

Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation

Duy Le, Kent Ziti, Evan Girard-Sun, Bakr Bouhaya, Sean O'Brien, Vasu Sharma, Kevin Zhu

机构 * Algoverse AI Research(Algoverse AI研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Paper was accepted in to NeurIPS 2025 Workshop GenProCC

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14146 2025-10-09 cs.CL 57%

MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation

Xian Gao, Jiacheng Ruan, Zongyun Zhang, Jingsheng Gao, Ting Liu, Yuzhuo Fu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06411 2025-10-09 cs.CL 57%

Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?

R. Alexander Knipper, Indrani Dey, Souvika Sarkar, Hari Narayanan, Sadhana Puntambekar, Santu Karmaker

机构 * Department of EdPsych, University of Wisconsin-Madison(威斯康星大学麦迪逊分校教育心理学系) Department of CS, Wichita State University(威斯康星州立大学Wichita分校计算机科学系) Department of CSSE, Auburn University(阿伯茨罕大学计算机科学与工程系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25282 2025-10-09 cs.AI cs.HC cs.SE 57%

Toward Causal-Visual Programming: Enhancing Agentic Reasoning in Low-Code Environments

Jiexi Xu, Jiaqi Liu, Lanruo Wang, Su Liu

机构 * School of Information \& Computer Science University of California, Irvine Irvine, CA, USA Independent Researcher University of Texas at Dallas Dallas, TX, USA Georgia Institute of Technology Atlanta, GA, USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 5 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19333 2025-10-09 cs.IR cs.CL 57%

Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation

Minghao Tang, Shiyu Ni, Jiafeng Guo, Keping Bi

机构 * State Key Laboratory of AI Safety(人工智能安全国家重点实验室) Institute of Computing Technology(计算技术研究所) Chinese Academy of Sciences(中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments SIGIR-AP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19017 2025-10-09 cs.CL 57%

Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models

Bin Zhu, Yinxuan Gui, Huiyan Qi, Jingjing Chen, Chong-Wah Ngo, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学) Fudan University(复旦大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Project website: https://yxg1005.github.io/GaslightingNegationAttacks/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03413 2025-10-08 cs.CE cs.AI 57%

Report of the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

Lois Curfman McInnes, Dorian Arnold, Prasanna Balaprakash, Mike Bernhardt, Beth Cerny, Anshu Dubey, Roscoe Giles, Denice Ward Hood, Mary Ann Leung, Vanessa Lopez-Marrero, Paul Messina, Olivia B. Newton, Chris Oehmen, Stefan M. Wild, Jim Willenbring, Lou Woodley, Tony Baylis, David E. Bernholdt, Chris Camano, Johannah Cohoon, Charles Ferenbaugh, Stephen M. Fiore, Sandra Gesing, Diego Gomez-Zara, James Howison, Tanzima Islam, David Kepczynski, Charles Lively, Harshitha Menon, Bronson Messer, Marieme Ngom, Umesh Paliath, Michael E. Papka, Irene Qualters, Elaine M. Raybourn, Katherine Riley, Paulina Rodriguez, Damian Rouson, Michelle Schwalbe, Sudip K. Seal, Ozge Surer, Valerie Taylor, Lingfei Wu

机构 * Argonne National Laboratory(阿贡国家实验室) Emory University(埃默里大学) Oak Ridge National Laboratory(橡树岭国家实验室) Team Libra(团队Libra) Boston University(波士顿大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Sustainable Horizons Institute(可持续远景研究所) Stony Brook University(石溪大学) University of Montana(蒙大拿大学) Pacific Northwest National Laboratory(太平洋西北国家实验室) Lawrence Berkeley National Laboratory(伯克利国家实验室) Sandia National Laboratories(桑塔那国家实验室) Center for Scientific Collaboration and Community Engagement(科学协作与社区参与中心) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) Californi(加利福尼亚)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 38 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05519 2025-10-08 cs.CY 57%

Assessing Human Rights Risks in AI: A Framework for Model Evaluation

Vyoma Raman, Camille Chabot, Betsy Popken

专题命中 安全评测 :safety(abstract);分类 cs.CY

Comments AAAI/ACM Conference on AI, Ethics, and Society (AIES) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05185 2025-10-08 cs.MA cs.CE cs.CY cs.NE cs.SI 57%

AgentZero++: Modeling Fear-Based Behavior

Vrinda Malhotra, Jiaman Li, Nandini Pisupati

专题命中 安全评测 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05142 2025-10-08 cs.CL cond-mat.mtrl-sci 57%

Reliable End-to-End Material Information Extraction from the Literature with Source-Tracked Multi-Stage Large Language Models

Xin Wang, Anshu Raj, Matthew Luebbe, Haiming Wen, Shuozhi Xu, Kun Lu

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 27 pages, 4 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04093 2025-10-08 cs.AI 57%

Harnessing LLM for Noise-Robust Cognitive Diagnosis in Web-Based Intelligent Education Systems

Guixian Zhang, Guan Yuan, Ziqi Xu, Yanmei Zhang, Jing Ren, Zhenyun Deng, Debo Cheng

机构 * School of Computer Science and Technology/School of Artificial Intelligence, China University of Mining and Technology(计算机科学与技术学院/人工智能学院,中国矿业大学) School of Computing Technologies, RMIT University(计算技术学院,拉筹伯大学) Department of Computer Science and Technology, University of Cambridge(计算机科学与技术系,剑桥大学) School of Computer Science and Technology, Hainan University(计算机科学与技术学院,海南大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04648 2025-10-07 cs.CV cs.CY 57%

EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents

Buyuan Zhu, Shiyu Hu, Yiping Ma, Yuanming Zhang, Kang Hao Cheong

机构 * School of Physical and Mathematical Sciences, Nanyang Technological University(南洋理工大学物理与数学科学学院) Lab of Artificial Intelligence for Education, East China Normal University(华东师范大学教育人工智能实验室) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(哈尔滨工业大学机器人系统国家重点实验室) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

Comments Preprint, Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09630 2025-10-07 cs.CV cs.AI 57%

Brain Stroke Detection and Classification Using CT Imaging with Transformer Models and Explainable AI

Shomukh Qari, Maha A. Thafar

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 5 figures

Journal ref https://www.mdpi.com/2075-4418/15/19/2486

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04380 2025-10-07 cs.SE cs.AI cs.HC 57%

Reconsidering Requirements Engineering: Human-AI Collaboration in AI-Native Software Development

Mateen Ahmed Abbasi, Petri Ihantola, Tommi Mikkonen, Niko Mäkitalo

机构 * Faculty of Information Technology, University of Jyväskylä(信息技术学院,约赫斯库利亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted at SEAA 2025. Appearing in Springer LNCS 16081, pages 164-180

Journal ref In: SEAA 2025 proceedings, LNCS vol. 16081, Springer

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04201 2025-10-07 cs.CV cs.AI 57%

World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge

Moo Hyun Son, Jintaek Oh, Sun Bin Mun, Jaechul Roh, Sehyun Choi

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Georgia Institute of Technology(佐治亚理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) TwelveLabs

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04081 2025-10-07 cs.CL cs.PL 57%

Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning

Honglin Lin, Qizhi Pei, Xin Gao, Zhuoshi Pan, Yu Li, Juntao Li, Conghui He, Lijun Wu

机构 * OpenDataLab, Shanghai Artificial Intelligence Laboratory(OpenDataLab,上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Soochow University(苏州大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Accepted by NeurIPS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03719 2025-10-07 cs.CY cs.HC 57%

A Survey of LLM-Based Applications in Programming Education: Balancing Automation and Human Oversight

Griffin Pitts, Anurata Prabha Hridi, Arun-Balajiee Lekshmi-Narayanan

专题命中 安全评测 :alignment(abstract);分类 cs.CY

Comments 2025 EMNLP HCI+NLP Workshop Short Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14574 2025-10-07 cs.CV cs.AI 57%

Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark

Rashid Mushkani

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19568 2025-10-07 cs.CV cs.AI 57%

How Far are AI-generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach

Chirui Chang, Jiahui Liu, Zhengzhe Liu, Xiaoyang Lyu, Yi-Hua Huang, Xin Tao, Pengfei Wan, Di Zhang, Xiaojuan Qi

机构 * The University of Hong Kong(香港大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Lingnan University(岭大)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16076 2025-10-06 cs.CL 57%

The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models

Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, Markus Strohmaier

机构 * University of Mannheim(曼海姆大学) GESIS - Leibniz Institute for the Social Sciences(莱布尼茨社会科学研究所) Complexity Science Hub Vienna(维也纳复杂科学中心)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to EMNLP Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02075 2025-10-06 cs.RO cs.LG cs.SY eess.SY 57%

Active Alignments of Lens Systems with Reinforcement Learning

Matthias Burkhardt, Tobias Schmähling, Pascal Stegmann, Michael Layh, Tobias Windisch

机构 * Institute for Machine Vision, University of Applied Sciences Kempten(机器视觉研究所,应用科技大学凯普腾)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21185 2025-10-06 cs.LG 57%

Amelia: A Large Dataset and Benchmark for Airport Surface Movement Forecasting

Ingrid Navarro, Pablo Ortega-Kral, Jay Patrikar, Haichuan Wang, Alonso Cano, Zelin Ye, Jong Hoon Park, Sebastian Scherer, Jean Oh

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 40 pages, 19 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01671 2025-10-03 cs.AI cs.HC 57%

A Locally Executable AI System for Improving Preoperative Patient Communication: A Multi-Domain Clinical Evaluation

Motoki Sato, Yuki Matsushita, Hidekazu Takahashi, Tomoaki Kakazu, Sou Nagata, Mizuho Ohnuma, Atsushi Yoshikawa, Masayuki Yamamura

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 32 pages, 4 figures, 10 tables 32 pages, 4 figures, 10 tables. This paper is currently under review at ACM Transactions on Computing for Healthcare. Reproducibility resources: http://github.com/motokinaru/LENOHA-medical-dialogue

详情

展开后加载摘要…

URL PDF HTML 收藏