arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-14 至 2025-10-14 共收录 97 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 36 篇

2506.01339 2025-10-14 cs.LG 57%

Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning

Changsheng Wang, Yihua Zhang, Jinghan Jia, Parikshit Ram, Dennis Wei, Yuguang Yao, Soumyadeep Pal, Nathalie Baracaldo, Sijia Liu

机构 * Michigan State University(密歇根州立大学) IBM Research(IBM研究院)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11534 2025-10-14 cs.RO cs.SY eess.SY 50%

IntersectioNDE: Learning Complex Urban Traffic Dynamics based on Interaction Decoupling Strategy

Enli Lin, Ziyuan Yang, Qiujing Lu, Jianming Hu, Shuo Feng

机构 * Department of Mechanical Engineering, Tsinghua University(清华大学机械工程系) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 安全评测 :safety(abstract)

Comments Accepted by ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08789 2025-10-14 cs.CV 50%

Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization

Shuo Xing, Soumik Dey, Mingyang Wu, Ashirbad Mishra, Naveen Ravipati, Binbin Li, Hansi Wu, Zhengzhong Tu

机构 * Department of Computer Science(计算机科学系) Cranberry-Lemon University(Cranberry-Lemon 大学) Texas A&M University(德克萨斯A&M大学) eBay Inc.(eBay公司)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11414 2025-10-14 cs.CR 50%

Uncertainty-Aware, Risk-Adaptive Access Control for Agentic Systems using an LLM-Judged TBAC Model

Charles Fleming, Ashish Kundu, Ramana Kompella

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10037 2025-10-14 cs.CE 50%

Automated Glaucoma Report Generation via Dual-Attention Semantic Parallel-LSTM and Multimodal Clinical Data Integration

Cheng Huang, Weizheng Xie, Zeyu Han, Tsengdar Lee, Karanjit Kooner, Jui-Ka Wang, Ning Zhang, Jia Zhang

专题命中 安全评测 :alignment(abstract)

Comments Accepted by IEEE 25th BIBE

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 6 篇

2510.10998 2025-10-14 cs.CL cs.AI cs.CY cs.HC cs.LG 70%

ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios

Mahika Phutane, Hayoung Jung, Matthew Kim, Tanushree Mitra, Aditya Vashistha

机构 * Cornell University(康奈尔大学) Princeton University(普林斯顿大学) University of Washington(华盛顿大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 28 pages, 11 figures, 16 tables. In submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09871 2025-10-14 cs.CL 70%

CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs

Nafiseh Nikeghbal, Amir Hossein Kargaran, Jana Diesner

机构 * Technical University of Munich(慕尼黑技术大学) LMU Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09634 2025-10-14 cs.CY cs.AI 62%

Responsible AI Adoption in the Public Sector: A Data-Centric Taxonomy of AI Adoption Challenges

Anastasija Nikiforova, Martin Lnenicka, Ulf Melin, David Valle-Cruz, Asif Gill, Cesar Casiano Flores, Emyana Sirait, Mariusz Luterek, Richard Michael Dreyling, Barbora Tesarova

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11595 2025-10-14 cs.AI cs.GL 57%

Reproducibility: The New Frontier in AI Governance

Israel Mason-Williams, Gabryel Mason-Williams

机构 * Imperial College London(帝国理工学院) King's College London(国王学院) Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 12 pages,6 figures,Workshop on Technical AI Governance at ICML

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11079 2025-10-14 cs.AI 57%

Argumentation-Based Explainability for Legal AI: Comparative and Regulatory Perspectives

Andrada Iulia Prajescu, Roberto Confalonieri

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18837 2025-10-14 cs.SI 50%

Sentiment and Social Signals in the Climate Crisis: A Survey on Analyzing Social Media Responses to Extreme Weather Events

Pouya Shaeri, Yasaman Mohammadpour, Alimohammad Beigi, Ariane Middel

专题命中 AI治理与伦理 :alignment(abstract)

Comments Accepted and Published in SBP-BRiMS 2025. 18th International Conference on Social Computing, Behavioral-Cultural Modeling & Prediction and Behavior Representation in Modeling and Simulation

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 26 篇

2505.20333 2025-10-14 cs.CL cs.AI 81%

Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework

Yukun Zhang, Qi Dong

机构 * The Chinese University of Hong Kong(香港中文大学) Fudan University(复旦大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05129 2025-10-14 cs.CL cs.LG 81%

Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models

Qingshu Xu, Hong Jiao, Tianyi Zhou, Ming Li, Nan Zhang, Sydney Peters, Yanbin Fu

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Shandong Jiaotong University(山东交通大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03227 2025-10-14 cs.LG cs.AI 81%

Bridging Neural ODE and ResNet: A Formal Error Bound for Safety Verification

Abdelrahman Sayed Sayed, Pierre-Jean Meyer, Mohamed Ghazel

机构 * Univ Gustave Eiffel(法国埃菲尔大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 18 pages, 5 figures, Accepted for publication in the proceedings of the 8th International Symposium on AI Verification SAIV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10732 2025-10-14 cs.CY 79%

When Openness Fails: Lessons from System Safety for Assessing Openness in AI

Tamara Paris, Shalaleh Rismani

专题命中 其他安全 :safety(title,abstract);分类 cs.CY

Comments Accepted to Symposium on Model Accountability, Sustainability and Healthcare (SMASH) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10280 2025-10-14 cs.CL 79%

On the Entity-Level Alignment in Crosslingual Consistency

Yihong Liu, Mingyang Wang, François Yvon, Hinrich Schütze

机构 * Center for Information and Language Processing, LMU Munich(信息与语言处理中心,慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心(MCML)) Sorbonne Université, CNRS, ISIR, France(索邦大学,CNRS,ISIR,法国)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26431 2025-10-14 cs.CL 79%

Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests

Yanbin Fu, Hong Jiao, Tianyi Zhou, Nan Zhang, Ming Li, Qingshu Xu, Sydney Peters, Robert W. Lissitz

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments need updates

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11040 2025-10-14 cs.CL 70%

Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks

Wenya Xie, Qingying Xiao, Yu Zheng, Xidong Wang, Junying Chen, Ke Ji, Anningzhe Gao, Prayag Tiwari, Xiang Wan, Feng Jiang, Benyou Wang

机构 * Shenzhen Research Institute of Big Data(大数据研究 institute) National Health Data Institute(国家健康数据研究所) Halmstad University(哈马碧大学) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11693 2025-10-14 cs.CL cs.AI cs.CV 62%

Scaling Language-Centric Omnimodal Representation Learning

Chenghao Xiao, Hou Pong Chan, Hao Zhang, Weiwen Xu, Mahani Aljunied, Yu Rong

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23703 2025-10-14 cs.AI cs.CL 62%

Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability

Ruida Wang, Yuxin Li, Yi R. Fung, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10398 2025-10-14 cs.CL cs.AI 62%

STEAM: A Semantic-Level Knowledge Editing Framework for Large Language Models

Geunyeong Jeong, Juoh Sun, Seonghee Lee, Harksoo Kim

机构 * Konkuk University(韩国康康大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18361 2025-10-14 q-bio.NC cs.AI cs.LG cs.RO 62%

Task-Optimized Convolutional Recurrent Networks Align with Tactile Processing in the Rodent Brain

Trinity Chung, Yuchen Shen, Nathan C. L. Kong, Aran Nayebi

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Machine Learning Department, Carnegie Mellon University(卡内基梅隆大学机器学习系) Department of Psychology, University of Pennsylvania(宾夕法尼亚大学心理学系) Neuroscience Institute, Carnegie Mellon University(卡内基梅隆大学神经科学研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 8 figures, 7 tables, NeurIPS 2025 Camera Ready Version (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15607 2025-10-14 cs.CL cs.AI 62%

From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning

David Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi, Iryna Gurevych, Mrinmaya Sachan

机构 * Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系) ETH AI Center(ETH人工智能中心) Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, Technical University of Darmstadt and National Research Center for Applied Cybersecurity ATHENE, Germany(达姆施塔特技术大学计算机科学系、应用网络安全国家研究中心ATHENE)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Main as an oral presentation. David Dinucu-Jianu and Jakub Macina contributed equally. Code available: https://github.com/eth-lre/PedagogicalRL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09762 2025-10-14 cs.LG cs.AI 62%

PatentVision: A multimodal method for drafting patent applications

Ruo Yang, Sai Krishna Reddy Mudhiganti, Manali Sharma

机构 * Samsung Semiconductor, Inc.(三星半导体公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12963 2025-10-14 cs.AI cs.LG 62%

Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills

Changsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia, Dennis Wei, Parikshit Ram, Nathalie Baracaldo, Sijia Liu

机构 * Michigan State University(密歇根州立大学) IBM Research(IBM研究院)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11502 2025-10-14 cs.LG 57%

Learning to Make MISTAKEs: Modeling Incorrect Student Thinking And Key Errors

Alexis Ross, Jacob Andreas

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10397 2025-10-14 cs.CL 57%

AssoMem: Scalable Memory QA with Multi-Signal Associative Retrieval

Kai Zhang, Xinyuan Zhang, Ejaz Ahmed, Hongda Jiang, Caleb Kumar, Kai Sun, Zhaojiang Lin, Sanat Sharma, Shereen Oraby, Aaron Colak, Ahmed Aly, Anuj Kumar, Xiaozhong Liu, Xin Luna Dong

机构 * Meta Reality Labs(Meta现实实验室) Worcester Polytechnic Institute(沃斯特理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08075 2025-10-14 cs.AI 57%

Multi-Condition Conformal Selection

Qingyang Hao, Wenbo Liao, Bingyi Jing, Hongxin Wei

机构 * Southern University of Science and Technology(南方科技大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05444 2025-10-14 cs.CL 57%

PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs

Sana Kang, Myeongseok Gwon, Su Young Kwon, Jaewook Lee, Andrew Lan, Bhiksha Raj, Rita Singh

机构 * KAIST(韩国科学技术院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12450 2025-10-14 cs.CL 57%

Language Surgery in Multilingual Large Language Models

Joanito Agili Lopo, Muhammad Ravi Shulthan Habibi, Tack Hwa Wong, Muhammad Ilham Ghozali, Fajri Koto, Genta Indra Winata, Peerat Limkonchotiwat, Alham Fikri Aji, Samuel Cahyawijaya

机构 * SEACrowd Kreasof AI Universitas Indonesia(印度尼西亚大学) MBZUAI Capital One AI Singapore(AI新加坡) Cohere

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏