arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-22 至 2025-10-22 共收录 15 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 15 篇

2503.04150 2025-10-22 cs.CL cs.AI 81%

Temporal Alignment of LLMs through Cycle Encoding for Long-Range Time Representations

Xue Han, Qian Hu, Yitong Wang, Wenchun Gao, Lianlian Zhang, Qing Wang, Lijun Mei, Chao Deng, Junlan Feng

机构 * JIUTIAN Team China Mobile Research Institute(中国移动研究院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12494 2025-10-22 cs.CL 79%

BIRD: A Trustworthy Bayesian Inference Framework for Large Language Models

Yu Feng, Ben Zhou, Weidong Lin, Dan Roth

机构 * University of Pennsylvania(宾夕法尼亚大学) Arizona State University(亚利桑那州立大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

Journal ref ICLR 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18550 2025-10-22 cs.NI 78%

JAUNT: Joint Alignment of User Intent and Network State for QoE-centric LLM Tool Routing

Enhan Li, Hongyang Du

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17575 2025-10-22 cs.HC 67%

DeTAILS: Deep Thematic Analysis with Iterative LLM Support

Ansh Sharma, Karen Cochrane, James R. Wallace

专题命中 安全评测 :alignment(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17910 2025-10-22 cs.CY cs.AI cs.CL 67%

Interpretability Framework for LLMs in Undergraduate Calculus

Sagnik Dakshit, Sushmita Sinha Roy

机构 * University of Texas at Tyler(德克萨斯理工大学) Florida Gulf Coast University(佛罗里达盖恩斯维尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18581 2025-10-22 cs.CY cs.AI 62%

The Cost-Benefit of Interdisciplinarity in AI for Mental Health

Katerina Drakos, Eva Paraschou, Simay Toplu, Line Harder Clemmensen, Christoph Lütge, Nicole Nadine Lønfeldt, Sneha Das

机构 * Center for Social Data Science, Faculty of Social Sciences, University of Copenhagen(哥本哈根大学社会科学学院社会数据科学中心) Dept. of Applied Mathematics and Computer Science, Technical University of Denmark(丹麦技术大学应用数学与计算机科学系) Institute for Ethics in Artificial Intelligence, School of Social Sciences and Technology, Technical University of Munich(慕尼黑技术大学社会科学与技术学院人工智能伦理研究所) Dept. of Mathematical Sciences, University of Copenhagen(哥本哈根大学数学科学系) Child and Adolescent Mental Health Center, Copenhagen University Hospital(哥本哈根大学医院青少年与儿童心理健康中心)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

Comments Accepted for poster presentation at the AI in Science Summit 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18898 2025-10-22 cs.CV cs.AI cs.LG cs.RO 62%

Interpretable Decision-Making for End-to-End Autonomous Driving

Mona Mirzaie, Bodo Rosenhahn

机构 * Institute for Information Processing, Leibniz University Hannover(信息处理研究所,汉诺威莱布尼茨大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted to the ICCV 2025 2nd Workshop on the Challenge Of Out-of-Label Hazards in Autonomous Driving (2COOOL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18425 2025-10-22 cs.AI 57%

Automated urban waterlogging assessment and early warning through a mixture of foundation models

Chenxu Zhang, Fuxiang Huang, Lei Zhang

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Submitted to Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18417 2025-10-22 cs.NI cs.AI 57%

On AI Verification in Open RAN

Rahul Soundrarajan, Claudio Fiandrino, Michele Polese, Salvatore D'Oro, Leonardo Bonati, Tommaso Melodia

机构 * Tejas Networks IMDEA Networks Institute Northeastern University

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06239 2025-10-22 cs.AI 57%

Proof2Silicon: Prompt Repair for Verified Code and Hardware Generation via Reinforcement Learning

Manvi Jha, Jiaxin Wan, Deming Chen

机构 * Electrical and Computer Engineering(电气与计算机工程系) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11690 2025-10-22 cs.LG cs.CV 57%

The Impact of Coreset Selection on Spurious Correlations and Group Robustness

Amaya Dharmasiri, William Yang, Polina Kirichenko, Lydia Liu, Olga Russakovsky

机构 * Princeton University(普林斯顿大学) FAIR at Meta(Meta 的 FAIR 实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 10 pages, 9 additional pages for Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18563 2025-10-22 cs.CR 50%

The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability

Zijie Xu, Minfeng Qi, Shiqing Wu, Lefeng Zhang, Qiwen Wei, Han He, Ningran Li

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18364 2025-10-22 cs.IR cs.SE 50%

Evaluating LLM-Based Mobile App Recommendations: An Empirical Study

Quim Motger, Xavier Franch, Vincenzo Gervasi, Jordi Marco

专题命中 安全评测 :alignment(abstract)

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18169 2025-10-22 eess.AS cs.SD 50%

Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction

Yu-Wen Chen, William Ho, Sasha M. Vergez, Grace Flaherty, Pallavi Gupta, Zhihong Zhang, Maryam Zolnoori, Margaret V. McDonald, Maxim Topaz, Zoran Kostic, Julia Hirschberg

机构 * The Fu Foundation School of Engineering and Applied Science, Columbia University(哥伦比亚大学福基金会工程与应用科学学院) School of Nursing, Columbia University(哥伦比亚大学护理学院) Center for Home Care Policy & Research, VNS Health(VNS健康居家护理政策与研究中心)

专题命中 安全评测 :alignment(abstract)

Comments The Second Workshop on GenAI for Health at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02236 2025-10-22 cs.CV cs.MM cs.SD eess.AS 50%

3D Audio-Visual Segmentation

Artem Sokolov, Swapnil Bhosale, Xiatian Zhu

机构 * University of Surrey, UK(Surrey大学)

专题命中 安全评测 :alignment(abstract)

Comments Accepted at the NeurIPS 2024 Workshop on Audio Imagination; this version updates the project page link

详情

展开后加载摘要…

URL PDF HTML 收藏