arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-07 至 2025-10-07 共收录 82 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 7 篇

2510.04357 2025-10-07 cs.LG q-fin.CP 57%

From News to Returns: A Granger-Causal Hypergraph Transformer on the Sphere

Anoushka Harit, Zhongtian Sun, Jongmin Yu

机构 * University of Cambridge(剑桥大学) University of Kent(肯特大学) Department of Computer Science, University of Cambridge(剑桥大学计算机科学系)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments 6th ACM International Conference on AI in Finance

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10246 2025-10-07 cs.LG 57%

Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions

Hazel Kim, Tom A. Lamb, Adel Bibi, Philip Torr, Yarin Gal

机构 * University of Oxford(牛津大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

Comments Accepted to EMNLP(main)2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20148 2025-10-07 cs.CV 50%

Smaller is Better: Enhancing Transparency in Vehicle AI Systems via Pruning

Sanish Suwal, Shaurya Garg, Dipkamal Bhusal, Michael Clifford, Nidhi Rastogi

专题命中 幻觉与事实性 :safety(abstract)

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 隐私与版权 2 篇

2510.04609 2025-10-07 cs.CY cs.AI 62%

Accountability Capture: How Record-Keeping to Support AI Transparency and Accountability (Re)shapes Algorithmic Oversight

Shreya Chappidi, Jennifer Cobbe, Chris Norval, Anjali Mazumder, Jatinder Singh

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI、cs.CY

Comments To appear at 8th AAAI/ACM Conference on AI, Ethics, and Society (AIES 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01344 2025-10-07 cs.HC cs.AI cs.CR 57%

Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent

Zhiping Zhang, Bingcan Guo, Tianshi Li

机构 * Northeastern University(东北大学) University of Washington(华盛顿大学)

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 安全评测 22 篇

2510.04320 2025-10-07 cs.CL cs.LG 84%

Read the Scene, Not the Script: Outcome-Aware Safety for LLMs

Rui Wu, Yihao Quan, Zeru Shi, Zhenting Wang, Yanshu Li, Ruixiang Tang

机构 * Rutgers University(新泽西罗格斯大学)

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03309 2025-10-07 cs.LG q-bio.BM 79%

Thin Bridges for Drug Text Alignment: Lightweight Contrastive Learning for Target Specific Drug Retrieval

Mallikarjuna Tupakula

机构 * Rochester Institute of Technology(罗切斯特技术研究所)

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04023 2025-10-07 cs.AI cs.CL 79%

LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions

Mizanur Rahman, Amran Bhuiyan, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Ridwan Mahbub, Ahmed Masry, Shafiq Joty, Enamul Hoque

机构 * York University(约克大学) Vector Institute for AI(人工智能矢量研究所) Dialpad Inc.(Dialpad公司) Nanyang Technological University(南洋理工大学) Salesforce AI Research(Salesforce人工智能研究)

专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments Survey paper; 45 data science agents; under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03399 2025-10-07 cs.AI cs.CL cs.CY cs.LG 77%

Know Thyself? On the Incapability and Implications of AI Self-Recognition

Xiaoyan Bai, Aryan Shrivastava, Ari Holtzman, Chenhao Tan

机构 * University of Chicago(芝加哥大学)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Our code is available, see https://github.com/ChicagoHAI/self-recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03815 2025-10-07 eess.SY cs.LG cs.SY eess.SP 74%

A Trustworthy Industrial Fault Diagnosis Architecture Integrating Probabilistic Models and Large Language Models

Yue wu

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments 1tables,6 figs,11pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14744 2025-10-07 cs.LG 74%

Beyond the Single-Best Model: Rashomon Partial Dependence Profile for Trustworthy Explanations in AutoML

Mustafa Cavus, Jan N. van Rijn, Przemysław Biecek

机构 * Department of Statistics, Eskisehir Technical University, Turkiye(埃斯基谢普大学统计系) Leiden Institute of Advanced Computer Science, Leiden University, the Netherlands(莱顿大学高级计算机科学研究所) Faculty of Mathematics and Information Science, Warsaw University of Technology, Poland(华沙理工大学数学与信息科学学院) Informatics and Mechanics, University of Warsaw, Faculty of Mathematics, Poland(华沙大学信息技术与力学系)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Accepted at 28th International Conference on Discovery Science 2025

Journal ref In: Džeroski, S., Levatić, J., Pio, G., Simidjievski, N. (eds) Discovery Science. DS 2025. Lecture Notes in Computer Science, vol 16090. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03604 2025-10-07 cs.LG cs.AI 73%

Deep Domain Adaptation for Turbofan Engine Remaining Useful Life Prediction: Methodologies, Evaluation and Future Trends

Yucheng Wang, Mohamed Ragab, Yubo Hou, Zhenghua Chen, Min Wu, Xiaoli Li

机构 * Institute for Infocomm Research, Agency for Science, Technology and Research, Singapore(新加坡资讯通信研究院,科技研究局) Propulsion and Space Research Center, Technology Innovation Institute, UAE(阿联酋技术创新研究所推进与航天研究中心) James Watt School of Engineering, University of Glasgow, UK(格拉斯哥大学詹姆斯·瓦特工程学院) ISTD Pillar at SUTD(新加坡科技设计大学ISTD支柱)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18710 2025-10-07 cs.AI 70%

Autonomous Data Agents: A New Opportunity for Smart Data

Yanjie Fu, Dongjie Wang, Wangyang Ying, Xinyuan Wang, Xiangliang Zhang, Huan Liu, Jian Pei

机构 * Arizona State University(亚利桑那州立大学) University of Kansas(堪萨斯大学) University of Notre Dame(圣母大学) Duke University(杜克大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03688 2025-10-07 cs.AI cs.CL cs.LG 67%

AgentBench: Evaluating LLMs as Agents

Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, Jie Tang

机构 * Tsinghua University(清华大学) The Ohio State University(俄亥俄州立大学) UC Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published in ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01238 2025-10-07 cs.CL cs.LG 62%

Silent Tokens, Loud Effects: Padding in LLMs

Rom Himelstein, Amit LeVi, Yonatan Belinkov, Avi Mendelson

机构 * Department of Data and Decision Science, Technion - Israel Institute of Technology(数据与决策科学系,技术离子理工学院) Department of Computer Science, Technion - Israel Institute of Technology(计算机科学系,技术离子理工学院)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

Comments Accepted to NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04648 2025-10-07 cs.CV cs.CY 57%

EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents

Buyuan Zhu, Shiyu Hu, Yiping Ma, Yuanming Zhang, Kang Hao Cheong

机构 * School of Physical and Mathematical Sciences, Nanyang Technological University(南洋理工大学物理与数学科学学院) Lab of Artificial Intelligence for Education, East China Normal University(华东师范大学教育人工智能实验室) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(哈尔滨工业大学机器人系统国家重点实验室) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

Comments Preprint, Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09630 2025-10-07 cs.CV cs.AI 57%

Brain Stroke Detection and Classification Using CT Imaging with Transformer Models and Explainable AI

Shomukh Qari, Maha A. Thafar

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 5 figures

Journal ref https://www.mdpi.com/2075-4418/15/19/2486

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04380 2025-10-07 cs.SE cs.AI cs.HC 57%

Reconsidering Requirements Engineering: Human-AI Collaboration in AI-Native Software Development

Mateen Ahmed Abbasi, Petri Ihantola, Tommi Mikkonen, Niko Mäkitalo

机构 * Faculty of Information Technology, University of Jyväskylä(信息技术学院,约赫斯库利亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted at SEAA 2025. Appearing in Springer LNCS 16081, pages 164-180

Journal ref In: SEAA 2025 proceedings, LNCS vol. 16081, Springer

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04201 2025-10-07 cs.CV cs.AI 57%

World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge

Moo Hyun Son, Jintaek Oh, Sun Bin Mun, Jaechul Roh, Sehyun Choi

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Georgia Institute of Technology(佐治亚理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) TwelveLabs

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04081 2025-10-07 cs.CL cs.PL 57%

Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning

Honglin Lin, Qizhi Pei, Xin Gao, Zhuoshi Pan, Yu Li, Juntao Li, Conghui He, Lijun Wu

机构 * OpenDataLab, Shanghai Artificial Intelligence Laboratory(OpenDataLab,上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Soochow University(苏州大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Accepted by NeurIPS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03719 2025-10-07 cs.CY cs.HC 57%

A Survey of LLM-Based Applications in Programming Education: Balancing Automation and Human Oversight

Griffin Pitts, Anurata Prabha Hridi, Arun-Balajiee Lekshmi-Narayanan

专题命中 安全评测 :alignment(abstract);分类 cs.CY

Comments 2025 EMNLP HCI+NLP Workshop Short Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14574 2025-10-07 cs.CV cs.AI 57%

Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark

Rashid Mushkani

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19568 2025-10-07 cs.CV cs.AI 57%

How Far are AI-generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach

Chirui Chang, Jiahui Liu, Zhengzhe Liu, Xiaoyang Lyu, Yi-Hua Huang, Xin Tao, Pengfei Wan, Di Zhang, Xiaojuan Qi

机构 * The University of Hong Kong(香港大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Lingnan University(岭大)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03909 2025-10-07 cs.CV 50%

Generating Human Motion Videos using a Cascaded Text-to-Video Framework

Hyelin Nam, Hyojun Go, Byeongjun Park, Byung-Hoon Kim, Hyungjin Chung

机构 * EverEx University of Michigan(密歇根大学) ETH Zurich(苏黎世联邦理工学院) Yonsei University(延世大学)

专题命中 安全评测 :alignment(abstract)

Comments 18 pages, 7 figures, Project Page:https://hyelinnam.github.io/Cameo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24311 2025-10-07 cs.CV 50%

Towards Foundation Models for Cryo-ET Subtomogram Analysis

Runmin Jiang, Wanyue Feng, Yuntian Yang, Shriya Pingulkar, Hong Wang, Xi Xiao, Xiaoyu Cao, Genpei Zhang, Xiao Wang, Xiaolong Wu, Tianyang Wang, Yang Liu, Xingjian Li, Min Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Harvard University(哈佛大学) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) Oak Ridge National Laboratory(橡树岭国家实验室) K. J. Somaiya College of Engineering(K.J. Somaiya 工程学院)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23054 2025-10-07 cs.CV 50%

Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis

Ruilang Wang, Shuotong Xu, Bowen Liu, Runlin Huang, Donglong Chen, Weifeng Su

机构 * Beijing Normal–Hong Kong Baptist University(北京师范大学-香港 Baptist大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10808 2025-10-07 cs.CR cs.SY eess.SP eess.SY 50%

Contrastive-KAN: A Semi-Supervised Intrusion Detection Framework for Cybersecurity with scarce Labeled Data

Mohammad Alikhani, Reza Kazemi

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. AI治理与伦理 9 篇

2510.04528 2025-10-07 cs.CR cs.AI 79%

Unified Threat Detection and Mitigation Framework (UTDMF): Combating Prompt Injection, Deception, and Bias in Enterprise-Scale Transformers

Santhosh KumarRavindran

机构 * Microsoft Corporation(微软公司)

专题命中 AI治理与伦理 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04073 2025-10-07 cs.AI 79%

Moral Anchor System: A Predictive Framework for AI Value Alignment and Drift Prevention

Santhosh Kumar Ravindran

机构 * Microsoft Corporation(微软公司)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

Comments 11 pages Includes simulations with over 4 million steps

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06303 2025-10-07 cs.CY cs.AI cs.CL cs.LG 70%

On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions

Dang Nguyen, Chenhao Tan

机构 * Department of Computer Science University of Chicago(计算机科学系芝加哥大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 21 pages, 15 figures, 14 tables. Accepted as a conference paper at COLM 2025. Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏