arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-16 至 2025-09-16 共收录 61 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 17 篇

2509.11376 2025-09-16 cs.LG cs.AI cs.CE 62%

Intelligent Reservoir Decision Support: An Integrated Framework Combining Large Language Models, Advanced Prompt Engineering, and Multimodal Data Fusion for Real-Time Petroleum Operations

Seyed Kourosh Mahjour, Seyed Saman Mahjour

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10802 2025-09-16 q-fin.RM cs.CL cs.LG q-fin.CP 62%

Why Bonds Fail Differently? Explainable Multimodal Learning for Multi-Class Default Prediction

Yi Lu, Aifan Ling, Chaoqun Wang, Yaxin Xu

机构 * School of Economics and Finance, Shanghai International Studies University(经济金融学院,上海国际问题研究大学) School of AI and Advanced Computing, Xi’an Jiaotong-Liverpool University(人工智能与先进计算学院,西安交通大学利物浦大学) School of Foreign Studies, Shanghai University of Finance and Economics(外国语言学院,上海金融学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12137 2025-09-16 eess.SY cs.AI cs.SY 57%

Control Analysis and Design for Autonomous Vehicles Subject to Imperfect AI-Based Perception

Tao Yan, Zheyu Zhang, Jingjing Jiang, Wen-Hua Chen

机构 * Department of Aeronautical and Automotive Engineering, Loughborough University(航空与汽车工程系,洛辛厄姆大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08638 2025-09-16 eess.AS cs.AI cs.MM cs.SD 57%

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Ruibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang, Jiahao Pan, Yongyi Zang, Haohe Liu, Yiming Liang, Wenye Ma, Xingjian Du, Xinrun Du, Zhen Ye, Tianyu Zheng, Zhengxuan Jiang, Yinghao Ma, Minghao Liu, Zeyue Tian, Ziya Zhou, Liumeng Xue, Xingwei Qu, Yizhi Li, Shangda Wu, Tianhao Shen, Ziyang Ma, Jun Zhan, Chunhui Wang, Yatian Wang, Xiaowei Chi, Xinyue Zhang, Zhenzhu Yang, Xiangzhou Wang, Shansong Liu, Lingrui Mei, Peng Li, Junjie Wang, Jianwei Yu, Guojian Pang, Xu Li, Zihao Wang, Xiaohuan Zhou, Lijun Yu, Emmanouil Benetos, Yong Chen, Chenghua Lin, Xie Chen, Gus Xia, Zhaoxiang Zhang, Chao Zhang, Wenhu Chen, Xinyu Zhou, Xipeng Qiu, Roger Dannenberg, Jiaheng Liu, Jian Yang, Wenhao Huang, Wei Xue, Xu Tan, Yike Guo

机构 * HKUST(香港科技大学) MAP(多模态艺术投影)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments https://github.com/multimodal-art-projection/YuE

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11802 2025-09-16 cs.CL 57%

When Curiosity Signals Danger: Predicting Health Crises Through Online Medication Inquiries

Dvora Goncharok, Arbel Shifman, Alexander Apartsin, Yehudit Aperstein

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10937 2025-09-16 cs.CL 57%

An Interpretable Benchmark for Clickbait Detection and Tactic Attribution

Lihi Nofar, Tomer Portal, Aviv Elbaz, Alexander Apartsin, Yehudit Aperstein

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10843 2025-09-16 cs.CL 57%

Evaluating Large Language Models for Evidence-Based Clinical Question Answering

Can Wang, Yiqun Chen

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06336 2025-09-16 cs.CV cs.AI cs.CR 57%

Multi-View Slot Attention Using Paraphrased Texts for Face Anti-Spoofing

Jeongmin Yu, Susang Kim, Kisu Lee, Taekyoung Kwon, Won-Yong Shin, Ha Young Kim

机构 * Yonsei University(延世大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06646 2025-09-16 econ.GN cs.CL q-fin.EC 57%

Evaluating and Aligning Human Economic Risk Preferences in LLMs

Jiaxin Liu, Yixuan Tang, Yi Yang, Kar Yan Tam

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14633 2025-09-16 q-bio.NC cs.AI cs.CV 57%

Evaluating Representational Similarity Measures from the Lens of Functional Correspondence

Yiqing Bo, Ansh Soni, Sudhanshu Srivastava, Meenakshi Khosla

机构 * University of California, San Diego(加州大学圣地亚哥分校) University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Published in CCN 2025 Proceedings (Talk & Poster), May 14, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12076 2025-09-16 cs.IR 50%

AEFS: Adaptive Early Feature Selection for Deep Recommender Systems

Fan Hu, Gaofeng Lu, Jun Chen, Chaonan Guo, Yuekui Yang, Xirong Li

专题命中 安全评测 :alignment(abstract)

Comments Accepted by TKDE

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11885 2025-09-16 cs.CV 50%

BREA-Depth: Bronchoscopy Realistic Airway-geometric Depth Estimation

Francis Xiatian Zhang, Emile Mackute, Mohammadreza Kasaei, Kevin Dhaliwal, Robert Thomson, Mohsen Khadem

机构 * Baillie Gifford Pandemic Science Hub(巴利尔·吉福德疫情科学中心) Institute of Regeneration and Repair(再生与修复研究所) Heriot-Watt University(赫罗特-瓦特大学) Institute of Photonics and Quantum Science(光子与量子科学研究所)

专题命中 安全评测 :safety(abstract)

Comments The paper has been accepted to MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19260 2025-09-16 cs.CR 50%

ALRPHFS: Adversarially Learned Risk Patterns with Hierarchical Fast \& Slow Reasoning for Robust Agent Defense

Shiyu Xiang, Tong Zhang, Ronghao Chen

专题命中 安全评测 :safety(abstract)

Comments EMNLP 2025 findings, 20 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20065 2025-09-16 physics.app-ph physics.optics 50%

Fast approximate solvers for metamaterials design in electromagnetism

Raphael Pestourie

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 5 篇

2509.11620 2025-09-16 cs.CL cs.CY 81%

AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment

Kun Li, Lai-Man Po, Hongzheng Yang, Xuyuan Xu, Kangcheng Liu, Yuzhi Zhao

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11648 2025-09-16 cs.CL cs.AI cs.CY 67%

EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI

Sai Kartheek Reddy Kasu

机构 * IIIT Dharwad(德瓦德理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17044 2025-09-16 cs.CY cs.AI cs.LG 67%

Approaches to Responsible Governance of GenAI in Organizations

Dhari Gandhi, Himanshu Joshi, Lucas Hartman, Shabnam Hassani

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) Western University(西部大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10653 2025-09-16 cs.CY cs.AI 62%

SCOR: A Framework for Responsible AI Innovation in Digital Ecosystems

Mohammad Saleh Torkestani, Taha Mansouri

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Proceeding of The British Academy of Management Conference 2025, University of Kent, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12104 2025-09-16 cs.AI 57%

JustEva: A Toolkit to Evaluate LLM Fairness in Legal Knowledge Inference

Zongyue Xue, Siyuan Zheng, Shaochun Wang, Yiran Hu, Shenran Wang, Yuxin Yao, Haitao Li, Qingyao Ai, Yiqun Liu, Yun Liu, Weixing Shen

机构 * Tsinghua University(清华大学) Yale Law School(耶鲁法学院) Shanghai Jiaotong University(上海交通大学) University of Waterloo(滑铁卢大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments This paper has been accepted at CIKM 2025 (Demo Track)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 12 篇

2410.15633 2025-09-16 cs.CL cs.AI 81%

GATEAU: Selecting Influential Samples for Long Context Alignment

Shuzheng Si, Haozhe Zhao, Gang Chen, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Peking University(北京大学) DeepLang AI Institute for AI, Tsinghua University(清华大学人工智能研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09931 2025-09-16 cs.LG cs.AI 81%

Mechanistic Interpretability of LoRA-Adapted Language Models for Nuclear Reactor Safety Applications

Yoon Pyo Lee

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in Nuclear Technology. 24 pages, 2 tables, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10509 2025-09-16 cs.LG cs.AI 73%

The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback

Sai Teja Reddy Adapala

机构 * University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments 5 pages, 3 figures, 2 tables. Code is available at: https://github.com/imsaitejareddy/ouroboros-effect-experiment

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10529 2025-09-16 cs.LG cs.AI cs.CV 62%

Mitigating Catastrophic Forgetting and Mode Collapse in Text-to-Image Diffusion via Latent Replay

Aoi Otani

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10523 2025-09-16 cs.LG cs.AI 62%

From Predictions to Explanations: Explainable AI for Autism Diagnosis and Identification of Critical Brain Regions

Kush Gupta, Amir Aly, Emmanuel Ifeachor, Rohit Shankar

机构 * University of Plymouth, Plymouth, UK(普利茅斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01042 2025-09-16 cs.LG 57%

SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals

Peixuan Han, Cheng Qian, Xiusi Chen, Yuji Zhang, Heng Ji, Denghui Zhang

机构 * Siebel School of Computing(计算科学系) Data Science, University of Illinois Urbana-Champaign(数据科学,伊利诺伊大学厄巴纳-香槟分校) School of Business, Stevens Institute of Technology(商学院,史蒂文斯理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11661 2025-09-16 cs.CV cs.AI 57%

DTGen: Generative Diffusion-Based Few-Shot Data Augmentation for Fine-Grained Dirty Tableware Recognition

Lifei Hao, Yue Cheng, Baoqi Huang, Bing Jia, Xuandong Zhao

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00521 2025-09-16 cs.SE cs.AI 57%

Automated detection of atomicity violations in large-scale systems

Hang He, Yixing Luo, Chengcheng Wan, Ting Su, Haiying Sun, Geguang Pu

机构 * East China Normal University(华东师范大学) Beijing Institute of Control Engineering(北京控制工程研究所) East China Normal University Shanghai Innovation Institute(华东师范大学上海创新研究院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11840 2025-09-16 cs.CV 50%

Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation

Tim Lebailly, Vijay Veerabadran, Satwik Kottur, Karl Ridgeway, Michael Louis Iuzzolino

机构 * Meta KU Leuven(鲁汶大学)

专题命中 其他安全 :alignment(abstract)

Comments ICCV 2025 CDEL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12356 2025-09-16 eess.IV cs.CV 50%

Regist3R: Incremental Registration with Stereo Foundation Model

Sidun Liu, Wenyu Li, Peng Qiao, Yong Dou

机构 * College of Computer Science and Technology(计算机科学与技术学院) National Key Laboratory of Parallel and Distributed Computing(并行与分布式计算国家重点实验室) National University of Defense Technology(国防科技大学)

专题命中 其他安全 :alignment(abstract)

Comments Accepted by ACM Multimedia 2025. github link: https://github.com/Liu-SD/Regist3R

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11783 2025-09-16 cs.RO 50%

Augmented Reality-Enhanced Robot Teleoperation for Collecting User Demonstrations

Shiqi Gong, Sebastian Zudaire, Chi Zhang, Zhen Li

机构 * Aalto University(阿alto大学) ABB Corporate Research(ABB企业研究)

专题命中 其他安全 :safety(abstract)

Comments Accepted by 2025 8th International Conference on Robotics, Control and Automation Engineering (RCAE 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏