arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-12 至 2025-11-12 共收录 55 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2511.07794 2025-11-12 cs.CL 57%

Design, Results and Industry Implications of the World's First Insurance Large Language Model Evaluation Benchmark

Hua Zhou, Bing Ma, Yufei Zhang, Yi Zhao

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments 16 pages, 11 models,1 set of evaluation framework,5 core dimensions, 54 sub-indicators, 14,430 high-quality questions

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05529 2025-11-12 q-bio.QM cs.AI cs.CV 57%

Selective Diabetic Retinopathy Screening with Accuracy-Weighted Deep Ensembles and Entropy-Guided Abstention

Jophy Lin

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05459 2025-11-12 cs.SE cs.AI 57%

SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models

Jingxuan Xu, Ken Deng, Weihao Li, Songwei Yu, Huaixi Tang, Haoyang Huang, Zhiyi Lai, Zizheng Zhan, Yanan Wu, Chenchen Zhang, Kepeng Lei, Yifan Yao, Xinping Lei, Wenqiang Zhu, Zongxian Feng, Han Li, Junqi Xiong, Dailin Li, Zuchen Gao, Kun Wu, Wen Xiang, Ziqi Zhan, Yuanxing Zhang, Wuxuan Gong, Ziyuan Gao, Guanxiang Wang, Yirong Xue, Mengtong Li, Mengfei Xie, Xiaojiang Zhang, Jinghui Wang, Wenhao Zhuang, Zheng Lin, Huiming Wang, Zhaoxiang Zhang, Yuqun Zhang, Haotian Zhang, Bin Chen, Jiaheng Liu

机构 * Kuaishou Technology(快手科技) Nanjing University(南京大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24378 2025-11-12 cs.LG 57%

AXIS: Explainable Time Series Anomaly Detection with Large Language Models

Tian Lan, Hao Duong Le, Jinbo Li, Wenjun He, Meng Wang, Chenghao Liu, Chen Zhang

机构 * Huawei(华为)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03764 2025-11-12 cs.IR cs.LG 57%

LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest

Han Wang, Alex Whitworth, Pak Ming Cheung, Zhenjie Zhang, Krishna Kamath

机构 * Pinterest(Pininterest)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments RecSys 2025 EARL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01849 2025-11-12 cs.HC cs.AI 57%

A Multi-Agent Conversational Bandit Approach to Online Evaluation and Selection of User-Aligned LLM Responses

Xiangxiang Dai, Yuejin Xie, Maoli Liu, Xuchuang Wang, Zhuohua Li, Huanyu Wang, John C. S. Lui

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06771 2025-11-12 cs.RO cs.LG cs.MA 57%

JaxRobotarium: Training and Deploying Multi-Robot Policies in 10 Minutes

Shalin Anand Jain, Jiazhen Liu, Siva Kailas, Harish Ravichandar

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 22 pages, 14 figures, 10 tables. https://github.com/GT-STAR-Lab/JaxRobotarium. Manuscript accepted for publication at the 9th Conference on Robot Learning (CoRL 2025), Seoul, Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07744 2025-11-12 cs.CV 50%

VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics

Daniel Cher, Brian Wei, Srikumar Sastry, Nathan Jacobs

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 6 篇

2511.08567 2025-11-12 cs.LG cs.AI 62%

The Path Not Taken: RLVR Provably Learns Off the Principals

Hanqing Zhu, Zhenyu Zhang, Hanxian Huang, DiJia Su, Zechun Liu, Jiawei Zhao, Igor Fedorov, Hamed Pirsiavash, Zhizhou Sha, Jinwon Lee, David Z. Pan, Zhangyang Wang, Yuandong Tian, Kai Sheng Tai

机构 * Meta AI The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments Preliminary version accepted as a spotlight in NeurIPS 2025 Workshop on Efficient Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08082 2025-11-12 cs.AI cs.LG econ.GN q-fin.EC 62%

Prudential Reliability of Large Language Models in Reinsurance: Governance, Assurance, and Capital Efficiency

Stella C. Dong

机构 * Reinsurance Analytics(再保险分析)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments 48 pages, 9 figures, 5 tables. Submitted to the Journal of Risk and Insurance (JRI), November 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07803 2025-11-12 cs.CY cs.AI 62%

Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring

Wenhao Xu, Akshatha Arodi, Jian-Yun Nie, Arsene Fansi Tchango

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments To appear at AAAI-26 (Social Impact Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07941 2025-11-12 cs.CV cs.AI 57%

Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image Classification

Zhenfeng Zhuang, Fangyu Zhou, Liansheng Wang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02241 2025-11-12 cs.CL 57%

Isolating Culture Neurons in Multilingual Large Language Models

Danial Namazifard, Lukas Galke Poech

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted at IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14686 2025-11-12 cs.CV 50%

From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition

Chen Cai, Tianyi Liu, Jianjun Gao, Wenyang Liu, Kejun Wu, Ruoyu Wang, Yi Wang, Soo Chin Liew

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Huazhong University of Science and Technology(华中科技大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 AI治理与伦理 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 11 篇

2511.08080 2025-11-12 cs.LG cs.AI 81%

Hierarchical Structure-Property Alignment for Data-Efficient Molecular Generation and Editing

Ziyu Fan, Zhijian Huang, Yahan Li, Xiaowen Hu, Siyuan Shen, Yunliang Wang, Zeyu Zhong, Shuhong Liu, Shuning Yang, Shangqian Wu, Min Wu, Lei Deng

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Central South University(中南大学) Institute for Infocomm Research Agency for Science, Technology and Research (A* STAR)(信息通信研究所(A* STAR))

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08402 2025-11-12 cs.CV cs.AI cs.LG 62%

Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation

Difei Gu, Yunhe Gao, Mu Zhou, Dimitris Metaxas

机构 * Rutgers University(新泽西罗格斯大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08052 2025-11-12 cs.AI cs.CL cs.SE 62%

Dual-Process Scaffold Reasoning for Enhancing LLM Code Debugging

Po-Chung Hsieh, Chin-Po Chen, Jeng-Lin Li, Ming-Ching Chang

机构 * National Taiwan University(国立台湾大学) AI Research Center, Inventec Corporation(Inventec公司人工智能研究中心) University at Albany, SUNY(纽约州立大学阿尔巴尼分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07699 2025-11-12 econ.GN cs.LG q-fin.EC 57%

Misaligned by Design: Incentive Failures in Machine Learning

David Autor, Andrew Caplin, Daniel Martin, Philip Marx

机构 * Massachusetts Institute of Technology, Google Technology and Society Fellows program, and NBER(麻省理工学院、谷歌技术与社会 fellows 程序及国家经济研究局) New York University and NBER(纽约大学及国家经济研究局) University of California, Santa Barbara(加州大学圣塔芭芭拉分校) Louisiana State University(路易斯安那州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17510 2025-11-12 cs.CL 57%

Large Language Models Do Multi-Label Classification Differently

Marcus Ma, Georgios Chochlakis, Niyantha Maruthu Pandiyan, Jesse Thomason, Shrikanth Narayanan

机构 * University of Southern California(南加州大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments To be published in the Main Conference Proceedings of EMNLP 2025, 24 pages, 16 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15056 2025-11-12 cs.LG cs.CV cs.GR 57%

ElastoGen: 4D Generative Elastodynamics

Yutao Feng, Yintong Shang, Xiang Feng, Lei Lan, Shandian Zhe, Tianjia Shao, Hongzhi Wu, Kun Zhou, Chenfanfu Jiang, Yin Yang

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17120 2025-11-12 cs.CL 57%

Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions

Dillon Plunkett, Adam Morris, Keerthi Reddy, Jorge Morales

机构 * Northeastern University(东北大学) Princeton University(普林斯顿大学) Independent Researcher(独立研究者)

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09205 2025-11-12 cs.MM cs.CL cs.IR cs.SD eess.AS 57%

Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model

Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper

机构 * University of Rochester(罗切斯特大学) Microsoft Research(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted at EUSIPCO 2025 - 5 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11639 2025-11-12 cs.IR 50%

OneRec-Think: In-Text Reasoning for Generative Recommendation

Zhanyu Liu, Shiyao Wang, Xingmei Wang, Rongzhou Zhang, Jiaxin Deng, Honghui Bao, Jinghao Zhang, Wuchao Li, Pengfei Zheng, Xiangyu Wu, Yifei Hu, Qigen Hu, Xinchen Luo, Lejian Ren, Zixing Zhang, Qianqian Wang, Kuo Cai, Yunfan Wu, Hongtao Cheng, Zexuan Cheng, Lu Ren, Huanjie Wang, Yi Su, Ruiming Tang, Kun Gai, Guorui Zhou

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00598 2025-11-12 cs.CV 50%

DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation

Boyi Li, Ce Zhang, Richard M. Timmerman, Wenxuan Bao

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07598 2025-11-12 physics.app-ph 50%

Passive Acoustic Monitoring of Underwater Well Leakages with Machine Learning: A Review

Guanlin Zhu, Zechun Deng, Jiaxin Shen, Junchi Yang

专题命中 其他安全 :safety(abstract)

Comments 10 pages, 5 figures, 4 equations, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏