arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-29 至 2025-07-29 共收录 57 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2507.20774 2025-07-29 cs.AI 57%

evalSmarT: An LLM-Based Framework for Evaluating Smart Contract Generated Comments

Fatou Ndiaye Mbodji

机构 * SnT – University of Luxembourg(卢森堡大学SnT学院) Université Cheikh Anta Diop(谢赫·安塔·迪奥普大学) Huawei(华为)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 4 pages, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20451 2025-07-29 cs.AI 57%

STARN-GAT: A Multi-Modal Spatio-Temporal Graph Attention Network for Accident Severity Prediction

Pritom Ray Nobin, Imran Ahammad Rifat

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20046 2025-07-29 cs.CL 57%

Infogen: Generating Complex Statistical Infographics from Documents

Akash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay, Sriparna Saha

机构 * Indian Institute of Technology Patna(印度帕纳大学理工学院) Adobe Research(Adobe研究)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments ACL Main 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16674 2025-07-29 cs.LG q-bio.NC 57%

GASPnet: Global Agreement to Synchronize Phases

Andrea Alamia, Sabine Muzellec, Thomas Serre, Rufin VanRullen

机构 * CerCo, CNRS(CerCo与CNRS) Carney Institute for Brain Science, Brown University(脑科学研究所与布朗大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07317 2025-07-29 cs.CV 50%

ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation

Sherry X. Chen, Yi Wei, Luowei Zhou, Suren Kumar

机构 * Samsung AI Center(三星AI中心) Mountain View University of California, Santa Barbara(山景城加州大学圣巴巴拉分校)

专题命中 安全评测 :alignment(abstract)

Comments International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19965 2025-07-29 math.OC 50%

A Convex Optimization Approach to Model-Free Inverse Optimal Control with Provable Convergence

Meiling Yu, Lechen Feng, Lei Jiang, Yuan-Hua Ni

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14919 2025-07-29 cs.CV 50%

GenM$^3$: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation

Junyu Shi, Lijiang Liu, Yong Sun, Zhiyuan Zhang, Jinni Zhou, Qiang Nie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11824 2025-07-29 cs.CV 50%

KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities

Hsin-Ping Huang, Xinyi Wang, Yonatan Bitton, Hagai Taitelbaum, Gaurav Singh Tomar, Ming-Wei Chang, Xuhui Jia, Kelvin C. K. Chan, Hexiang Hu, Yu-Chuan Su, Ming-Hsuan Yang

专题命中 安全评测 :alignment(abstract)

Comments Project page: https://kitten-project.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 4 篇

2506.12088 2025-07-29 cs.CR cs.CY 86%

Risks & Benefits of LLMs & GenAI for Platform Integrity, Healthcare Diagnostics, Financial Trust and Compliance, Cybersecurity, Privacy & AI Safety: A Comprehensive Survey, Roadmap & Implementation Blueprint

Kiarash Ahi

专题命中 AI治理与伦理 :safety(title,abstract);AI safety(title);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19962 2025-07-29 cs.CL 57%

KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models

Seorin Kim, Dongyoung Lee, Jaejin Lee

机构 * Dept. of Data Science, Seoul National University(数据科学系,首尔国立大学) Dept. of Computer Science and Engineering, Seoul National University(计算机科学与工程系,首尔国立大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13138 2025-07-29 cs.CL 57%

Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation

Hadi Mohammadi, Tina Shahedi, Pablo Mosteiro, Massimo Poesio, Ayoub Bagheri, Anastasia Giachanou

机构 * Department of Methodology and Statistics, Utrecht University, The Netherlands(方法论与统计学系,乌特列支大学,荷兰) Department of Information and Computing Sciences, Utrecht University, The Netherlands(信息与计算科学系,乌特列支大学,荷兰) Queen Mary University of London, London, United Kingdom(伦敦女王玛丽大学,伦敦,英国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13175 2025-07-29 cs.AI 57%

Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era

Matthew E. Brophy

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 42 pages. Supplementary material included at end of article

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 15 篇

2507.20994 2025-07-29 cs.CV cs.AI 83%

Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM

Shen Li, Liuyi Yao, Wujia Niu, Lan Zhang, Yaliang Li

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI

Comments Codes and data are available at https://github.com/listen0425/Security-Tensors

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20028 2025-07-29 cs.CV cs.AI 70%

TAPS : Frustratingly Simple Test Time Active Learning for VLMs

Dhruv Sarkar, Aprameyo Chakrabartty, Bibhudatta Bhanja

机构 * IIT Kharagpur(印度理工学院Kharagpur分校)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19568 2025-07-29 cs.CY cs.AI cs.CE cs.LG 67%

Programmable Virtual Humans Toward Human Physiologically-Based Drug Discovery

You Wu, Philip E. Bourne, Lei Xie

机构 * Ph.D. Program in Computer Science The Graduate Center The City University of New York(计算机科学博士项目 美国城市大学研究生中心 新 York 美国) School of Data Science & Department of Biomedical Engineering University of Virginia(数据科学学院与生物医学工程系 美国弗吉尼亚大学) School of Pharmacy and Pharmaceutical Sciences & Center for Drug Discovery Northeastern University(药学与药学科学学院及药物发现中心 美国东北大学) Department of Computer Science Hunter College The City University of New York(计算机科学系 美国城市大学霍顿学院) Helen & Robert Appel Alzheimer’s Disease Research Institute Feil Family Brain & Mind Research Institute Weill Cornell Medicine Cornell University(海伦与罗伯特·阿普尔阿尔茨海默病研究所 费尔家族脑与心灵研究研究所 威尔·康奈尔医学院 康奈尔大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20191 2025-07-29 cs.LG cs.AI 62%

Partial Domain Adaptation via Importance Sampling-based Shift Correction

Cheng-Jun Guo, Chuan-Xian Ren, You-Wei Luo, Xiao-Lin Xu, Hong Yan

机构 * School of Mathematics, Sun Yat-Sen University(中山大学数学学院) School of Statistics and Mathematics, Guangdong University of Finance and Economics(广东财经大学统计与数学学院) Department of Electrical Engineering, City University of Hong Kong(香港城市大学电子工程系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13887 2025-07-29 cs.HC cs.CL cs.CY 62%

AI as a deliberative partner fosters intercultural empathy for Americans but fails for Latin American participants

Isabel Villanueva, Tara Bobinac, Binwei Yao, Junjie Hu, Kaiping Chen

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19526 2025-07-29 cs.LG cs.AI 62%

Quantizing Text-attributed Graphs for Semantic-Structural Integration

Jianyuan Bo, Hao Wu, Yuan Fang

机构 * Singapore Management University(新加坡国立管理大学) Beijing Normal University(北京师范大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted at KDD'2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20786 2025-07-29 cs.CL 57%

Automating Thematic Review of Prevention of Future Deaths Reports: Replicating the ONS Child Suicide Study using Large Language Models

Sam Osian, Arpan Dutta, Sahil Bhandari, Iain E. Buchan, Dan W. Joyce

专题命中 其他安全 :safety(abstract);分类 cs.CL

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20156 2025-07-29 cs.CV cs.AI 57%

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality

Daulet Toibazar, Kesen Wang, Sherif Mohamed, Abdulaziz Al-Badawi, Abdulrahman Alfulayt, Pedro J. Moreno

机构 * Humain Riyadh, KSA(利雅得人类,沙特阿拉伯)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01615 2025-07-29 cs.CL 57%

Large Language Models Are Human-Like Internally

Tatsuki Kuribayashi, Yohei Oseki, Souhaib Ben Taieb, Kentaro Inui, Timothy Baldwin

机构 * MBZUAI The University of Tokyo(东京大学) University of Mons(蒙斯大学) Tohoku University(东北大学) RIKEN(日本理化学研究所) The University of Melbourne(墨尔本大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments This is a pre-MIT Press publication version of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21033 2025-07-29 cs.CV 50%

GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

Yuhan Wang, Siwei Yang, Bingchen Zhao, Letian Zhang, Qing Liu, Yuyin Zhou, Cihang Xie

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校) The University of Edinburgh(爱丁堡大学) Adobe Project Page(Adobe项目页面)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20147 2025-07-29 cs.IR 50%

Integrating LLM-Derived Multi-Semantic Intent into Graph Model for Session-based Recommendation

Shuo Zhang, Xiao Li, Jiayi Wu, Fan Yang, Xiang Li, Ming Gao

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20025 2025-07-29 cs.CV 50%

Region-based Cluster Discrimination for Visual Representation Learning

Yin Xie, Kaicheng Yang, Xiang An, Kun Wu, Yongle Zhao, Weimo Deng, Zimin Ran, Yumeng Wang, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng

机构 * DeepGlint University of Technology Sydney(悉尼科技大学) Huawei London Research Center(华为伦敦研究中心) Imperial College London(伦敦帝国理工学院)

专题命中 其他安全 :alignment(abstract)

Comments Accepted as a highlight paper at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19860 2025-07-29 cs.RO cs.MA cs.SY eess.SY 50%

Homotopy-aware Multi-agent Navigation via Distributed Model Predictive Control

Haoze Dong, Meng Guo, Chengyi He, Zhongkui Li

机构 * School of Advanced Manufacturing and Robotics, Peking University(先进制造与机器人学院,北京大学) School of Computer Science and Engineering, Beihang University(计算机科学与工程学院,北航)

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18181 2025-07-29 eess.AS cs.SD 50%

SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding

Linye Wei, Shuzhang Zhong, Songqiang Xu, Runsheng Wang, Ru Huang, Meng Li

机构 * Institute for Artificial Intelligence, Peking University, Beijing, China(人工智能研究院,北京大学,北京) School of Integrated Circuits, Peking University, Beijing, China(集成电路学院,北京大学,北京) School of Software and Microelectronics, Peking University, Beijing, China(软件与微电子学院,北京大学,北京) Institute of Electronic Design Automation, Peking University, Wuxi, China(电子设计自动化研究院,北京大学,无锡) Beijing Advanced Innovation Center for Integrated Circuits, Beijing, China(北京集成电路先进创新中心,北京)

专题命中 其他安全 :alignment(abstract)

Comments Accepted by Design Automation Conference (DAC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16466 2025-07-29 cs.HC 50%

SceneLoom: Communicating Data with Scene Context

Lin Gao, Leixian Shen, Yuheng Zhao, Jiexiang Lan, Huamin Qu, Siming Chen

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏