arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-17 至 2025-12-17 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 10 篇

2512.14019 2025-12-17 cs.LG q-bio.QM 79%

EXAONE Path 2.5: Pathology Foundation Model with Multi-Omics Alignment

EXAONE Path 2.5:多组学对齐的病理基础模型

Juseung Yun, Sunwoo Yu, Sumin Ha, Jonghyun Kim, Janghyeon Lee, Jongseong Jang, Soonyoung Lee

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

AI总结 EXAONE Path 2.5通过多组学对齐构建病理基础模型,实现更全面的肿瘤生物学建模,展现高效率和适应性,推动精准肿瘤学发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11827 2025-12-17 cs.CY cs.AI cs.CV 73%

Assessing Greenspace Attractiveness with ChatGPT, Claude, and Gemini: Do AI Models Reflect Human Perceptions?

利用ChatGPT、Claude和Gemini评估绿地吸引力:AI模型能反映人类感知吗?

Milad Malekzadeh, Magdalena Biernacka, Elias Willberg, Jussi Torkko, Edyta Łaszkiewicz, Tuuli Toivonen

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

AI总结 本文研究了AI模型在评估绿地吸引力方面的表现,发现其在正式绿地和非正式空间的判断一致性较高,但存在对安全性和本地嵌入质量的低估问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21112 2025-12-17 cs.AI cs.CE cs.CL cs.CY cs.IR 67%

Optimizing Large Language Models for ESG Activity Detection in Financial Texts

优化大型语言模型以检测金融文本中的ESG活动

Mattia Birti, Andrea Maurino, Francesco Osborne

机构 * Department of Informatics, Systems and Communication, University of Milano-Bicocca(信息学、系统与通信系,米兰-比科卡大学) University of Milano-Bicocca(米兰-比科卡大学) The Open University(开放大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出通过微调优化大型语言模型,提升金融文本中ESG活动检测的准确性。

Comments Published in the Proceedings of the ACM International Conference on AI in Finance (ICAIF). ACM version

Journal ref Proceedings of the ACM International Conference on AI in Finance (ICAIF), 2024, ACM

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14562 2025-12-17 cs.CL cs.AI 62%

Polypersona: Persona-Grounded LLM for Synthetic Survey Responses

Polypersona:基于身份的LLM生成合成调查回应

Tejaswani Dash, Dinesh Karri, Anudeep Vurity, Gautam Datla, Tazeem Ahmad, Saima Rafi, Rohith Tangudu

机构 * George Mason University, Virginia, USA(乔治·马歇尔大学) New Jersey Institute of Technology (NJIT), New Jersey, USA(新泽西理工学院) University of Southern Queensland, Queensland, Australia(南方昆士兰大学) Edinburgh Napier University, Edinburgh, Scotland(爱丁堡纳皮尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 PolyPersona通过基于身份的微调,使小型语言模型能生成可靠且一致的合成调查数据,实现多领域高效调查数据生成。

Comments Accepted in IEEE Bigdata 2025- LLMs4ALL

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14330 2025-12-17 cs.CY cs.AI cs.CR 62%

Criminal Liability in AI-Enabled Autonomous Vehicles: A Comparative Study

人工智能赋能的自动驾驶车辆中的刑事责任:比较研究

Sahibpreet Singh, Manjit Singh

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文通过比较研究分析人工智能赋能自动驾驶车辆中的刑事责任问题,探讨不同国家的监管差异及责任归属机制,强调全球法律统一对技术发展的重要性。

Comments Published in Journal of University Institute of Legal Studies, Vol. 18, Issue 1, pp. 57-78, 2025

Journal ref Journal of University Institute of Legal Studies 18(1), 57-78 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08743 2025-12-17 cs.AI cs.MA 57%

Single-Agent Scaling Fails Multi-Agent Intelligence: Towards Foundation Models with Native Multi-Agent Intelligence

单智能体扩展无法实现多智能体智能:迈向具有原生多智能体智能的基础模型

Shuyue Hu, Haoyang Yan, Yiqun Zhang, Yang Chen, Dongzhan Zhou, Lei Bai

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 本文指出单智能体扩展无法实现多智能体智能,提出构建具有原生多智能体智能的基础模型的关键方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17399 2025-12-17 cs.CL 57%

DIWALI: Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context

DIWALI:多样性与包容性-aware的文化特定项目用于印度:数据集和对LLM进行文化文本适应的评估

Pramit Sahoo, Maharaj Brahma, Maunendra Sankar Desarkar

机构 * Natural Language and Information Processing Lab (NLIP) Indian Institute of Technology Hyderabad(自然语言与信息处理实验室(NLIP)印度理工学院海得拉巴)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 DIWALI数据集通过17个文化维度的8,000个文化概念,评估LLM在印度文化文本适应中的文化意识与一致性。

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14499 2025-12-17 cs.CV 50%

Native Intelligence Emerges from Large-Scale Clinical Practice: A Retinal Foundation Model with Deployment Efficiency

原生智能源自大规模临床实践:一种具有部署效率的视网膜基础模型

Jia Guo, Jiawei Du, Shengzhu Yang, Shuai Lu, Wenquan Cheng, Kaiwen Zhang, Yihua Sun, Chuhong Yang, Weihang Zhang, Fang Chen, Yilan Wu, Lie Ju, Guochen Ning, Longfei Ma, Huiping Yao, Jinyuan Wang, Peilun Shi, Yukun Zhou, Jie Xu, Pearse A. Keane, Hanruo Liu, Hongen Liao, Ningli Wang, Huiqi Li

机构 * School of Biomedical Engineering, Tsinghua Medicine, Tsinghua University(生物医学工程学院,清华大学医学部,清华大学) School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) Beijing Key Laboratory of Intelligent Diagnosis Technology and Equipment for Optic Nerve-Related Eye Diseases(北京智能诊断技术与设备重点实验室) School of Medical Technology, Beijing Institute of Technology(医学技术学院,北京理工大学) Beijing Tongren Hospital, Capital Medical University(北京同仁医院,首都医科大学) School of Biomedical Engineering, Shanghai Jiaotong University(生物医学工程学院,上海交通大学) Beijing Visual Science and Translational Eye Research Institute (BERI), Beijing Tsinghua Changgung Hospital Eye Center, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University Institute of Ophthalmology(北京视觉科学与转化眼研所(BERI),北京清华长庚医院眼科中心,临床医学学院,清华大学医学部,清华大学眼科学院)

专题命中 安全评测 :alignment(abstract)

AI总结 ReVision通过直接从真实临床数据中学习,实现了在低资源环境下高效部署的视网膜基础模型,无需额外标注即可在多种医学任务中取得优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11065 2025-12-17 cs.HC 50%

Immutable Explainability: Fuzzy Logic and Blockchain for Verifiable Affective AI

不可变的可解释性:模糊逻辑与区块链用于可验证的情感人工智能

Marcelo Fransoy, Alejandro Hossian, Hernán Merlino

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出不可变可解释性架构,结合模糊逻辑与区块链,实现情感AI的透明决策和可信审计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13977 2025-12-17 cs.CV 50%

XAI-Driven Diagnosis of Generalization Failure in State-Space Cerebrovascular Segmentation Models: A Case Study on Domain Shift Between RSNA and TopCoW Datasets

基于XAI的州空间脑血管分割模型泛化失败诊断:在RSNA与TopCoW数据集域偏移案例研究

Youssef Abuzeid, Shimaa El-Bana, Ahmad Al-Kabbany

机构 * Department of Electronics and Electrical Communications Engineering(电子与电气通信工程系) Cairo University(开罗大学) Multimedia Interaction and Communication Lab(多媒体交互与通信实验室) Arab Academy for Science and Technology(阿拉伯科学与技术学院) Wearables, Biosensing, and Biosignal Processing Research lab(可穿戴设备、生物传感与生物信号处理研究实验室)

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出基于XAI的两阶段方法,诊断州空间模型在脑血管分割中的泛化失败原因,揭示模型因关注机制失效而学习虚假相关性。

详情

展开后加载摘要…

URL PDF HTML 收藏