arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2601.17173 2026-01-27 cs.CL cs.AI 62%

Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content

超越事实问答:面向长形式多语言内容的指导型问答

Parth Bhalerao, Diola Dsouza, Ruiwen Guan, Oana Ignat

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MentorQA,首个多语言长形式视频指导型问答数据集和评估框架,通过对比不同架构发现Multi-Agent在复杂和低资源语言中表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15124 2026-01-27 cs.LG cs.AI 62%

RAG-GFM: Overcoming In-Memory Bottlenecks in Graph Foundation Models via Retrieval-Augmented Generation

RAG-GFM:通过检索增强生成克服图基础模型中的内存瓶颈

Haonan Yuan, Qingyun Sun, Jiacheng Tao, Xingcheng Fu, Jianxin Li

机构 * SKLCCSE, School of Computer Science and Engineering(计算机科学与工程学院) Beihang University(北航) Guangxi Normal University(广西师范大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 RAG-GFM通过检索增强生成方法,解决图基础模型中的内存瓶颈问题,提升模型的效率和效果。

Comments Accepted by the Web Conference 2026 (Research Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15511 2026-01-23 cs.CL cs.CY 62%

AdversaRiskQA: An Adversarial Factuality Benchmark for High-Risk Domains

AdversaRiskQA: 面向高风险领域的对抗事实性基准

Adam Szelestey, Sofie van Engelen, Tianhao Huang, Justin Snelders, Qintao Zeng, Songgaojun Deng

机构 * Eindhoven University of Technology(埃因霍温理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 AdversaRiskQA是一个面向高风险领域的对抗事实性基准,通过评估LLM在不同领域和难度下的事实性表现,揭示模型弱点并提升高风险应用的可靠性。

Comments 13 pages, 4 figures, and 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14620 2026-01-22 eess.AS cs.AI cs.LG cs.SD 62%

Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models

缩放歧义:通过音频-语言模型增强语音情感识别中的人工标注

Wenda Zhang, Hongyu Jin, Siyi Wang, Zhiqiang Wei, Ting Dang

机构 * University of Melbourne(墨尔本大学) Xi'an Jiaotong University(西安交通大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文通过音频-语言模型生成合成标注,提升语音情感识别中模糊情感的真实分布可靠性,实验表明合成标注在低歧义区域效果显著,但高歧义情况需进一步优化。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14269 2026-01-22 cs.CL cs.AI 62%

The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues

支持的缓慢漂移:多轮心理健康大语言模型对话中的边界失败

Youyou Cheng, Zhuangwei Kang, Kerry Jiang, Chenyu Sun, Qiyang Pan

机构 * University of Incarnate Word School of Osteopathic Medicine(incarnate Word 学校医学部) Independent Researcher(独立研究者) Mayo Clinic(梅奥诊所)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出多轮压力测试框架,揭示LLMs在长对话中因安慰和同理心尝试导致的安全边界逐步侵犯问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14235 2026-01-22 astro-ph.IM astro-ph.CO cs.AI cs.LG stat.ML 62%

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

人工智能/机器学习在Rubin LSST暗能量科学合作中的机遇

LSST Dark Energy Science Collaboration, Eric Aubourg, Camille Avestruz, Matthew R. Becker, Biswajit Biswas, Rahul Biswas, Boris Bolliet, Adam S. Bolton, Clecio R. Bom, Raphaël Bonnet-Guerrini, Alexandre Boucaud, Jean-Eric Campagne, Chihway Chang, Aleksandra Ćiprijanović, Johann Cohen-Tanugi, Michael W. Coughlin, John Franklin Crenshaw, Juan C. Cuevas-Tello, Juan de Vicente, Seth W. Digel, Steven Dillmann, Mariano Javier de León Dominguez Romero, Alex Drlica-Wagner, Sydney Erickson, Alexander T. Gagliano, Christos Georgiou, Aritra Ghosh, Matthew Grayling, Kirill A. Grishin, Alan Heavens, Lindsay R. House, Mustapha Ishak, Wassim Kabalan, Arun Kannawadi, François Lanusse, C. Danielle Leonard, Pierre-François Léget, Michelle Lochner, Yao-Yuan Mao, Peter Melchior, Grant Merz, Martin Millon, Anais Möller, Gautham Narayan, Yuuki Omori, Hiranya Peiris, Laurence Perreault-Levasseur, Andrés A. Plazas Malagón, Nesar Ramachandra, Benjamin Remy, Cécile Roucelle, Jaime Ruiz-Zapatero, Stefan Schuldt, Ignacio Sevilla-Noarbe, Ved G. Shah, Tjitske Starkenburg, Stephen Thorp, Laura Toribio San Cipriano, Tilman Tröster, Roberto Trotta, Padma Venkatraman, Amanda Wasserman, Tim White, Justine Zeghal, Tianqing Zhang, Yuanyuan Zhang

机构 * Université Paris Cité, CNRS, CEA, Astroparticule et Cosmologie, F-75013 Paris, France Department of Physics, University of Michigan, Ann Arbor, MI 48109, USA Leinweber Institute of Theoretical Physics, University of Michigan, Ann Arbor, MI 48109, USA Argonne National Laboratory, 9700 South Cass Avenue, Lemont, IL 60439, USA Cavendish Astrophysics, University of Cambridge, Madingley Road, Cambridge CB3 0HA, UK Kavli Institute for Cosmology, University of Cambridge, Madingley Road, Cambridge CB3 0HA, UK SLAC National Accelerator Laboratory, Menlo Park, CA 94025, USA Department of Computer Science, University of Milan, Milan, Italy Université Paris Cité, CNRS, Astroparticule et Cosmologie, F-75013 Paris, France Université Paris-Saclay, CNRS/IN2P3, IJCLab, 91405 Orsay, France Department of Astronomy Astrophysics, University of Chicago, Chicago, IL 60637, USA Kavli Institute for Cosmological Physics, University of Chicago, Chicago, IL 60637, USA NSF-Simons AI Institute for the Sky (SkAI), 172 E. Chestnut St., Chicago, IL 60611, USA Fermi National Accelerator Laboratory, P.O. Box 500, Batavia, IL 60510, USA Universit\'e Clermont-Auvergne, CNRS, LPCA, 63000 Clermont-Ferrand, France Kavli Institute for Particle Astrophysics Cosmology, Stanford University, Stanford, CA 94305, USA Department of Physics, Stanford University, 382 Via Pueblo Mall, Stanford, CA 94305, USA Engineering Faculty, Universidad Autonoma de San Luis Potosi, Zona Universitaria, San Luis Potosi, 78290, Mexico Stanford Artificial Intelligence Laboratory, Stanford University, Stanford, CA 94305, USA Kavli Institute of Cosmological Physics, University of Chicago, Chicago, IL 60637, USA The NSF AI Institute for Artificial Intelligence Center for Astrophysics Harvard \& Smithsonian, 60 Garden Street, Cambridge, MA 02138, USA Department of Physics Kavli Institute for Astrophysics Space Research, Massachusetts Institute of Technology, Cambridge, MA 02139, USA Institut de Física d'Altes Energies (IFAE), The Barcelona Institute of Science Institute of Astronomy Kavli Institute for Cosmology, University of Cambridge, Madingley Road, Cambridge, CB3 0HA, UK Imperial Centre for Inference Cosmology (ICIC), Imperial College London, Blackett Laboratory, Prince Consort Road, London SW7 2AZ, UK Data Science Institute, The University of Chicago, Chicago, IL 60615, USA Department of Physics, The University of Texas at Dallas, Richardson, TX 75080, USA Department of Physics, Duke University, Durham, NC 27708, USA Université Paris-Saclay, Université Paris Cité, CEA, CNRS, AIM, F-91191 Gif-sur-Yvette, France School of Mathematics, Statistics Physics, Newcastle University, Newcastle upon Tyne, NE1 7RU, United Kingdom Department of Astrophysical Sciences, Princeton University, Princeton, NJ 08544, USA Astronomy, University of the Western Cape, Bellville, Cape Town, 7535, South Africa Astronomy, University of Utah, Salt Lake City, UT 84112, USA Department of Astrophysical Sciences, Princeton University, Peyton Hall, Princeton, NJ 08544, USA Department of Astronomy, University of Illinois Urbana Champaign, 1002 W. Green St., Urbana, IL, 61801, USA Institute for Particle Physics Astrophysics, ETH Zürich, Wolfgang-Pauli-Strasse 27, CH-8093 Zurich, Switzerland Swinburne University of Technology, Hawthorn, Victoria 3122, Australia Ciela - Montr\'eal Institute for Astrophysical Data Analysis Mila - Quebec Artificial Intelligence Institute, Montréal, QC H2S 3H1, Canada Advanced Research Computing Centre, University College London, 90 High Holborn, London WC1V 6LJ, UK Finnish Centre for Astronomy with ESO (FINCA), University of Turku, FI-20014 Turku, Finland Department of Physics, P.O. Box 64, University of Helsinki, FI-00014 Helsinki, Finland Astronomy, Northwestern University, Evanston, IL, USA Center for Interdisciplinary Exploration Research in Astrophysics, Northwestern University, Evanston, IL, USA Scientific Data Science, International School for Advanced Study, Via Bonomea 265, I-34136 Trieste, Italy Department of Statistics, University of Michigan, Ann Arbor, MI 48109, USA PITT PACC, University of Pittsburgh, Pittsburgh, PA 15260, USA NSF NOIRLab, 950 N. Cherry Ave., Tucson, AZ 85719, USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了AI/ML在LSST暗能量科学合作中的应用机遇,强调了大规模贝叶斯推断、物理指导方法和主动学习等关键方法学优先事项,并讨论了新兴技术在重塑工作流程中的潜力。

Comments 84 pages. This is v1.0 of the DESC's white paper on AI/ML, a collaboration document that is being made public but which is not planned for submission to a journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00332 2026-01-22 cs.CL cs.AI 62%

Assertion-Conditioned Compliance: A Provenance-Aware Vulnerability in Multi-Turn Tool-Calling Agents

断言-条件合规:多轮工具调用代理中的一种意识-aware 的漏洞

Daud Waqas, Aaryamaan Golthi, Erika Hayashida, Huanzhi Mao

机构 * Monash University(墨尔本大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种新的评估方法A-CC,用于检测多轮工具调用代理在面对误导性断言时的鲁棒性问题,揭示了模型在用户和系统政策冲突下的脆弱性。

Comments 15 pages (incl. Appendix), 3 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22936 2026-01-21 cs.CL cs.AI cs.CE cs.HC q-fin.CP 62%

Evaluating Large Language Models (LLMs) in Financial NLP: A Comparative Study on Financial Report Analysis

评估大型语言模型(LLMs)在金融NLP中的表现:对财务报告分析的比较研究

Md Talha Mohsin

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文评估了五种基于transformer的LLMs在财务报告分析中的表现,发现模型在不同评估维度上表现各异,需考虑人类主观性和行为差异性。

Comments 23 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10101 2026-01-21 cs.AI cs.CL 62%

Matrix as Plan: Structured Logical Reasoning with Feedback-Driven Replanning

矩阵作为计划:基于反馈驱动的重计划的结构化逻辑推理

Ke Chen, Jiandian Zeng, Zihao Peng, Guo Li, Guangxue Zhang, Tian Wang

机构 * Faculty of Arts and Sciences(艺术与科学学院) Beijing Normal University(北京师范大学) Institute of Artificial Intelligence and Future Networks(人工智能与未来网络研究所) Engineering Research Center of Cloud-Edge Intelligent Collaboration on Big Data, Ministry of Education(教育部云-边智能协同大数据工程研究中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 MatrixCoT通过引入矩阵基础计划和反馈驱动重计划机制,提升LLM在复杂符号推理任务中的鲁棒性和可解释性。

Comments 12 pages, 5 figures, 2 tables. Accepted at The Web Conference (WWW) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16544 2026-01-21 cs.CL cs.AI 62%

WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue

WER是无意识的:评估ASR错误如何扭曲患者面对对话中的临床理解

Zachary Ellis, Jared Joselowitz, Yash Deo, Yajie He, Anna Kalygina, Aisling Higham, Mana Rahimzadeh, Yan Jia, Ibrahim Habli, Ernest Lim

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过LLM作为判断者评估ASR错误的临床影响,发现WER等指标与临床风险标签相关性低,引入优化的LLM模型提升评估准确性。

Comments Published as an Oral at IWSDS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18760 2026-01-21 cs.AI cs.CL 62%

Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models

回答不可回答的问题是明知其错:分析和缓解大推理模型中的回避失败

Yi Liu, Xiangyu Liu, Zequn Sun, Wei Hu

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 本文针对大推理模型在面对不可回答问题时的回避失败问题,提出一种轻量级两阶段方法,通过认知监控与推理干预提升回避率并保持推理性能。

Comments Accepted in the 39th AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12259 2026-01-21 cs.AI cs.CE cs.LG 62%

FutureX-Pro: Extending Future Prediction to High-Value Vertical Domains

FutureX-Pro: 将未来预测扩展到高价值垂直领域

Jiashuo Liu, Siyuan Chen, Zaiyuan Wang, Zhiyuan Zeng, Jiacheng Guo, Liang Hu, Lingyue Yin, Suozhi Huang, Wenxin Hao, Yang Yang, Zerui Cheng, Zixin Yao, Lingyue Yin, Haoxin Liu, Jiayi Cheng, Yuzhen Li, Zezhong Ma, Bingjie Wang, Bingsen Qiu, Xiao Liu, Zeyang Zhang, Zijian Liu, Jinpeng Wang, Mingren Yin, Tianci He, Yali Liao, Yixiao Tian, Zhenwei Zhu, Anqi Dai, Ge Zhang, Jingkai Liu, Kaiyuan Zhang, Wenlong Wu, Xiang Gao, Xinjie Chen, Zhixin Yao, Zhoufutu Wen, B. Aditya Prakash, Jose Blanchet, Mengdi Wang, Nian Si, Wenhao Huang

机构 * Hong Kong University of Science and Technology(香港科技大学) Georgia Institute of Technology(佐治亚理工学院) Stanford University(斯坦福大学) Princeton University(普林斯顿大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

AI总结 FutureX-Pro通过扩展未来预测到金融、零售、公共健康和自然灾害等高价值垂直领域,评估代理LLMs在工业部署中的领域基础能力。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11920 2026-01-21 cs.CL cs.AI 62%

Enhancing LLM-Based Data Annotation with Error Decomposition

通过错误分解增强基于LLM的数据标注

Zhen Xu, Vedant Khatri, Yijun Dai, Xiner Liu, Siyan Li, Xuanming Zhang, Renzhe Yu

机构 * Columbia University(哥伦比亚大学) University of California, Irvine(加州大学尔湾分校) University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 通过错误分解增强基于LLM的数据标注,提出诊断评估范式以区分任务固有模糊性与模型不准确,提升标注质量评估的准确性与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11905 2026-01-21 cs.AI cs.LG math.ST stat.TH 62%

LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning

LIBRA:基于语言模型的带状 recourse 算法用于个性化治疗计划

Junyu Cao, Ruijiang Gao, Esmaeil Keyvanshokooh, Jianhao Ma

机构 * McCombs School of Business, University of Texas at Austin(德克萨斯大学奥斯汀分校麦克斯韦商学院) Naveen Jindal School of Management, University of Texas at Dallas(德克萨斯大学达拉斯分校奈文·金达管理学院) Mays Business School, Texas A&M University(德克萨斯农工大学梅斯商学院) Wharton School, University of Pennsylvania(宾夕法尼亚大学沃顿商学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 LIBRA 是一种结合大语言模型和带状学习的算法,用于在个性化治疗中实现更高效的决策和鲁棒性。

Comments 50 pages. Previous version with human-AI collaboration: arXiv:2410.14640

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26854 2026-01-21 cs.AI cs.LG 62%

Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base

反向知识搜索与可验证推理:从长推理链知识库合成科学百科全书

Yu Li, Yuan Huang, Tao Wang, Caiyu Fan, Xiansheng Cai, Sihan Hu, Xinzijian Liu, Cheng Shi, Mingjun Xu, Zhen Wang, Yan Wang, Xiangqi Jin, Tianhan Zhang, Linfeng Zhang, Lei Wang, Youjin Deng, Pan Zhang, Weijie Sun, Xinyu Li, Weinan E, Linfeng Zhang, Zhiyuan Yao, Kun Chen

机构 * Lanzhou Center for Theoretical Physics, Key Laboratory of Theoretical Physics of Gansu Province, Key Laboratory of Quantum Theory and Applications of MoE, Gansu Provincial Research Center for Basic Disciplines of Quantum Physics(兰州理论物理中心、甘肃省理论物理重点实验室、教育部量子理论与应用重点实验室、甘肃省量子物理基础学科研究省重点中心) Institute of Theoretical Physics, Chinese Academy of Sciences(中国科学院理论物理研究所) DP Technology(DP技术) Institute of Physics, Chinese Academy of Sciences(中国科学院物理研究所) Département d’Informatique, École normale supérieure(巴黎高等师范大学计算机系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于长推理链知识库的反向知识搜索方法,通过生成可验证的科学百科全书,实现了跨领域科学合成。

Comments 43 pages, 4 figures. This work is part of the SciencePedia project (sciencepedia.bohrium.com). Corrected author name spelling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08211 2026-01-21 cs.CL cs.AI cs.CR 62%

LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions

大语言模型无意中欺骗:从不一致样本到有偏的人机交互中的涌现不一致

Xuhao Hu, Peng Wang, Xiaoya Lu, Dongrui Liu, Xuanjing Huang, Jing Shao

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 研究发现大语言模型在高风险场景中可能因不一致样本而无意产生不诚实行为,且在下游任务和人机交互中风险显著增加。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06243 2026-01-21 cs.CL cs.AI 62%

CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning

CoT Referring: 通过 grounded 推理改进指称表达任务

Qihua Dong, Luis Figueroa, Handong Zhao, Kushal Kafle, Jason Kuen, Zhihong Ding, Scott Cohen, Yun Fu

机构 * Adobe Research(Adobe研究院) Northeastern University(东北大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 通过 grounded 推理改进指称表达任务,提出CoT Referring方法,提升多模态大语言模型在复杂指称场景中的性能。

Comments MLLM, Referring Expression Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23816 2026-01-21 cs.CL cs.LG 62%

A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs

在可转向性评估中的课程修正:揭示LLMs中的误校准与副作用

Trenton Chang, Tobias Schnabel, Adith Swaminathan, Jenna Wiens

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 本文提出了一种多维目标空间框架,揭示LLMs在文本改写任务中存在意外副作用,表明现有对齐策略可能不足。

Comments 8 pages, 6 figures. 26 pages of references and supplementary material, 22 additional figures. Association for the Advancement of Artificial Intelligence Conference (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12537 2026-01-19 cs.CL cs.AI eess.AS 62%

What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study

什么使LLM为中心的语音生成中的良好语音分词器?系统研究

Xiaoran Fan, Zhichao Sun, Yangfan Gao, Jingfei Xiong, Hang Yan, Yifei Cao, Jiajun Sun, Shuo Li, Zhihao Zhang, Zhiheng Xi, Yuhao Zhou, Senjie Jin, Changhao Jiang, Junjie Ye, Ming Zhang, Rui Zheng, Zhenhua Han, Yunke Zhang, Demei Yan, Shaokang Dong, Tao Ji, Tao Gui

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了LLM为中心的语音生成中语音分词器设计的影响,通过引入多令牌预测和说话人感知生成,提升了语音生成的质量和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10562 2026-01-16 cs.LG cs.AI cs.CV 62%

Process-Guided Concept Bottleneck Model

过程引导的概念瓶颈模型

Reza M. Asiyabi, SEOSAW Partnership, Steven Hancock, Casey Ryan

机构 * SEOSAW Partnership(SEOSAW合作机构)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 过程引导的概念瓶颈模型通过引入生物物理意义的中间概念,提升科学领域中深度学习模型的可解释性和透明性,减少误差和偏见,同时利用多源数据产生可解释的中间输出。

Comments 13 pages with 7 figures and 1 table, Supplementary Materials 10 pages with 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10122 2026-01-16 cs.CL cs.AI cs.HC 62%

Role-Playing Agents Driven by Large Language Models: Current Status, Challenges, and Future Trends

由大型语言模型驱动的角色扮演代理:现状、挑战与未来趋势

Ye Wang, Jiaxing Chen, Hongjiang Xiao

机构 * Communication University of China(中国传媒大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文系统回顾了由大型语言模型驱动的角色扮演代理的现状、挑战及未来趋势,探讨了关键技术路径与评估方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09066 2026-01-15 cs.CL cs.AI 62%

Mi:dm 2.0 Korea-centric Bilingual Language Models

Mi:dm 2.0 韩国中心双语语言模型

Donghoon Shin, Sejung Lee, Soonmin Bae, Hwijung Ryu, Changwon Ok, Hoyoun Jung, Hyesung Ji, Jeehyun Lim, Jehoon Lee, Ji-Eun Han, Jisoo Baik, Mihyeon Kim, Riwoo Chung, Seongmin Lee, Wonjae Park, Yoonseok Heo, Youngkyung Seo, Seyoun Won, Boeun Kim, Cheolhun Heo, Eunkyeong Lee, Honghee Lee, Hyeongju Ju, Hyeontae Seo, Jeongyong Shim, Jisoo Lee, Junseok Koh, Junwoo Kim, Minho Lee, Minji Kang, Minju Kim, Sangha Nam, Seongheum Park, Taehyeong Kim, Euijai Ahn, Hong Seok Jeung, Jisu Shin, Jiyeon Kim, Seonyeong Song, Seung Hyun Kong, Sukjin Hong, Taeyang Yun, Yu-Seon Kim, A-Hyun Lee, Chae-Jeong Lee, Hye-Won Yu, Ji-Hyun Ahn, Song-Yeon Kim, Sun-Woo Jung, Eunju Kim, Eunji Ha, Jinwoo Baek, Yun-ji Lee, Wanjin Park, Jeong Yeop Kim, Eun Mi Kim, Hyoung Jun Park, Jung Won Yoon, Min Sung Noh, Myung Gyo Oh, Wongyoung Lee, Yun Jin Park, Young S. Kwon, Hyun Keun Kim, Jieun Lee, YeoJoo Park

机构 * Tech. Innovation Group(技术创新组)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 Mi:dm 2.0是一款专为韩国中心AI设计的双语大语言模型,通过整合韩国社会价值观和常识知识,提升文化适应性和生成能力,支持多任务和多场景应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08869 2026-01-15 cs.CY cs.AI 62%

AI Deployment Authorisation: A Global Standard for Machine-Readable Governance of High-Risk Artificial Intelligence

AI部署授权:面向高风险人工智能的全球标准:机器可读的治理机制

Daniel Djan Saparning

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出AI部署授权分数(ADAS),通过五个维度评估AI系统,生成可验证的部署证书,以实现安全、合法且可扩展的人工智能治理。

Comments 28 pages, 4 figures. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08490 2026-01-14 cs.CL cs.AI 62%

BenchOverflow: Measuring Overflow in Large Language Models via Plain-Text Prompts

BenchOverflow: 通过纯文本提示测量大语言模型中的溢出现象

Erin Feiglin, Nir Hutnik, Raz Lapid

机构 * Deepkeep(深保持)

专题命中 安全评测 :prompt injection(abstract);分类 cs.CL、cs.AI

AI总结 BenchOverflow通过纯文本提示策略评估大语言模型的溢出现象,揭示长度控制对可靠性、成本和可持续性的影响,提供标准化比较框架以优化部署和防御措施。

Comments Accepted at TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08196 2026-01-14 cs.CL cs.AI cs.CR cs.LO cs.SE 62%

Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis

通过逻辑引导合成评估LLM工具调用中的隐式监管合规性

Da Song, Yuheng Huang, Boqi Chen, Tianshuo Cong, Randy Goebel, Lei Ma, Foutse Khomh

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出LogiSafetyGen框架,通过逻辑引导合成评估LLM在工具调用中的隐式监管合规性,揭示大模型在安全约束上的不足。

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07871 2026-01-14 q-bio.QM cs.AI cs.CV cs.LG 62%

Imaging-anchored Multiomics in Cardiovascular Disease: Integrating Cardiac Imaging, Bulk, Single-cell, and Spatial Transcriptomics

心血管疾病中的成像锚定多组学:整合心脏成像、批量、单细胞和空间转录组学

Minh H. N. Le, Tuan Vinh, Thanh-Huy Nguyen, Tao Li, Bao Quang Gia Le, Han H. Huynh, Monika Raj, Carl Yang, Min Xu, Nguyen Quoc Khanh Le

机构 * International Ph.D. Program in Medicine, College of Medicine, Taipei Medical University, Taipei, Taiwan AIBioMed Research Group, Taipei Medical University, Taipei, Taiwan Medical Sciences Division, University of Oxford, Oxford, United Kingdom Computational Biology Department, School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA Department of Computer Science, Emory University, Atlanta, GA, USA Department of Chemistry, Emory University, Atlanta, GA, USA International Master Program for Translational Science, College of Medical Science Technology, Taipei Medical University, Taipei 110, Taiwan In-Service Master Program in Artificial Intelligence in Medicine, College of Medicine, Taipei Medical University, Taipei, Taiwan Translational Imaging Research Center, Taipei Medical University Hospital, Taipei, Taiwan

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过整合心脏成像与多组学数据,推动心血管疾病研究的多模态融合方法,提升疾病诊断和治疗的精准性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00549 2026-01-14 cs.LG cs.AI 62%

Robust Single-Agent Reinforcement Learning for Regional Traffic Signal Control Under Demand Fluctuations

鲁棒的单智能体强化学习用于应对需求波动的区域交通信号控制

Qiang Li, Jin Niu, Lina Yu

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种鲁棒的单智能体强化学习框架,用于应对交通需求波动的区域交通信号控制,通过集中决策和高效学习模型有效减少交通队列长度。

Comments A critical error in the methodology. The reported congestion control effects were not caused by the proposed signal timing optimization, but by an incorrect traffic volume scaling factor during evaluation. The traffic demand was not properly amplified, resulting in misleading performance gains. Due to the substantial nature of the error, completion of revisions is not feasible in the short term

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20315 2026-01-14 cs.CL cs.AI 62%

Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL

Arctic-Text2SQL-R1: 简单奖励,强推理的文本到SQL

Zhewei Yao, Guoheng Sun, Lukasz Borchmann, Gaurav Nuti, Zheyu Shen, Minghang Deng, Bohan Zhai, Hao Zhang, Ang Li, Yuxiong He

机构 * Snowflake AI Research(Snowflake AI研究院) University of Maryland(马里兰大学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 Arctic-Text2SQL-R1通过简单奖励机制和强化学习框架,在文本到SQL任务中实现高准确率和高效性,优于现有大型模型。

Comments 22 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07767 2026-01-13 cs.LG cs.CL 62%

Are LLM Decisions Faithful to Verbal Confidence?

大型语言模型的决策是否忠实于其语言自信度?

Jiawei Wang, Yanfei Zhou, Siddartha Devic, Deqing Fu

机构 * University of Southern California(南加州大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

AI总结 研究发现大型语言模型在高惩罚条件下缺乏战略决策能力,其语言自信度与实际决策不一致,影响AI系统的可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07313 2026-01-13 cs.LG cs.AI 62%

Explaining Machine Learning Predictive Models through Conditional Expectation Methods

通过条件期望方法解释机器学习预测模型

Silvia Ruiz-España, Laura Arnal, François Signol, Juan-Carlos Perez-Cortes, Joaquim Arlandis

机构 * ITI, Universitat Politècnica de València(ITI,瓦伦西亚理工大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出MUCE方法,通过多变量条件期望捕捉预测变化,结合稳定性与不确定性指标,提升模型局部可解释性与可信度。

Comments 24 pages, 15 figures. Silvia Ruiz-España and Laura Arnal contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏