arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-21 至 2026-01-21 共收录 131 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 51 篇

2601.10101 2026-01-21 cs.AI cs.CL 62%

Matrix as Plan: Structured Logical Reasoning with Feedback-Driven Replanning

矩阵作为计划:基于反馈驱动的重计划的结构化逻辑推理

Ke Chen, Jiandian Zeng, Zihao Peng, Guo Li, Guangxue Zhang, Tian Wang

机构 * Faculty of Arts and Sciences(艺术与科学学院) Beijing Normal University(北京师范大学) Institute of Artificial Intelligence and Future Networks(人工智能与未来网络研究所) Engineering Research Center of Cloud-Edge Intelligent Collaboration on Big Data, Ministry of Education(教育部云-边智能协同大数据工程研究中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 MatrixCoT通过引入矩阵基础计划和反馈驱动重计划机制,提升LLM在复杂符号推理任务中的鲁棒性和可解释性。

Comments 12 pages, 5 figures, 2 tables. Accepted at The Web Conference (WWW) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16544 2026-01-21 cs.CL cs.AI 62%

WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue

WER是无意识的:评估ASR错误如何扭曲患者面对对话中的临床理解

Zachary Ellis, Jared Joselowitz, Yash Deo, Yajie He, Anna Kalygina, Aisling Higham, Mana Rahimzadeh, Yan Jia, Ibrahim Habli, Ernest Lim

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过LLM作为判断者评估ASR错误的临床影响,发现WER等指标与临床风险标签相关性低,引入优化的LLM模型提升评估准确性。

Comments Published as an Oral at IWSDS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18760 2026-01-21 cs.AI cs.CL 62%

Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models

回答不可回答的问题是明知其错:分析和缓解大推理模型中的回避失败

Yi Liu, Xiangyu Liu, Zequn Sun, Wei Hu

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 本文针对大推理模型在面对不可回答问题时的回避失败问题,提出一种轻量级两阶段方法,通过认知监控与推理干预提升回避率并保持推理性能。

Comments Accepted in the 39th AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12259 2026-01-21 cs.AI cs.CE cs.LG 62%

FutureX-Pro: Extending Future Prediction to High-Value Vertical Domains

FutureX-Pro: 将未来预测扩展到高价值垂直领域

Jiashuo Liu, Siyuan Chen, Zaiyuan Wang, Zhiyuan Zeng, Jiacheng Guo, Liang Hu, Lingyue Yin, Suozhi Huang, Wenxin Hao, Yang Yang, Zerui Cheng, Zixin Yao, Lingyue Yin, Haoxin Liu, Jiayi Cheng, Yuzhen Li, Zezhong Ma, Bingjie Wang, Bingsen Qiu, Xiao Liu, Zeyang Zhang, Zijian Liu, Jinpeng Wang, Mingren Yin, Tianci He, Yali Liao, Yixiao Tian, Zhenwei Zhu, Anqi Dai, Ge Zhang, Jingkai Liu, Kaiyuan Zhang, Wenlong Wu, Xiang Gao, Xinjie Chen, Zhixin Yao, Zhoufutu Wen, B. Aditya Prakash, Jose Blanchet, Mengdi Wang, Nian Si, Wenhao Huang

机构 * Hong Kong University of Science and Technology(香港科技大学) Georgia Institute of Technology(佐治亚理工学院) Stanford University(斯坦福大学) Princeton University(普林斯顿大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

AI总结 FutureX-Pro通过扩展未来预测到金融、零售、公共健康和自然灾害等高价值垂直领域,评估代理LLMs在工业部署中的领域基础能力。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11920 2026-01-21 cs.CL cs.AI 62%

Enhancing LLM-Based Data Annotation with Error Decomposition

通过错误分解增强基于LLM的数据标注

Zhen Xu, Vedant Khatri, Yijun Dai, Xiner Liu, Siyan Li, Xuanming Zhang, Renzhe Yu

机构 * Columbia University(哥伦比亚大学) University of California, Irvine(加州大学尔湾分校) University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 通过错误分解增强基于LLM的数据标注,提出诊断评估范式以区分任务固有模糊性与模型不准确,提升标注质量评估的准确性与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11905 2026-01-21 cs.AI cs.LG math.ST stat.TH 62%

LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning

LIBRA:基于语言模型的带状 recourse 算法用于个性化治疗计划

Junyu Cao, Ruijiang Gao, Esmaeil Keyvanshokooh, Jianhao Ma

机构 * McCombs School of Business, University of Texas at Austin(德克萨斯大学奥斯汀分校麦克斯韦商学院) Naveen Jindal School of Management, University of Texas at Dallas(德克萨斯大学达拉斯分校奈文·金达管理学院) Mays Business School, Texas A&M University(德克萨斯农工大学梅斯商学院) Wharton School, University of Pennsylvania(宾夕法尼亚大学沃顿商学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 LIBRA 是一种结合大语言模型和带状学习的算法,用于在个性化治疗中实现更高效的决策和鲁棒性。

Comments 50 pages. Previous version with human-AI collaboration: arXiv:2410.14640

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26854 2026-01-21 cs.AI cs.LG 62%

Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base

反向知识搜索与可验证推理:从长推理链知识库合成科学百科全书

Yu Li, Yuan Huang, Tao Wang, Caiyu Fan, Xiansheng Cai, Sihan Hu, Xinzijian Liu, Cheng Shi, Mingjun Xu, Zhen Wang, Yan Wang, Xiangqi Jin, Tianhan Zhang, Linfeng Zhang, Lei Wang, Youjin Deng, Pan Zhang, Weijie Sun, Xinyu Li, Weinan E, Linfeng Zhang, Zhiyuan Yao, Kun Chen

机构 * Lanzhou Center for Theoretical Physics, Key Laboratory of Theoretical Physics of Gansu Province, Key Laboratory of Quantum Theory and Applications of MoE, Gansu Provincial Research Center for Basic Disciplines of Quantum Physics(兰州理论物理中心、甘肃省理论物理重点实验室、教育部量子理论与应用重点实验室、甘肃省量子物理基础学科研究省重点中心) Institute of Theoretical Physics, Chinese Academy of Sciences(中国科学院理论物理研究所) DP Technology(DP技术) Institute of Physics, Chinese Academy of Sciences(中国科学院物理研究所) Département d’Informatique, École normale supérieure(巴黎高等师范大学计算机系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于长推理链知识库的反向知识搜索方法,通过生成可验证的科学百科全书,实现了跨领域科学合成。

Comments 43 pages, 4 figures. This work is part of the SciencePedia project (sciencepedia.bohrium.com). Corrected author name spelling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08211 2026-01-21 cs.CL cs.AI cs.CR 62%

LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions

大语言模型无意中欺骗:从不一致样本到有偏的人机交互中的涌现不一致

Xuhao Hu, Peng Wang, Xiaoya Lu, Dongrui Liu, Xuanjing Huang, Jing Shao

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 研究发现大语言模型在高风险场景中可能因不一致样本而无意产生不诚实行为,且在下游任务和人机交互中风险显著增加。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06243 2026-01-21 cs.CL cs.AI 62%

CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning

CoT Referring: 通过 grounded 推理改进指称表达任务

Qihua Dong, Luis Figueroa, Handong Zhao, Kushal Kafle, Jason Kuen, Zhihong Ding, Scott Cohen, Yun Fu

机构 * Adobe Research(Adobe研究院) Northeastern University(东北大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 通过 grounded 推理改进指称表达任务,提出CoT Referring方法,提升多模态大语言模型在复杂指称场景中的性能。

Comments MLLM, Referring Expression Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23816 2026-01-21 cs.CL cs.LG 62%

A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs

在可转向性评估中的课程修正:揭示LLMs中的误校准与副作用

Trenton Chang, Tobias Schnabel, Adith Swaminathan, Jenna Wiens

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 本文提出了一种多维目标空间框架,揭示LLMs在文本改写任务中存在意外副作用,表明现有对齐策略可能不足。

Comments 8 pages, 6 figures. 26 pages of references and supplementary material, 22 additional figures. Association for the Advancement of Artificial Intelligence Conference (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13882 2026-01-21 cs.CL 57%

OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models

OpenLearnLM基准:一个评估教育大型语言模型知识、技能和态度的统一框架

Unggi Lee, Sookbun Lee, Heungsoo Choi, Jinseo Lee, Haeun Park, Younghoon Jeon, Sungmin Cho, Minju Kang, Junbo Koh, Jiyeong Bae, Minwoo Nam, Juyeon Eun, Yeonji Jung, Yeil Jeong

机构 * Chosun University(chosun大学) Korea University(韩国大学) Ewha Womans University(成均馆大学) Korea Institute for Curriculum and Evaluation(韩国课程评价院) Seoul National University(首尔国立大学) Texas A&M University(德克萨斯农工大学) Indiana University Bloomington(印第安纳大学布卢明顿分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 OpenLearnLM基准通过统一框架评估教育大型语言模型的知识、技能和态度,揭示不同模型在多维度上的能力差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13481 2026-01-21 cs.AI 57%

Towards Efficient and Robust Linguistic Emotion Diagnosis for Mental Health via Multi-Agent Instruction Refinement

通过多智能体指令细化实现高效的稳健语言情绪诊断以促进心理健康

Jian Zhang, Zhangqi Wang, Zhiyuan Wang, Weiping Fu, Yu He, Haiping Zhu, Qika Lin, Jun Liu

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) Saw Swee Hock School of Public Health, National University of Singapore(Saw Swee Hock 公共卫生学院,新加坡国立大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 APOLO通过多智能体协作优化提示空间,提升语言情绪诊断的效率与鲁棒性,适用于心理健康领域的可信LLM应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13443 2026-01-21 cs.AI 57%

Explicit Cognitive Allocation: A Principle for Governed and Auditable Inference in Large Language Models

显式认知分配:一种用于受控和可审计的大语言模型推理的原则

Héctor Manuel Manzanilla-Granados, Zaira Navarrete-Cazales, Miriam Pescador-Rojas, Tonahtiu Ramírez-Romero

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出显式认知分配原则,通过认知通用代理架构实现结构化推理,提升大语言模型在高责任场景下的可追溯性和可审计性。

Comments Preprint. This version corresponds to the initial public release of the CUA architecture and associated evaluation metrics

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08710 2026-01-21 cs.CL 57%

Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning

深入思考,而非总是更聪明:评估大语言模型在分层法律推理中的能力

Li Zhang, Matthias Grabmair, Morgan Gray, Kevin Ashley

机构 * University of Pittsburgh(匹兹堡大学) Technical University of Munich(慕尼黑技术大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

AI总结 本文提出一个框架评估大语言模型在分层法律推理中的能力,发现模型在表面推理准确但分层推理表现下降,揭示了模型在复杂领域中的局限性。

Comments 15 pages, 7 figures, Proceedings of the 2026 Symposium on Computer Science and Law (CSLAW '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10795 2026-01-21 cs.CL 57%

Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback

超越‘不够新颖’:利用LLM辅助反馈丰富学术批评

Osama Mohammed Afzal, Preslav Nakov, Tom Hope, Iryna Gurevych

机构 * UKP Lab, TU Darmstadt and Hessian Center for AI (hessian.AI)(图林人工智能实验室、德累斯顿理工大学和黑森人工智能中心) MBZUAI(马克斯·普朗克人工智能研究所) The Allen Institute for AI (AI2)(人工智能研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 本文提出一种结构化LLM辅助方法,通过三个阶段评估论文新颖性,实现86.5%的人类推理一致性,显著优于现有基线方法,提升同行评审的严谨性和透明度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23292 2026-01-21 cs.CV cs.AI 57%

Federated Unsupervised Semantic Segmentation

联邦无监督语义分割

Evangelos Charalampakis, Vasileios Mygdalis, Ioannis Pitas

机构 * Department of Informatics, Aristotle University of Thessaloniki(信息学院,阿基米德大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 本文提出FUSS框架,通过联邦学习实现去中心化无监督语义分割,优于传统方法和本地训练。

Comments Accepted for publication in Neurocomputing

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13162 2026-01-21 cs.LG cs.ET 57%

NeuroShield: A Neuro-Symbolic Framework for Adversarial Robustness

NeuroShield:一种用于对抗鲁棒性的神经符号框架

Ali Shafiee Sarvestani, Jason Schmidt, Arman Roohi

机构 * University of Illinois Chicago(伊利诺伊大学香槟分校) Department of Electrical and Computer Engineering(电气与计算机工程系) Department of Computer Science(计算机科学系)

专题命中 安全评测 :safety(abstract);分类 cs.LG

AI总结 NeuroShield通过整合符号规则监督提升深度神经网络的对抗鲁棒性和可解释性,其方法在对抗攻击测试中表现出显著优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12974 2026-01-21 cs.CL 57%

Bridging the Knowledge-Action Gap by Evaluating LLMs in Dynamic Dental Clinical Scenarios

通过评估LLMs在动态牙科临床场景中弥合知识-行动鸿沟

Hongyang Ma, Tiantian Gu, Huaiyuan Sun, Huilin Zhu, Yongxin Wang, Jie Li, Wubin Sun, Zeliang Lian, Yinghong Zhou, Yi Gao, Shirui Wang, Zhihui Tang

专题命中 安全评测 :safety(abstract);分类 cs.CL

AI总结 本文通过SCMPE基准评估LLMs在动态牙科临床场景中的表现,发现其在动态对话中存在主动信息收集和状态跟踪的瓶颈,揭示了外部知识不足以弥补推理差距,需领域自适应预训练。

Comments 29 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12661 2026-01-21 cs.AI 57%

MedConsultBench: A Full-Cycle, Fine-Grained, Process-Aware Benchmark for Medical Consultation Agents

MedConsultBench: 一个完整周期、细粒度、过程感知的医疗咨询代理基准

Chuhan Qiao, Jianghua Huang, Daxing Zhao, Ziding Liu, Yanjun Shen, Bing Cheng, Wei Lin, Kai Wu

机构 * Meituan(美团)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 MedConsultBench通过细粒度过程感知评估医疗咨询代理的完整咨询周期,揭示高诊断准确性背后的效率与安全问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12560 2026-01-21 cs.AI cs.MA 57%

Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents

代理型人工智能(AI):架构、分类及大语言模型代理的评估

Arunkumar V, Gangadharan G. R., Rajkumar Buyya

机构 * University College of Engineering, Anna University(安娜大学工程学院) National Institute of Technology Tiruchirappalli(Tiruchirappalli 国家理工学院) School of Computing and Information Systems University of Melbourne(墨尔本大学计算机与信息系统学院)

专题命中 安全评测 :prompt injection(abstract);分类 cs.AI

AI总结 本文探讨了代理型人工智能的架构、分类及评估,提出统一分类框架,分析从线性推理到原生推理模型的转变,并指出未来研究方向以提升自主系统的可靠性。

Comments 28 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12440 2026-01-21 cs.SD cs.AI cs.CV cs.DL eess.AS 57%

Style-based Composer Identification and Attribution of Symbolic Music Scores: a Systematic Survey

基于风格的作曲者识别与符号音乐谱作者归属:系统性综述

Federico Simonetta

机构 * GSSI – Gran Sasso Science Institute(格拉萨索科学研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 本文通过系统综述,探讨了符号音乐谱中基于风格的作曲者识别与作者归属问题,提出改进验证方法和模型的指导原则,以提高研究的可靠性与可重复性。

Comments Accepted at the TISMIR

Journal ref 2025 Transactions of the International Society for Music Information Retrieval 8 p. 213-235

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18056 2026-01-21 cs.CL 57%

MathEDU: Feedback Generation on Problem-Solving Processes for Mathematical Learning Support

MathEDU: 为数学学习支持生成问题解决过程的反馈

Wei-Ling Hsu, Yu-Chien Tang, An-Zi Yen

机构 * Department of Computer Science, National Yang Ming Chiao Tung University, Taiwan(计算机科学系,National Yang Ming Chiao Tung大学,台湾)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

AI总结 MathEDU研究通过构建数学问题解决过程反馈数据集,评估不同模型在正确性分类、错误识别和反馈生成任务中的可靠性,发现生成反馈存在显著差距,需提升教学意识的AI反馈。

Comments Accepted by EACL 2026 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05080 2026-01-21 cs.CL 57%

An Architectural Advantage of The Instruction-Tuned LLM in Containing The Readability-Accuracy Tension in Text Simplification

指令调优LLM在文本简化中缓解可读性-准确性张力的优势

P. Bilha Githinji, Aikaterini Meilliou, Zeming Liang, Lian Zhang, Peiwu Qin

专题命中 安全评测 :safety(abstract);分类 cs.CL

AI总结 本研究通过比较指令调优和推理增强的LLM,发现指令调优在文本简化中能更有效平衡可读性与准确性,提升话语忠实度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23045 2026-01-21 cs.AI 57%

A Survey of AI Scientists

AI科学家的综述

Guiyao Tie, Pan Zhou, Lichao Sun

机构 * Huazhong University of Science and Technology(华中科技大学) Lehigh University(莱斯大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 本文综述了AI科学家的发展历程,提出六阶段方法论框架,分析了从基础模块到闭环系统再到可扩展性与人机协作的演进,为未来系统发展提供路线图。

Comments 28 pages, 9 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00332 2026-01-21 cs.AI cs.CE 57%

When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets

当幻觉成本百万:在高风险对抗性金融市场中基准测试AI代理

Zeshi Dai, Zimo Peng, Zerui Cheng, Ryan Yihe Li

机构 * Surf AI, Cybertino Lab(Surf AI,Cybertino 实验室) Princeton University(普林斯顿大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 CAIA基准测试揭示了AI在对抗性金融市场中的能力缺口,指出当前模型在面对虚假信息和不可逆决策时表现不佳,强调对抗鲁棒性对可信AI的重要性。

Comments 15 pages, 5 figures, 4 tables; Accepted to AAAI 2026 (AI-4-Finance Workshop - Oral, top 10%); In submission to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24294 2026-01-21 cs.CL cs.HC 57%

LOGOS: LLM-driven End-to-End Grounded Theory Development and Schema Induction for Qualitative Research

LOGOS: 基于大语言模型的端到端 grounded theory 开发与模式诱导用于定性研究

Xinyu Pi, Qisen Yang, Chuong Nguyen

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 LOGOS 是一种基于大语言模型的端到端框架,通过自动化 grounded theory 工作流程,实现定性研究的结构化理论开发和模式诱导,显著提升了研究的可扩展性和理论精确度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21046 2026-01-21 cs.AI 57%

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

自我进化代理的综述:何时、何地、如何进化以实现人工超级智能

Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, Hongru Wang, Han Xiao, Yuhang Zhou, Shaokun Zhang, Jiayi Zhang, Jinyu Xiang, Yixiong Fang, Qiwen Zhao, Dongrui Liu, Qihan Ren, Cheng Qian, Zhenhailong Wang, Minda Hu, Huazheng Wang, Qingyun Wu, Heng Ji, Mengdi Wang

机构 * Princeton University(普林斯顿大学) Princeton AI Lab(普林斯顿人工智能实验室) Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学) University of Sydney(悉尼大学) Shanghai Jiao Tong University(上海交通大学) Pennsylvania State University(宾夕法尼亚州立大学) University of Michigan(密歇根大学) Oregon State University(俄勒冈州立大学) The Chinese University of Hong Kong(香港中文大学) Fudan University(复旦大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The University of Hong Kong(香港大学) University of California, Santa Barbara(加州大学圣芭芭拉分校) University of California San Diego(加州大学圣地亚哥分校) University of Edinburgh(爱丁堡大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 本文综述了自我进化代理的现状,探讨了进化机制、适应方法及挑战,为实现人工超级智能提供路线图。

Comments 77 pages, 9 figures, Transactions on Machine Learning Research (01/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09660 2026-01-21 astro-ph.CO 50%

Measuring the Dark Matter Self-Interaction Cross-Section with Deep Compact Clustering for Robust Machine Learning Inference

利用深度紧致聚类测量暗物质自相互作用截面以实现稳健的机器学习推断

Ethan Tregidga, David Harvey, Luca Biggio, Felix Vecchi

专题命中 安全评测 :trustworthy(abstract)

AI总结 该研究提出利用深度紧致聚类技术,通过分析星系团图像来测量暗物质自相互作用截面,从而实现稳健的机器学习推断。

Comments 11 pages, 7 figures, submitted to A&A

Journal ref A&A 705, A152 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14163 2026-01-21 cs.SE cs.CR 50%

An Empirical Study on Remote Code Execution in Machine Learning Model Hosting Ecosystems

关于机器学习模型托管生态系统中远程代码执行的实证研究

Mohammed Latif Siddiq, Tanzim Hossain Romel, Natalie Sekerak, Beatrice Casey, Joanna C. S. Santos

专题命中 安全评测 :safety(abstract)

AI总结 本研究通过实证分析揭示了机器学习模型托管平台中远程代码执行的安全风险及开发者认知问题,提出安全与可用性平衡的改进方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13933 2026-01-21 cs.SE 50%

VulnResolver: A Hybrid Agent Framework for LLM-Based Automated Vulnerability Issue Resolution

VulnResolver: 一种基于大语言模型的混合代理框架用于LLM驱动的自动漏洞问题解决

Mingming Zhang, Xu Wang, Jian Zhang, Xiangxin Meng, Jiayi Zhang, Chunming Hu

专题命中 安全评测 :safety(abstract)

AI总结 VulnResolver是一种基于大语言模型的混合代理框架,通过上下文预收集和安全属性分析代理,实现自动化漏洞问题的高效解决。

详情

展开后加载摘要…

URL PDF HTML 收藏