arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15756 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15756 篇

2602.02567 2026-02-04 cs.LG cs.AI eess.IV 62%

IceBench-S2S: A Benchmark of Deep Learning for Challenging Subseasonal-to-Seasonal Daily Arctic Sea Ice Forecasting in Deep Latent Space

IceBench-S2S:一种深度学习用于挑战性亚季节至季节每日北极海冰预测的基准

Jingyi Xu, Shengnan Wang, Weidong Yang, Siwei Tu, Lei Bai, Ben Fei

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) Chinese University of Hong Kong(香港中文大学)

专题命中 Agent评测 :planning(abstract);分类 cs.AI、cs.LG

AI总结 IceBench-S2S通过深度学习框架提升北极海冰预测的季节尺度能力,为极地环境监测提供统一的训练和评估流程。

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01995 2026-02-04 cs.AI cs.CL 62%

Thinking Like a Doctor: Conversational Diagnosis through the Exploration of Diagnostic Knowledge Graphs

像医生一样思考:通过探索诊断知识图谱进行对话诊断

Jeongmoon Won, Seungwon Kook, Yohan Jo

机构 * Graduate School of Data Science, Seoul National University(数据科学研究生院,首尔国立大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

AI总结 本文提出一种通过探索诊断知识图谱进行对话诊断的系统,通过两步推理生成和验证诊断假设,提高诊断准确性和效率,并通过医生评估验证其临床实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00408 2026-02-04 cs.LG cs.AI 62%

Variational Approach for Job Shop Scheduling

变分方法用于作业车间调度

Seung Heon Oh, Jiwon Baek, Ki Young Cho, Hee Chang Yoon, Jong Hun Woo

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出变分图到调度器框架,通过解耦表示学习与策略优化,提升作业车间调度问题的训练稳定性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19120 2026-02-04 cs.IR cs.AI cs.LG 62%

RobustExplain: Evaluating Robustness of LLM-Based Explanation Agents for Recommendation

RobustExplain: 评估基于LLM的推荐解释代理的鲁棒性

Guilin Zhang, Kai Zhao, Jeffrey Friedman, Xu Chu

机构 * Workday

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 RobustExplain提出首个评估框架,用于衡量LLM生成推荐解释的鲁棒性,揭示当前模型鲁棒性较低,较大模型稳定性更高。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10304 2026-02-03 cs.LG cs.AI cs.GT 62%

Large-Scale Auto-bidding with Nash Equilibrium Constraints

大规模自动出价与纳什均衡约束

Zhiyu Mou, Miao Xu, Rongquan Bai, Zhuoran Yang, Chuan Yu, Jian Xu, Bo Zheng

机构 * Alibaba Group(阿里巴巴集团) Department of Statistics and Data Science(统计与数据科学系) Yale University(耶鲁大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出纳什均衡约束出价框架,通过理论严谨的惩罚对偶梯度方法实现大规模自动出价的稳定性和最优性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01367 2026-02-03 cs.LG cs.AI 62%

Deep Variational Contrastive Learning for Joint Risk Stratification and Time-to-Event Estimation

深度变分对比学习用于联合风险分层和生存时间估计

Pinar Erbil, Alberto Archetti, Eugenio Lomurno, Matteo Matteucci

机构 * Politecnico di Milano, AIRLab(米兰理工学院,AIRLab)

专题命中 Agent评测 :planning(abstract);分类 cs.AI、cs.LG

AI总结 CONVERSE通过结合变分自编码器和对比学习,实现可解释的风险分层和生存时间估计,提升深度生存模型的性能与可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16584 2026-02-03 cs.CL cs.AI 62%

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations

从分数到步骤:诊断和改进证据医学计算中LLM的性能

Benlu Wang, Iris Xia, Yifan Zhang, Junda Wang, Feiyun Ouyang, Shuo Han, Arman Cohan, Hong Yu, Zonghai Yao

机构 * Department of Computer Science, Yale University, CT, USA(耶鲁大学计算机科学系) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗中心健康组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(UMass洛厄尔矿尔计算机与信息科学学院) Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(UMass阿默斯特马宁信息与计算机科学学院)

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.CL

AI总结 本文提出MedRaC框架,通过分步评估和代码执行提升LLM在证据医学计算中的准确性,揭示现有评估方法的不足,并推动临床可信度的提升。

Comments Equal contribution for the first two authors. To appear as an Oral presentation in the proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09751 2026-02-03 cs.LG cs.AI 62%

Meta-Learning Reinforcement Learning for Crypto-Return Prediction

元学习强化学习用于加密货币收益预测

Junqiao Wang, Zhaoyang Guan, Guanyu Liu, Tianze Xia, Xianzhi Li, Shuo Yin, Xinyuan Song, Chuhan Cheng, Tianyu Shi, Alex Lee

机构 * Sichuan University(四川大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Northwestern University(西北大学) Huazhong University of Science and Technology(华中科技大学) Queen’s University(女王大学) University of Toronto(多伦多大学) TrueNorth Tsinghua University(清华大学) Emory University(埃默里大学) University of Macau(澳门大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 Meta-RL-Crypto通过结合元学习和强化学习,构建了一个自我改进的交易代理,有效提升加密货币收益预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20886 2026-02-02 cs.SE cs.LG 62%

IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks

IDE-Bench:通过IDE原生工具接口评估大型语言模型作为IDE代理在真实软件工程任务中的表现

Spencer Mateega, Jeff Yang, Tiana Costello, Shaurya Jadhav, Nicole Tian, Agustin Garcinuño

专题命中 Agent评测 :agent(abstract);分类 cs.LG、cs.SE

AI总结 IDE-Bench通过IDE原生工具接口评估大型语言模型作为IDE代理在真实软件工程任务中的表现,提供首个系统性的多语言全栈环境评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20905 2026-01-30 eess.IV cs.AI cs.CV cs.LG eess.SP 62%

Denoising and Baseline Correction of Low-Scan FTIR Spectra: A Benchmark of Deep Learning Models Against Traditional Signal Processing

低扫描傅里叶变换红外光谱的去噪与基线校正:深度学习模型与传统信号处理的基准测试

Azadeh Mokari, Shravan Raghunathan, Artem Shydliukh, Oleg Ryabchykov, Christoph Krafft, Thomas Bocklitz

专题命中 Agent评测 :workflow(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种物理引导的级联Unet模型,通过分离去噪和基线校正任务,有效提升FTIR成像速度和精度,实现51.3%的RMSE降低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20714 2026-01-29 cs.LG cs.AI 62%

Adapting the Behavior of Reinforcement Learning Agents to Changing Action Spaces and Reward Functions

适应强化学习智能体的行为以应对变化的动作空间和奖励函数

Raul de la Rosa, Ivana Dusparic, Nicolas Cardozo

机构 * Trinity College Dublin(都柏林三一学院)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 MORPHIN通过动态调整超参数和概念漂移检测,实现强化学习智能体在变化的奖励函数和动作空间中的自适应学习,提升收敛速度和学习效率。

Journal ref 2025 IEEE International Conference on Autonomic Computing and Self-Organizing Systems Companion (ACSOS-C), Tokyo, Japan, 2025, pp. 148-153

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00282 2026-01-29 cs.AI cs.CL 62%

Mind the Gap: The Divergence Between Human and LLM-Generated Tasks

注意差距:人类与LLM生成任务之间的差异

Yi-Long Lu, Jiajun Song, Chunhui Zhang, Wei Wang

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

AI总结 研究揭示了人类与LLM在任务生成中的核心差异,指出人类受心理驱动影响,而LLM在生成具身目标方面存在不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19738 2026-01-28 quant-ph cs.AI cs.LG 62%

Quantum Circuit Pre-Synthesis: Learning Local Edits to Reduce $T$-count

量子电路预合成:学习局部编辑以减少T门数量

Daniele Lizzio Bosco, Lukasz Cincio, Giuseppe Serra, M. Cerezo

机构 * Department of Mathematics, Computer Science Physics, University of Udine, Udine, Italy Department of Biology, University of Naples Federico II, Naples, Italy Theoretical Division, Los Alamos National Laboratory, Los Alamos, New Mexico 87545, USA Quantum Science Center, Oak Ridge, TN 37931, USA Information Sciences, Los Alamos National Laboratory, Los Alamos, New Mexico 87545, USA

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 Q-PreSyn通过强化学习优化局部编辑,有效减少量子电路的T门数量,提升编译效率

Comments 10+5 pages, 10 figures, 3 algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18827 2026-01-28 cs.SE cs.AI 62%

Automated structural testing of LLM-based agents: methods, framework, and case studies

基于大语言模型的智能体的自动化结构测试:方法、框架与案例研究

Jens Kohl, Otto Kruse, Youssef Mostafa, Andre Luckow, Karsten Schroer, Thomas Riedl, Ryan French, David Katz, Manuel P. Luitz, Tanrajbir Takher, Ken E. Friedl, Céline Laurent-Winter

机构 * Amazon Web Services(亚马逊网络服务)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.SE

AI总结 本文提出基于大语言模型的智能体结构测试方法,通过跟踪、模拟和断言实现自动化测试,提升测试效率与质量。

Comments 10 pages, 5 figures. Preprint of an accepted paper at IEEE BigData 2025 (main track). Source code for the introduced methods and framework available at https://github.com/awslabs/generative-ai-toolkit

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17588 2026-01-27 cs.AI cs.CL 62%

Intelligence Requires Grounding But Not Embodiment

智能需要具身但不需要身体

Marcus Ma, Shrikanth Narayanan

机构 * University of Southern California(南加州大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

AI总结 本文提出智能需要基础性而非具身,通过定义智能的四个属性并论证非具身智能体可实现这些属性,从而得出结论:基础性是智能的必要条件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09715 2026-01-27 cs.CL cs.AI cs.HC cs.IR 62%

Introducing Axlerod: An LLM-based Chatbot for Assisting Independent Insurance Agents

引入Axlerod:一种基于大语言模型的聊天机器人,用于协助独立保险代理

Adam Bradley, John Hastings, Khandaker Mamun Ahmed

机构 * The Beacom College of Computer and Cyber Sciences, Dakota State University(计算机与网络安全学院,达科他州立大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

AI总结 Axlerod是一种基于大语言模型的聊天机器人,旨在通过自然语言处理、检索增强生成和领域知识整合,提高独立保险代理的工作效率。

Comments 6 pages, 2 figures, 1 table

Journal ref 2025 IEEE Cyber Awareness and Research Symposium (CARS'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10994 2026-01-27 cs.CL cs.AI 62%

Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety

基于开放领域评估和多阶段防护机制的深度研究

Wei-Chieh Huang, Henry Peng Zou, Yaozu Wu, Dongyuan Li, Yankai Chen, Weizhi Zhang, Yangning Li, Angelo Zangari, Jizhou Guo, Chunyu Miao, Liancheng Fang, Langzhou He, Yinghui Li, Renhe Jiang, Philip S. Yu

专题命中 Agent评测 :planning(abstract);分类 cs.AI、cs.CL

AI总结 本文提出DeepResearchGuard和DRSafeBench,通过开放领域评估和多阶段防护机制提升深度研究的安全性和报告质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15487 2026-01-23 cs.AI cs.CL cs.MA 62%

MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation

MiRAGE:一种多智能体框架,用于生成多模态多跳问题-答案数据集以评估RAG系统

Chandan Kumar Sahu, Premith Kumar Chilukuri, Matthew Hetrich

机构 * ABB Inc(ABB公司)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

AI总结 MiRAGE通过多智能体框架生成多模态多跳问题-答案数据集,提升RAG系统评估的准确性和复杂性。

Comments 12 pages, 2 figures, Submitted to ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12471 2026-01-23 cs.CL cs.AI 62%

Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

知何时退避:医疗大语言模型在临床不确定性中的表现

Sravanthi Machcha, Sushrita Yerra, Sahil Gupta, Aishwarya Sahoo, Sharmin Sultana, Hong Yu, Zonghai Yao

机构 * Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(马萨诸塞大学阿姆赫斯特曼宁信息与计算机科学学院) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(医疗组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(米纳尔计算机与信息科学学院)

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.CL

AI总结 本文提出MedAbstain基准,探讨医疗LLM在临床不确定性中的退避能力,发现显式退避选项能显著提升安全性,而模型规模和提示方法效果有限。

Comments Equal contribution for the first two authors; To appear in proceedings of the Main Conference of the European Chapter of the Association for Computational Linguistics (EACL) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14235 2026-01-22 astro-ph.IM astro-ph.CO cs.AI cs.LG stat.ML 62%

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

人工智能/机器学习在Rubin LSST暗能量科学合作中的机遇

LSST Dark Energy Science Collaboration, Eric Aubourg, Camille Avestruz, Matthew R. Becker, Biswajit Biswas, Rahul Biswas, Boris Bolliet, Adam S. Bolton, Clecio R. Bom, Raphaël Bonnet-Guerrini, Alexandre Boucaud, Jean-Eric Campagne, Chihway Chang, Aleksandra Ćiprijanović, Johann Cohen-Tanugi, Michael W. Coughlin, John Franklin Crenshaw, Juan C. Cuevas-Tello, Juan de Vicente, Seth W. Digel, Steven Dillmann, Mariano Javier de León Dominguez Romero, Alex Drlica-Wagner, Sydney Erickson, Alexander T. Gagliano, Christos Georgiou, Aritra Ghosh, Matthew Grayling, Kirill A. Grishin, Alan Heavens, Lindsay R. House, Mustapha Ishak, Wassim Kabalan, Arun Kannawadi, François Lanusse, C. Danielle Leonard, Pierre-François Léget, Michelle Lochner, Yao-Yuan Mao, Peter Melchior, Grant Merz, Martin Millon, Anais Möller, Gautham Narayan, Yuuki Omori, Hiranya Peiris, Laurence Perreault-Levasseur, Andrés A. Plazas Malagón, Nesar Ramachandra, Benjamin Remy, Cécile Roucelle, Jaime Ruiz-Zapatero, Stefan Schuldt, Ignacio Sevilla-Noarbe, Ved G. Shah, Tjitske Starkenburg, Stephen Thorp, Laura Toribio San Cipriano, Tilman Tröster, Roberto Trotta, Padma Venkatraman, Amanda Wasserman, Tim White, Justine Zeghal, Tianqing Zhang, Yuanyuan Zhang

机构 * Université Paris Cité, CNRS, CEA, Astroparticule et Cosmologie, F-75013 Paris, France Department of Physics, University of Michigan, Ann Arbor, MI 48109, USA Leinweber Institute of Theoretical Physics, University of Michigan, Ann Arbor, MI 48109, USA Argonne National Laboratory, 9700 South Cass Avenue, Lemont, IL 60439, USA Cavendish Astrophysics, University of Cambridge, Madingley Road, Cambridge CB3 0HA, UK Kavli Institute for Cosmology, University of Cambridge, Madingley Road, Cambridge CB3 0HA, UK SLAC National Accelerator Laboratory, Menlo Park, CA 94025, USA Department of Computer Science, University of Milan, Milan, Italy Université Paris Cité, CNRS, Astroparticule et Cosmologie, F-75013 Paris, France Université Paris-Saclay, CNRS/IN2P3, IJCLab, 91405 Orsay, France Department of Astronomy Astrophysics, University of Chicago, Chicago, IL 60637, USA Kavli Institute for Cosmological Physics, University of Chicago, Chicago, IL 60637, USA NSF-Simons AI Institute for the Sky (SkAI), 172 E. Chestnut St., Chicago, IL 60611, USA Fermi National Accelerator Laboratory, P.O. Box 500, Batavia, IL 60510, USA Universit\'e Clermont-Auvergne, CNRS, LPCA, 63000 Clermont-Ferrand, France Kavli Institute for Particle Astrophysics Cosmology, Stanford University, Stanford, CA 94305, USA Department of Physics, Stanford University, 382 Via Pueblo Mall, Stanford, CA 94305, USA Engineering Faculty, Universidad Autonoma de San Luis Potosi, Zona Universitaria, San Luis Potosi, 78290, Mexico Stanford Artificial Intelligence Laboratory, Stanford University, Stanford, CA 94305, USA Kavli Institute of Cosmological Physics, University of Chicago, Chicago, IL 60637, USA The NSF AI Institute for Artificial Intelligence Center for Astrophysics Harvard \& Smithsonian, 60 Garden Street, Cambridge, MA 02138, USA Department of Physics Kavli Institute for Astrophysics Space Research, Massachusetts Institute of Technology, Cambridge, MA 02139, USA Institut de Física d'Altes Energies (IFAE), The Barcelona Institute of Science Institute of Astronomy Kavli Institute for Cosmology, University of Cambridge, Madingley Road, Cambridge, CB3 0HA, UK Imperial Centre for Inference Cosmology (ICIC), Imperial College London, Blackett Laboratory, Prince Consort Road, London SW7 2AZ, UK Data Science Institute, The University of Chicago, Chicago, IL 60615, USA Department of Physics, The University of Texas at Dallas, Richardson, TX 75080, USA Department of Physics, Duke University, Durham, NC 27708, USA Université Paris-Saclay, Université Paris Cité, CEA, CNRS, AIM, F-91191 Gif-sur-Yvette, France School of Mathematics, Statistics Physics, Newcastle University, Newcastle upon Tyne, NE1 7RU, United Kingdom Department of Astrophysical Sciences, Princeton University, Princeton, NJ 08544, USA Astronomy, University of the Western Cape, Bellville, Cape Town, 7535, South Africa Astronomy, University of Utah, Salt Lake City, UT 84112, USA Department of Astrophysical Sciences, Princeton University, Peyton Hall, Princeton, NJ 08544, USA Department of Astronomy, University of Illinois Urbana Champaign, 1002 W. Green St., Urbana, IL, 61801, USA Institute for Particle Physics Astrophysics, ETH Zürich, Wolfgang-Pauli-Strasse 27, CH-8093 Zurich, Switzerland Swinburne University of Technology, Hawthorn, Victoria 3122, Australia Ciela - Montr\'eal Institute for Astrophysical Data Analysis Mila - Quebec Artificial Intelligence Institute, Montréal, QC H2S 3H1, Canada Advanced Research Computing Centre, University College London, 90 High Holborn, London WC1V 6LJ, UK Finnish Centre for Astronomy with ESO (FINCA), University of Turku, FI-20014 Turku, Finland Department of Physics, P.O. Box 64, University of Helsinki, FI-00014 Helsinki, Finland Astronomy, Northwestern University, Evanston, IL, USA Center for Interdisciplinary Exploration Research in Astrophysics, Northwestern University, Evanston, IL, USA Scientific Data Science, International School for Advanced Study, Via Bonomea 265, I-34136 Trieste, Italy Department of Statistics, University of Michigan, Ann Arbor, MI 48109, USA PITT PACC, University of Pittsburgh, Pittsburgh, PA 15260, USA NSF NOIRLab, 950 N. Cherry Ave., Tucson, AZ 85719, USA

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了AI/ML在LSST暗能量科学合作中的应用机遇,强调了大规模贝叶斯推断、物理指导方法和主动学习等关键方法学优先事项,并讨论了新兴技术在重塑工作流程中的潜力。

Comments 84 pages. This is v1.0 of the DESC's white paper on AI/ML, a collaboration document that is being made public but which is not planned for submission to a journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12787 2026-01-22 cs.LG cs.AI 62%

Impartial Games: A Challenge for Reinforcement Learning

impartial games: 一种对强化学习的挑战

Bei Zhou, Søren Riis

机构 * Imperial College London(帝国理工学院伦敦分校) Queen Mary University of London(女王玛丽大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了AlphaZero风格强化学习在impartial games中的局限性,指出其在学习抽象数学原理如奇偶性时存在表示瓶颈,需发展新型算法以实现专家级AI。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13217 2026-01-21 cs.CL cs.AI 62%

Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision

超越单次写作:深度研究代理在多轮报告修订中不可靠

Bingsen Chen, Boyan Li, Ping Nie, Yuyu Zhang, Xi Ye, Chen Zhao

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

AI总结 本研究提出Mr Dre评估套件,揭示深度研究代理在多轮报告修订中存在内容退化和编辑丢失的问题,表明需更深入的方法改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13007 2026-01-21 cs.SE cs.AI 62%

ArchAgent: Scalable Legacy Software Architecture Recovery with LLMs

ArchAgent: 基于LLM的可扩展遗留软件架构恢复

Rusheng Pan, Bingcheng Mao, Tianyi Ma, Zhenhua Ling

机构 * HiThink Research(慧思研究) University of Science and Technology of China(中国科学技术大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.SE

AI总结 ArchAgent通过结合静态分析、自适应分段和LLM合成,实现大规模遗留软件架构的高效恢复与多视图业务对齐。

Comments to be published in ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12259 2026-01-21 cs.AI cs.CE cs.LG 62%

FutureX-Pro: Extending Future Prediction to High-Value Vertical Domains

FutureX-Pro: 将未来预测扩展到高价值垂直领域

Jiashuo Liu, Siyuan Chen, Zaiyuan Wang, Zhiyuan Zeng, Jiacheng Guo, Liang Hu, Lingyue Yin, Suozhi Huang, Wenxin Hao, Yang Yang, Zerui Cheng, Zixin Yao, Lingyue Yin, Haoxin Liu, Jiayi Cheng, Yuzhen Li, Zezhong Ma, Bingjie Wang, Bingsen Qiu, Xiao Liu, Zeyang Zhang, Zijian Liu, Jinpeng Wang, Mingren Yin, Tianci He, Yali Liao, Yixiao Tian, Zhenwei Zhu, Anqi Dai, Ge Zhang, Jingkai Liu, Kaiyuan Zhang, Wenlong Wu, Xiang Gao, Xinjie Chen, Zhixin Yao, Zhoufutu Wen, B. Aditya Prakash, Jose Blanchet, Mengdi Wang, Nian Si, Wenhao Huang

机构 * Hong Kong University of Science and Technology(香港科技大学) Georgia Institute of Technology(佐治亚理工学院) Stanford University(斯坦福大学) Princeton University(普林斯顿大学)

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.LG

AI总结 FutureX-Pro通过扩展未来预测到金融、零售、公共健康和自然灾害等高价值垂直领域,评估代理LLMs在工业部署中的领域基础能力。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07541 2026-01-19 cs.LG cs.AI 62%

A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement Learning

一个简单的统一不确定性引导框架用于离线到在线强化学习

Siyuan Guo, Yanchao Sun, Jifeng Hu, Sili Huang, Hechang Chen, Haiyin Piao, Lichao Sun, Yi Chang

机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE(知识驱动的人机智能工程研究中心,教育部) International Center of Future Science, Jilin University(未来科学国际中心,吉林大学) JPMorgan AI Research(摩根大通AI研究) Northwestern Polytechnical University(西北工业大学) Lehigh University(莱斯大学)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 SUNG框架通过不确定性引导解决离线到在线强化学习中的探索与分布偏移问题,实现高效微调和跨环境的高性能表现。

Comments The final published version is available at IEEE Xplore: https://ieeexplore.ieee.org/abstract/document/11267513/. We correct the GitHub repo url in this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09723 2026-01-16 cs.CL cs.AI 62%

SagaScale: A Realistic, Scalable, and High-Quality Long-Context Benchmark Built from Full-Length Novels

SagaScale: 一种真实、可扩展且高质量的长上下文基准,基于完整小说构建

Guancheng Du, Yong Hu, Wenqing Wang, Yaming Yang, Jiaheng Gao

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.CL

AI总结 SagaScale基于完整小说构建,提供长上下文基准,通过自动化数据收集管道,评估12个LLMs和三种方法,揭示LLM在长上下文处理中的表现差异及Agentic RAG的改进效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03055 2026-01-16 cs.LG cs.AI 62%

Permissive Information-Flow Analysis for Large Language Models

宽松的信息流分析用于大型语言模型

Shoaib Ahmed Siddiqui, Radhika Gaonkar, Boris Köpf, David Krueger, Andrew Paverd, Ahmed Salem, Shruti Tople, Lukas Wutschitz, Menglin Xia, Santiago Zanella-Béguelin

机构 * University of Cambridge(剑桥大学) Microsoft(微软公司) Mila

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种更宽松的信息流分析方法,通过传播对模型输出有影响的样本标签来提高大型语言模型的安全性和隐私保护效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09527 2026-01-15 cs.LG cs.AI cs.PF 62%

Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs

在消费级Blackwell GPU上进行私有LLM推理:为中小企业实现低成本本地部署的实用指南

Jonathan Knoop, Hendrik Holtmann

机构 * IE Business University(IE商学院)

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.LG

AI总结 本文评估了消费级Blackwell GPU在LLM推理中的性能,证明其在多数中小企业工作负载中可替代云服务,但长上下文RAG任务仍需高端GPU。

Comments 15 pages, 18 tables, 7 figures. Includes link to GitHub repository and Docker image for reproducibility

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08682 2026-01-14 cs.CL cs.AI 62%

Lessons from the Field: An Adaptable Lifecycle Approach to Applied Dialogue Summarization

领域经验:一种适用于应用对话摘要的可适应生命周期方法

Kushal Chawla, Chenyang Zhu, Pengshan Cai, Sangwoo Cho, Scott Novotney, Ayushman Singh, Jonah Lewis, Keasha Safewright, Alfy Samuel, Erin Babinsky, Shi-Xiong Zhang, Sambit Sahu

机构 * Capital One

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.CL

AI总结 本文提出了一种适用于应用对话摘要的可适应生命周期方法,通过行业案例研究探讨了在动态需求下构建可靠摘要系统的挑战与解决方案。

Comments EACL 2026 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08477 2026-01-14 cs.CL cs.HC cs.SE 62%

Do You Understand How I Feel?: Towards Verified Empathy in Therapy Chatbots

你了解我感受吗?:朝向疗法聊天机器人中的验证共情

Francesco Dettori, Matteo Forasassi, Lorenzo Veronese, Livia Lestingi, Vincenzo Scotti, Matteo Giovanni Rossi

机构 * Université Paris-Saclay, Centre National de la Recherche Scientifique(巴黎萨克雷大学,法国国家科学研究中心) TU Wien(维也纳技术大学) Politecnico di Milano(米兰理工学院) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 Agent评测 :agent(abstract);分类 cs.CL、cs.SE

AI总结 本文提出通过自然语言处理与形式验证结合的方法,开发具有共情能力的疗法聊天机器人,通过统计模型检查验证共情属性并指导行为策略。

详情

展开后加载摘要…

URL PDF HTML 收藏