arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-01-27 至 2026-01-27 共收录 172 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 多智能体 32 篇

2601.18573 2026-01-27 cs.GT cs.DS 67%

Stable Matching with Deviators and Conformists

带有偏离者和顺从者的稳定匹配

Frederik Glitzner, David Manlove

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 研究在存在偏离者和顺从者的情况下,如何高效决定是否存在无偏离者阻塞的稳定匹配,并探讨其计算复杂性。

Comments Preliminary version to appear at AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16643 2026-01-27 physics.soc-ph 67%

Evolutionary Dynamics of Reputation-Based Voluntary Prisoner's Dilemma Games

声誉基于的自愿囚徒困境博弈的演化动态

Chen Shen, Zhao Song, Xinyu Wang, Lei Shi, Matjaž Perc, Zhen Wang, Jun Tanimoto

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 本文提出基于声誉的自愿囚徒困境模型,通过分析退出激励机制在不同种群结构中的动态影响,揭示了合作维持的多种共存路径及稳定机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18218 2026-01-27 cs.HC cs.AI cs.CL 62%

PaperTok: Exploring the Use of Generative AI for Creating Short-form Videos for Research Communication

PaperTok: 探索生成式AI在科研传播中制作短视频的应用

Meziah Ruby Cristobal, Hyeonjeong Byeon, Tze-Yu Chen, Ruoxi Shang, Donghoon Shin, Ruican Zhong, Tony Zhou, Gary Hsieh

机构 * University of Washington(华盛顿大学)

专题命中 多智能体 :workflow(abstract);分类 cs.AI、cs.CL

AI总结 PaperTok利用生成式AI帮助研究人员将学术论文转化为短视频,通过自动化脚本和视听内容生成,提升科研传播效率。

Journal ref In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26), Apr 13-17, 2026, Barcelona, Spain. ACM, New York, NY, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10950 2026-01-27 cs.HC cs.CY 50%

Can GenAI Move from Individual Use to Collaborative Work? Experiences, Challenges, and Opportunities of Coordinating GenAI into Collaborative Newswork

生成式AI能否从个体使用转向协作工作?将生成式AI纳入协作新闻工作中的经验、挑战与机遇

Qing Xiao, Qing Hu, Jingjia Xiao, Hancheng Cao, Hong Shen

专题命中 多智能体 :workflow(abstract)

AI总结 研究探讨生成式AI在新闻工作中从个体使用向协作工作转型的挑战与机遇,发现价值观契合依赖个体自主性,组织层面存在结构性和文化性障碍。

Comments 22 pages, 1 figure, accepted by CHI'26

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 工作流自动化 18 篇

2601.17332 2026-01-27 cs.AI 88%

TheoremForge: Scaling up Formal Data Synthesis with Low-Budget Agentic Workflow

TheoremForge: 以低成本代理工作流扩展正式数据合成

Yicheng Tao, Hongteng Xu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学全球化人工智能学院) Beijing Key Laboratory of Research on Large Models(北京大模型研究关键实验室) Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索与推荐工程研究中心)

专题命中 工作流自动化 :workflow(title,abstract);agentic(title,abstract);分类 cs.AI

AI总结 TheoremForge通过低成本代理工作流提升形式化数据合成效率,实现12.6%的验证率和1.6倍的数据产量提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17413 2026-01-27 cs.SE 87%

When AI Agents Touch CI/CD Configurations: Frequency and Success

当AI代理触碰CI/CD配置时:频率和成功

Taher A. Ghaleb

专题命中 工作流自动化 :AI agent(title,abstract);agent(abstract);workflow(abstract);agentic(abstract)

AI总结 研究发现AI代理在CI/CD配置中主要聚焦GitHub Actions,其配置更改可靠性与常规代码相当,Copilot在CI/CD性能上表现突出。

Comments Accepted at the 23rd International Conference on Mining Software Repositories (MSR '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14308 2026-01-27 cs.HC 86%

ReUseIt: Synthesizing Reusable AI Agent Workflows for Web Automation

ReUseIt:合成可重用的AI代理工作流以实现网页自动化

Yimeng Liu, Misha Sra, Jeevana Priya Inala, Chenglong Wang

专题命中 工作流自动化 :agent(title,abstract);AI agent(title)

AI总结 ReUseIt通过自动合成可重用的工作流,使AI代理在网页自动化任务中实现更高的成功率和更少的用户干预。

Comments ACM IUI '26 | 31st International Conference on Intelligent User Interfaces

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18239 2026-01-27 cs.HC 85%

Probing the Future of Meta-Analysis: Eliciting Design Principles via an Agentic Research IDE

探索元分析的未来:通过代理研究IDE eliciting设计原则

Sizhe Cheng, Feng Liang, Yuhan Wen, Xipei Yu, Yong Wang

专题命中 工作流自动化 :agentic(title);agent(abstract);workflow(abstract);multi-agent(abstract)

AI总结 Research IDE通过'研究作为代码'隐喻,提供一种多代理后端支持的写作环境,通过'假设断点'实现现场验证,提升研究人员的自主权和智力所有权。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18517 2026-01-27 cs.CL 70%

GenAI for Social Work Field Education: Client Simulation with Real-Time Feedback

生成AI用于社会工作领域教育:带有实时反馈的客户模拟

James Sungarda, Hongkai Liu, Zilong Zhou, Tien-Hsuan Wu, Johnson Chun-Sing Cheung, Ben Kao

机构 * School of Computing(计算学院) Data Science The University of Hong Kong Hong Kong, China(数据科学香港大学香港) Dept. of Mathematics The University of Hong Kong Hong Kong, China(数学系香港大学香港) Faculty of Engineering The University of Hong Kong Hong Kong, China(工程学院香港大学香港) Social Administration The University of Hong Kong Hong Kong, China(社会行政香港大学香港)

专题命中 工作流自动化 :agent(abstract);workflow(abstract);分类 cs.CL

AI总结 SWITCH通过实时反馈和客户模拟,提供了一种低成本的社交工作培训解决方案,提升咨询技能分类和动机访谈进阶系统。

Comments 2025 IEEE International Conference on Big Data. ISBN: 979-8-3315-9447-3/25. Page numbers: 3544-3553

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15333 2026-01-27 cs.LG cs.AI q-bio.QM 62%

Empowering LLMs for Structure-Based Drug Design via Exploration-Augmented Latent Inference

通过探索增强的潜在推理增强LLM用于基于结构的药物设计

Xuanning Hu, Anchen Li, Qianli Xing, Jinglong Ji, Hao Tuo, Bo Yang

机构 * College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院) Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(教育部符号计算与知识工程重点实验室) Department of Computer Science, Aalto University(艾尔沃斯大学计算机科学系) College of Artificial Intelligence, Jilin University(吉林大学人工智能学院)

专题命中 工作流自动化 :workflow(abstract);分类 cs.AI、cs.LG

AI总结 ELILLM通过探索增强的潜在推理框架,提升LLM在基于结构的药物设计中的表现,实现更有效的分子生成和结合亲和力预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18638 2026-01-27 cs.LG physics.comp-ph 57%

Physics-Informed Uncertainty Enables Reliable AI-driven Design

物理启发的不确定性使AI驱动的设计更加可靠

Tingkai Xue, Chin Chun Ooi, Yang Jiang, Luu Trung Pham Duong, Pao-Hsiung Chiu, Weijiang Zhao, Nagarajan Raghavan, My Ha Dao

机构 * Department of Mechanical Engineering, National University of Singapore(新加坡国立大学机械工程系) Institute of High Performance Computing, Agency for Science Technology and Research(科技研究局高性能计算研究所) Centre for Frontier AI Research, Agency for Science Technology and Research(科技研究局前沿人工智能研究中心) College of Electronics and Information Engineering, Shenzhen University(深圳大学电子与信息工程学院) Engineering Product Design Pillar, Singapore University of Technology and Design(新加坡科技设计大学工程产品设计支柱) Technology Centre for Offshore and Marine, Singapore(新加坡海洋与近海技术中心)

专题命中 工作流自动化 :workflow(abstract);分类 cs.LG

AI总结 物理启发的不确定性方法提升AI驱动的反向设计效率与鲁棒性

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17987 2026-01-27 cs.LG cs.CV 57%

Systematic Characterization of Minimal Deep Learning Architectures: A Unified Analysis of Convergence, Pruning, and Quantization

对最小深度学习架构的系统表征:收敛性、剪枝和量化的一体化分析

Ziwei Zheng, Huizhi Liang, Vaclav Snasel, Vito Latora, Panos Pardalos, Giuseppe Nicosia, Varun Ojha

机构 * School of Computing, Newcastle University(新castle大学计算机学院) VSB-Technical University of Ostrava(奥斯特拉瓦技术大学) Queen Mary University of London(伦敦女王玛丽大学) University of Florida(佛罗里达大学) University of Catania(卡塔尼亚大学)

专题命中 工作流自动化 :workflow(abstract);分类 cs.LG

AI总结 本文通过系统分析深度学习架构的收敛性、剪枝和量化特性,揭示了最小可学习参数与模型稳定性之间的关系,为图像分类任务中的紧凑模型设计提供了指导。

Journal ref IEEE Conference on Artificial Intelligence 2026 (IEEE CAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17921 2026-01-27 cs.CL 57%

ShapLoRA: Allocation of Low-rank Adaption on Large Language Models via Shapley Value Inspired Importance Estimation

ShapLoRA: 通过受Shapley值启发的重要性估计在大型语言模型上分配低秩适应

Yi Zhao, Qinghua Yao, Xinyuan song, Wei Zhu

机构 * Singapore Management University(新加坡国立管理学院) University of Pennsylvania(宾夕法尼亚大学) Emory University(埃默里大学) University of Hong Kong(香港大学)

专题命中 工作流自动化 :workflow(abstract);分类 cs.CL

AI总结 ShapLoRA通过受Shapley值启发的重要性估计方法,改进大型语言模型的低秩适应分配,提升模型性能。

Comments accepted by CPAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01054 2026-01-27 cs.CR cs.AI cs.CY 57%

Autonomous Penetration Testing: Solving Capture-the-Flag Challenges with LLMs

自主渗透测试:利用大语言模型解决夺旗挑战

Isabelle Bakker, John Hastings

机构 * The Beacom College of Computer and Cyber Sciences(计算机与网络安全科学学院) Dakota State University(达科他州立大学)

专题命中 工作流自动化 :workflow(abstract);分类 cs.AI

AI总结 利用大语言模型自主解决初级渗透测试任务,展示其在自动化简单攻击流程中的潜力,同时揭示安全环境对LLM攻击的挑战。

Comments 6 pages, 2 figures, 3 tables

Journal ref 2025 IEEE Cyber Awareness and Research Symposium (CARS'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09053 2026-01-27 cs.HC 50%

Who Fails Where? LLM and Human Error Patterns in Endometriosis Ultrasound Report Extraction

谁在何处失败?LLM和人类在内膜异位症超声报告提取中的错误模式

Haiyi Li, Yutong Li, Yiheng Chi, Alison Deslandes, Mathew Leonardi, Shay Freger, Yuan Zhang, Jodie Avery, M. Louise Hull, Hsiang-Ting Chen

专题命中 工作流自动化 :workflow(abstract)

AI总结 本研究比较了LLM与人类在内膜异位症超声报告提取中的表现,发现LLM在语法一致性上优于人类,而人类在语义解释上更优,支持人机协作的工作流程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19821 2026-01-27 cs.NE 50%

Fully Tensorized GPU-accelerated Multi-population Evolutionary Algorithm for Constrained Multiobjective Optimization Problems

完全张量化的GPU加速多种群进化算法用于约束多目标优化问题

Weixiong Huang, Rui Wang, Wenhua Li, Sheng Qi, Tianyu Luo, Delong Chen, Tao Zhang, Ling Wang

专题命中 工作流自动化 :workflow(abstract)

AI总结 本文提出了一种完全张量化的GPU加速多种群进化算法GMPEA,用于高效解决时间敏感的约束多目标优化问题,通过并行化方法提升计算效率和解的质量。

Journal ref IEEE Transactions on Evolutionary Computation, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08843 2026-01-27 quant-ph econ.GN math.OC q-fin.EC q-fin.PM q-fin.RM 50%

End-to-End Portfolio Optimization with Quantum Annealing

端到端的量子退火投资组合优化

Sai Nandan Morapakula, Sangram Deshpande, Rakesh Yata, Rushikesh Ubale, Uday Wad, Kazuki Ikeda

专题命中 工作流自动化 :workflow(abstract)

AI总结 本文提出一种结合量子退火与经典优化的端到端投资组合优化方法,实证比较其与基金经理和指数的表现,展示量子辅助选择与经典分配在现实中的应用。

Comments 11 pages, 10 figures, 2 tables

Journal ref Adv Quantum Technol. (2025): e00753

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17418 2026-01-27 cs.HC 50%

GraphPilot: GUI Task Automation with One-Step LLM Reasoning Powered by Knowledge Graph

GraphPilot: 基于知识图谱的GUI任务自动化:由LLM推理一步完成

Mingxian Yu, Siqi Luo, Xu Chen

专题命中 工作流自动化 :agent(abstract)

AI总结 GraphPilot通过基于知识图谱的LLM推理一步完成GUI任务自动化,显著提升任务完成率并降低延迟。

Comments This paper is accepted by the Journal of Intelligent Computing and Networking (JICN) for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17153 2026-01-27 stat.ME stat.AP 50%

Evaluating Aggregated Relational Data Models with Simple Diagnostics

用简单诊断评估聚合关系数据模型

Ian Laga, Benjamin Vogel, Jieyun Wang, Anna Smith, Owen Ward

专题命中 工作流自动化 :workflow(abstract)

AI总结 本文提出了一种用于评估聚合关系数据模型的诊断框架,通过点估计和优化方法快速评估模型拟合情况,帮助研究人员选择合适的模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09899 2026-01-27 cs.RO 50%

Semantic2D: Enabling Semantic Scene Understanding with 2D Lidar Alone

Semantic2D: 仅使用2D激光雷达实现语义场景理解

Zhanteng Xie, Yipeng Pan, Yinqiang Zhang, Jia Pan, Philip Dames

机构 * School of Computing and Data Science, The University of Hong Kong(计算与数据科学学院,香港大学) Department of Mechanical Engineering, Temple University(机械工程系, Temple大学)

专题命中 工作流自动化 :workflow(abstract)

AI总结 Semantic2D通过仅使用2D激光雷达实现语义场景理解,提出首个公开数据集和细粒度分割算法,提升机器人导航性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17055 2026-01-27 cs.CY 50%

AI, Metacognition, and the Verification Bottleneck: A Three-Wave Longitudinal Study of Human Problem-Solving

人工智能、元认知与验证瓶颈:人类问题解决的三波纵向研究

Matthias Huemmer, Franziska Durner, Theophile Shyiramunda, Michelle J. Cummings-Koether

专题命中 工作流自动化 :workflow(abstract)

AI总结 本研究探讨了生成式AI对人类问题解决的影响,发现验证成为瓶颈,提出ACTIVE框架以应对认知负荷问题。

Comments 62 pages, 2 figures, 23 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17035 2026-01-27 cs.DL 50%

Deferred Acceptance Algorithm Improves Peer Review Process

延迟接受算法改进同行评审流程

Christoph Bartneck, Richard Watt, Etienne Borde, Pattara Klinpibul

专题命中 工作流自动化 :agent(abstract)

AI总结 延迟接受算法优化了同行评审流程,通过减少评审数量和延迟,提高了科学出版的效率。

Comments 25 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 软件智能体 13 篇

2601.04886 2026-01-27 cs.SE cs.AI 84%

Analyzing Message-Code Inconsistency in AI Coding Agent-Authored Pull Requests

分析AI编码代理生成的拉取请求中的消息-代码不一致

Jingzhi Gong, Giovanni Pinna, Yixin Bian, Jie M. Zhang

机构 * King's College London(伦敦大学国王学院) University of Trieste(特里斯特大学) Harbin Normal University(哈尔滨师范大学)

专题命中 软件智能体 :agent(title);AI agent(abstract);agentic(abstract);分类 cs.AI、cs.SE

AI总结 研究发现AI生成的PR描述中存在显著的消息-代码不一致问题,导致接受率降低和合并时间延长,需加强验证机制和生成优化以提升人机协作的可信度。

Comments Accepted by MSR'26 Mining Challenge Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18749 2026-01-27 cs.SE 83%

Let's Make Every Pull Request Meaningful: An Empirical Analysis of Developer and Agentic Pull Requests

让每个拉取请求都有意义:对开发者和代理拉取请求的实证分析

Haruhiko Yoshioka, Takahiro Monno, Haruka Tokumasu, Taiki Wakamatsu, Yuki Ota, Nimmi Weeraddana, Kenichi Matsumoto

专题命中 软件智能体 :agentic(title,abstract);AI agent(abstract);分类 cs.SE

AI总结 本研究通过分析40,214个PRs,发现提交者属性对合并结果影响最大,审查相关特征在人类和代理PRs中效果相反,揭示了人机协作提升PR质量的途径。

Comments Accepted for publication in the 23rd International Conference on Mining Software Repositories (MSR '26) : 5 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17627 2026-01-27 cs.SE 79%

Code Change Characteristics and Description Alignment: A Comparative Study of Agentic versus Human Pull Requests

代码变更特征与描述对齐:代理与人类拉取请求的比较研究

Dung Pham, Taher A. Ghaleb

专题命中 软件智能体 :agentic(title);agent(abstract);分类 cs.SE

AI总结 研究比较了代理与人类生成的拉取请求在代码变更特征和描述质量上的差异,发现代理在提交级消息质量上表现更好,但在PR级总结上不如人类,揭示了代理在微观精确性与宏观沟通之间的差距。

Comments Accepted at the 23rd International Conference on Mining Software Repositories (MSR '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13295 2026-01-27 cs.LG cs.AI cs.CL cs.MA cs.SI 75%

CooperBench: Why Coding Agents Cannot be Your Teammates Yet

CooperBench:为何编码代理还不能成为你的队友

Arpandeep Khatua, Hao Zhu, Peter Tran, Arya Prabhudesai, Frederic Sadrieh, Johann K. Lieberwirth, Xinkai Yu, Yicheng Fu, Michael J. Ryan, Jiaxin Pei, Diyi Yang

机构 * Stanford University(斯坦福大学) SAP Labs US(SAP美国实验室)

专题命中 软件智能体 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 CooperBench通过大规模协作编码任务测试,发现AI代理在团队协作中表现不佳,揭示了沟通障碍、承诺偏离和期望错误等关键问题,呼吁发展社交智能。

Comments https://cooperbench.com First two authors contribute equally. The 3th - 6th authors contribute equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04427 2026-01-27 cs.SE cs.AI 73%

Speed at the Cost of Quality: How Cursor AI Increases Short-Term Velocity and Long-Term Complexity in Open-Source Projects

以质量为代价:Cursor AI如何在开源项目中提升短期速度和长期复杂性

Hao He, Courtney Miller, Shyam Agarwal, Christian Kästner, Bogdan Vasilescu

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 软件智能体 :agent(abstract);agentic(abstract);分类 cs.AI、cs.SE

AI总结 本文研究Cursor AI对开源项目开发速度和质量的影响,发现其短期提升速度但长期增加代码复杂性,指出质量保障是早期采用者的主要瓶颈。

Journal ref 23rd International Conference on Mining Software Repositories (MSR '26), April 13--14, 2026, Rio de Janeiro, Brazil

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17406 2026-01-27 cs.SE 70%

Fingerprinting AI Coding Agents on GitHub

在GitHub上指纹识别AI编码代理

Taher A. Ghaleb

专题命中 软件智能体 :agent(abstract);AI agent(abstract);分类 cs.SE

AI总结 该研究通过分析GitHub上的拉取请求,识别AI编码代理的行为特征,揭示了AI生成代码的独特模式,为软件仓库中AI贡献的检测提供了新的方法。

Comments Accepted at the 23rd International Conference on Mining Software Repositories (MSR '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18241 2026-01-27 cs.SE cs.AI 62%

TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance

TAM-Eval: 评估用于自动单元测试维护的LLM

Elena Bruches, Vadim Alperovich, Dari Baturova, Roman Derunets, Daniil Grebenkin, Georgy Mkrtchyan, Oleg Sedukhin, Mikhail Klementev, Ivan Bondarenko, Nikolay Bushkov, Stanislav Moiseev

机构 * Both authors contributed equally to this research.(共同作者)

专题命中 软件智能体 :agentic(abstract);分类 cs.AI、cs.SE

AI总结 TAM-Eval提出一个评估LLM在自动单元测试维护能力的框架和基准,通过测试套件创建、修复和更新三个场景,揭示LLM在真实测试维护中的局限性。

Comments Accepted for publication at the 9th Workshop on Validation, Analysis and Evolution of Software Tests (VST 2026), co-located with the the 33rd IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17584 2026-01-27 cs.SE cs.AI 62%

Prompt Driven Development with Claude Code: Building a Complete TUI Framework for the Ring Programming Language

通过Claude Code驱动的开发:为环编程语言构建完整的终端用户界面框架

Mahmoud Samir Fayed, Ahmed Samir Fayed

专题命中 软件智能体 :workflow(abstract);分类 cs.AI、cs.SE

AI总结 通过Claude Code驱动开发,为环编程语言构建完整终端用户界面框架,展示了提示驱动方法在软件工程中的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏