arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

代码大模型 / AI 编程

代码生成、软件工程智能体、程序修复、测试生成和开发者工具。

2026-01-13 至 2026-01-13 共收录 20 信号源:cs.SE, cs.CL, cs.AI, cs.LG, cs.PL

1. 代码生成 8 篇

2601.07084 2026-01-13 cs.CR cs.SE 79%

How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test

生成安全代码的安全性如何?对抗性提示对LLM防御进行了测试

Melissa Tessa, Iyiola E. Olatunji, Aicha War, Jacques Klein, Tegawendé F. Bissyandé

专题命中 代码生成 :code generation(title,abstract);分类 cs.SE

AI总结 本文评估了生成安全代码方法在对抗性条件下的鲁棒性,发现现有方法在安全性和功能性上存在显著缺陷,提出改进最佳实践。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23010 2026-01-13 cs.SE 79%

TALM: Dynamic Tree-Structured Multi-Agent Framework with Long-Term Memory for Scalable Code Generation

TALM:具有长期记忆的动态树结构多智能体框架用于可扩展的代码生成

Ming-Tung Shen, Yuh-Jzer Joung

专题命中 代码生成 :code generation(title,abstract);分类 cs.SE

AI总结 TALM通过动态树结构多智能体框架结合长期记忆机制,提升复杂代码生成任务中的推理性能和令牌效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05623 2026-01-13 cs.SE cs.AI cs.CL 78%

Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation

以部署为中心的基础设施即代码生成:通过LLM赋能的DevOps模拟实现失败、学习、细化和成功

Tianyi Zhang, Shidong Pan, Zejun Zhang, Zhenchang Xing, Xiaoyu Sun

机构 * Australian National University Canberra Australia New York University \& Columbia University USA Nanyang Technological University Singapore Australian National University Australia Australian National University New York University \& Columbia University Nanyang Technological University

专题命中 代码生成 :code generation(title);分类 cs.SE、cs.CL、cs.AI

AI总结 本文提出IaCGen框架,通过迭代反馈机制提升IaC模板的部署性,实验表明其在10次迭代内可使模板部署成功率提升至91.6%,并进一步通过人工反馈将性能提升至90%以上。

Comments Accepted by FSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06543 2026-01-13 cs.CL cs.AI cs.LG 75%

SimLLM: Fine-Tuning Code LLMs for SimPy-Based Queueing System Simulation

SimLLM:用于基于SimPy的排队系统模拟的代码LLM微调

Jun-Qi Chen, Kun Zhang, Rui Zheng, Ying Zhong

机构 * Institute of Statistics and Big Data, Renmin University of China(中国人民大学统计与大数据研究院) School of Information, Renmin University of China(中国人民大学信息学院) School of Management and Economics, University of Electronic Science and Technology of China(电子科技大学管理学院)

专题命中 代码生成 :code generation(abstract);code model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 SimLLM通过微调开源LLM提升SimPy排队系统模拟代码生成能力,提供替代闭源模型的可行方案。

Comments 33 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07593 2026-01-13 cs.AR cs.CL cs.LG 62%

GRPO with State Mutations: Improving LLM-Based Hardware Test Plan Generation

GRPO与状态突变:改进基于LLM的硬件测试计划生成

Dimple Vijay Kochar, Nathaniel Pinckney, Guan-Ting Liu, Chia-Tung Ho, Chenhui Deng, Haoxing Ren, Brucek Khailany

机构 * Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA(麻省理工学院电子工程与计算机科学系) NVIDIA Research, Austin, TX(NVIDIA研究部) NVIDIA Research, Taiwan(NVIDIA台湾研究部) NVIDIA Research, Santa Clara, CA(NVIDIA圣克拉拉研究部)

专题命中 代码生成 :code generation(abstract);分类 cs.CL、cs.LG

AI总结 GRPO-SMu通过改进LLM训练方法,显著提升硬件验证中测试计划生成的准确率和突变检测率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06181 2026-01-13 cs.AI cs.FL cs.LG cs.LO 62%

Neuro-Symbolic Compliance: Integrating LLMs and SMT Solvers for Automated Financial Legal Analysis

神经符号合规:整合大语言模型与SMT求解器用于自动化金融法律分析

Yung-Shen Hsia, Fang Yu, Jie-Hong Roland Jiang

机构 * Department of Management Information Systems, National ChengChi University, Taipei, Taiwan(管理信息系,中华大学,台北,台湾) Department of Electrical Engineering, National Taiwan University, Taipei, Taiwan(电子工程系,台湾大学,台北,台湾)

专题命中 代码生成 :code generation(abstract);分类 cs.AI、cs.LG

AI总结 本研究整合大语言模型与SMT求解器,提出神经符号合规框架,用于自动化金融法律分析,实现形式可验证性和基于优化的合规修正。

Comments 10 pages, 6 tables, 3 figures, accepted by the 2nd ACM AIware Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06419 2026-01-13 cs.CR cs.PL 57%

Lightweight Yet Secure: Secure Scripting Language Generation via Lightweight LLMs

轻量却安全:通过轻量LLM实现安全脚本语言生成

Keyang Zhang, Zeyu Chen, Xuan Feng, Dongliang Fang, Yaowen Zheng, Zhi Li, Limin Sun

专题命中 代码生成 :code generation(abstract);分类 cs.PL

AI总结 本文提出PSSec框架,通过数据合成与微调提升轻量LLM在生成安全PowerShell脚本方面的能力,显著降低推理成本。

Comments 19 pages,8 figures,conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06126 2026-01-13 cs.AI 57%

NL2Dashboard: A Lightweight and Controllable Framework for Generating Dashboards with LLMs

NL2Dashboard: 一种轻量且可控的基于LLM生成仪表盘的框架

Boshen Shi, Kexin Yang, Yuanbo Yang, Guanguang Chang, Ce Chi, Zhendong Wang, Xing Wang, Junlan Feng

机构 * Jiutian Research, China Mobile(九天研究,中国移动)

专题命中 代码生成 :code generation(abstract);分类 cs.AI

AI总结 NL2Dashboard通过分析-呈现解耦原理,提出轻量可控的仪表盘生成框架,实现高效视觉质量和高可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 软件智能体 2 篇

2512.12216 2026-01-13 cs.SE cs.AI cs.CL 82%

Training Versatile Coding Agents in Synthetic Environments

在合成环境中训练多功能编码代理

Yiqi Zhu, Apurva Gandhi, Graham Neubig

机构 * Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 软件智能体 :coding agent(title,abstract);分类 cs.SE、cs.CL、cs.AI

AI总结 SWE-Playground通过从头生成项目和任务,训练多功能编码代理,能够处理更广泛的编码任务并提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07377 2026-01-13 cs.CV cs.AI 57%

Learning Dynamic Collaborative Network for Semi-supervised 3D Vessel Segmentation

学习动态协作网络用于半监督3D血管分割

Jiao Xu, Xin Chen, Lihe Zhang

机构 * Dalian University of Technology(大连理工大学) City University of Hong Kong(香港城市大学)

专题命中 软件智能体 :repository(abstract);分类 cs.AI

AI总结 DiCo通过动态协作网络和多视图整合模块,提升半监督3D血管分割性能

Comments Accepted to the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 测试生成 1 篇

2601.06185 2026-01-13 cs.SE cs.AI cs.CL 67%

Attention Mechanism and Heuristic Approach: Context-Aware File Ranking Using Multi-Head Self-Attention

注意力机制与启发式方法:基于多头自注意力的上下文感知文件排序

Pradeep Kumar Sharma, Shantanu Godbole, Sarada Prasad Jena, Hritvik Shrivastava

专题命中 测试生成 :repository(abstract);分类 cs.SE、cs.CL、cs.AI

AI总结 本文提出基于多头自注意力机制的上下文感知文件排序方法,通过动态调整特征重要性提升召回率,改进仓库感知的努力估计

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 代码评测 1 篇

2601.06786 2026-01-13 cs.CL 57%

EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs

EpiCaR: 了解未知对提高大语言模型的推理能力至关重要

Jewon Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park, Taesup Kim

机构 * Graduate School of Data Science, Seoul National University(数据科学研究生院,首尔国立大学) Department of Rural Systems Engineering, Seoul National University(农村系统工程系,首尔国立大学) Department of Aerospace Engineering, Seoul National University(航空航天工程系,首尔国立大学)

专题命中 代码评测 :code generation(abstract);分类 cs.CL

AI总结 EpiCaR通过知识校准推理提升大语言模型的推理准确性和校准性,减少推理计算量。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 仓库级理解 8 篇

2601.06109 2026-01-13 cs.AI cs.LG 62%

CBMAS: Cognitive Behavioral Modeling via Activation Steering

通过激活引导的认知行为建模:CBMAS

Ahmed H. Ismail, Anthony Kuang, Ayo Akinkugbe, Kevin Zhu, Sean O'Brien

专题命中 仓库级理解 :repository(abstract);分类 cs.AI、cs.LG

AI总结 CBMAS通过连续激活引导技术,提升大型语言模型的认知行为可解释性,提供诊断框架和数据集以分析模型行为演变。

Comments Accepted to CogInterp @ NeurIPS 2025. Equal contribution by Ahmed H. Ismail and Anthony Kuang

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04293 2026-01-13 cs.AI cond-mat.dis-nn cs.NE math.OC 57%

A Random-Key Optimizer for Combinatorial Optimization

一种用于组合优化的随机键优化器

Antonio A. Chaves, Mauricio G. C. Resende, Martin J. A. Schuetz, J. Kyle Brubaker, Helmut G. Katzgraber, Edilson F. de Arruda, Ricardo M. A. Silva

机构 * Federal U. of São Paulo(巴西圣保罗联邦大学) Amazon Advanced Solutions Lab(亚马逊高级解决方案实验室) DIMACS University of Southampton(南安普顿大学) Federal U. of Pernambuco(巴西佩雷布鲁克联邦大学)

专题命中 仓库级理解 :repository(abstract);分类 cs.AI

AI总结 本文提出了一种随机键优化器,通过模块化设计结合多种元启发式算法,高效解决组合优化问题。

Comments 54 pages, 16 figures, 8 tables

Journal ref J Heuristics 31, 32 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07189 2026-01-13 cs.LG 57%

Standardization of Post-Publication Code Verification by Journals is Possible with the Support of the Community

通过社区支持实现期刊的出版后代码验证标准化

Susana Lopez-Moreno, Eric Dolores-Cuenca, Sangil Kim

机构 * Department of Mathematics, Pusan National University, Busan, South Korea(釜山国立大学数学系) Industrial Mathematics Center, Pusan National University, Busan, South Korea(釜山国立大学工业数学中心) Humanoid Olfactory Display Center, Pusan National University, Yangsan, Gyeongsangnam-do, South Korea(釜山国立大学人形嗅觉显示中心) Department of Mathematics, Yonsei University, Seoul, South Korea(延世大学数学系) Institute for Future Earth, Pusan National University, Busan, South Korea(釜山国立大学未来地球研究所)

专题命中 仓库级理解 :repository(abstract);分类 cs.LG

AI总结 本文提出通过社区支持实现期刊出版后代码验证标准化,通过修改ACM验证徽章允许研究人员提交代码复现,提升机器学习研究的可重复性。

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03973 2026-01-13 cs.SD cs.CL 57%

Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control

Muse:迈向可重现的长篇歌曲生成与细粒度风格控制

Changhao Jiang, Jiahao Chen, Zhenghao Xiang, Zhixiong Yang, Hanchen Wang, Jiabao Zhuang, Xinmeng Che, Jiajun Sun, Hui Li, Yifei Cao, Shihan Dou, Ming Zhang, Junjie Ye, Tao Ji, Tao Gui, Qi Zhang, Xuanjing Huang

专题命中 仓库级理解 :repository(abstract);分类 cs.CL

AI总结 Muse通过开源系统实现可控长篇歌曲生成,利用合成数据集和微调模型,在有限数据下取得竞争性性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06115 2026-01-13 cs.AI 57%

Dreaming Is Not a Bug: A Jung-Inspired Dream Layer for Multi-Agent LLM Companions

梦境并非bug:一种基于荣格思想的多智能体LLM伴侣的梦境层

V. Cheung

专题命中 仓库级理解 :repository(abstract);分类 cs.AI

AI总结 本文提出一种基于荣格思想的梦境层,通过离线生成梦境叙事来增强多智能体LLM伴侣的学习与关系建立能力,将幻觉转化为资源而非bug。

Comments Preprint, 35 pages (5 pages of appendix), 2 figures, 3 tables. Conceptual and architectural proposal with preliminary simulation results

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11819 2026-01-13 physics.optics cond-mat.mtrl-sci 50%

Refractive Index, Its Chromatic Dispersion, and Thermal Coefficients of Four Less Common Glycols

折射率、其色散和热系数四种较少研究的甘油类化合物

Anastasiya Derkachova, Daniel Jakubczyk, Gennadiy Derkachov, Kwasi Nyandey

专题命中 仓库级理解 :repository(abstract)

AI总结 本研究首次系统测量了四种较少研究的甘油类化合物在宽广光谱和温度范围内的折射率、色散和热系数,提供了全面的数据和拟合模型。

Comments 16 pages, 9 figures, raw data in Mendeley Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07701 2026-01-13 cs.RO 50%

Autonomous Driving in Unstructured Environments: How Far Have We Come?

在无结构环境中实现自动驾驶:我们已经取得了多大进展?

Chen Min, Shubin Si, Xu Wang, Hanzhang Xue, Weizhong Jiang, Zitong Chen, Mengmeng Li, Jilin Mei, Erke Shang, Zhipeng Xiao, Bin Dai, Qi Zhu, Hao Fu, Dawei Zhao, Liang Xiao, Yiming Nie, Yu Hu

机构 * Research Center for Intelligent Computing Systems, SKLP, Institute of Computing Technology, Chinese Academy of Sciences(智能计算系统研究中心、SKLP、计算技术研究所、中国科学院) Harbin Engineering University(哈尔滨工程大学) Jianghuai Advance Technology Center, Anhui Provincial Key Laboratory of Humanoid Robot, Anhui Provincial Industry Innovation Center of Humanoid Robot(江淮先进技术中心、安徽省人形机器人重点实验室、安徽省人形机器人产业创新中心) Beihang University(北航) Test Center, National University of Defense Technology(测试中心、国防科技大学) National University of Defense Technology(国防科技大学) Unmanned Systems Technology Research Center, Defense Innovation Institute(无人系统技术研究中心、创新研究院)

专题命中 仓库级理解 :repository(abstract)

AI总结 本文综述了无结构户外环境中自动驾驶的研究进展,涵盖多个关键技术领域,并讨论了未来研究方向。

Comments Accepted by Journal of Field Robotics (JFR) 2025; Survey paper; 59 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06340 2026-01-13 physics.optics physics.comp-ph 50%

Simulation package for solving dynamic diffraction problems in deformed crystals. Bragg, Laue geometry, asymmetric reflections, bend crystals, dislocations, crystals with arbitrary shapes, strain distributions and time dependent problems

用于求解变形晶体中动态衍射问题的模拟包。布拉格、劳厄几何,非对称反射,弯曲晶体,位错,任意形状的晶体,应变分布和时间依赖问题

Jacek Krzywinski, Aliaksei Halavanau

专题命中 仓库级理解 :repository(abstract)

AI总结 本文提出了一种基于FFT BPM的模拟方法,用于高效求解变形晶体中的动态衍射问题,包括布拉格、劳厄和非对称几何中的散射效应,并提供了Python实现及并行计算结构。

详情

展开后加载摘要…

URL PDF HTML 收藏