arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

代码大模型 / AI 编程

代码生成、软件工程智能体、程序修复、测试生成和开发者工具。

共收录 1740 信号源:cs.SE, cs.CL, cs.AI, cs.LG, cs.PL

1. 代码评测 1740 篇

1501.05279 2015-01-22 cs.LG 57%

Extreme Entropy Machines: Robust information theoretic classification

Wojciech Marian Czarnecki, Jacek Tabor

专题命中 代码评测 :repository(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1412.7964 2014-12-30 cs.AI 57%

Knowledge Propagation in Contextualized Knowledge Repositories: an Experimental Evaluation

Loris Bozzato, Luciano Serafini

专题命中 代码评测 :repository(abstract);分类 cs.AI

Comments ARCOE-Logic 2014 Workshop Notes, pp. 13-24

详情

展开后加载摘要…

URL PDF HTML 收藏
1410.5467 2014-10-22 cs.LO cs.LG 57%

Machine Learning of Coq Proof Guidance: First Experiments

Cezary Kaliszyk, Lionel Mamane, Josef Urban

专题命中 代码评测 :repository(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1403.2372 2014-03-12 cs.LG 57%

A Hybrid Feature Selection Method to Improve Performance of a Group of Classification Algorithms

Mehdi Naseriparsa, Amir-Masoud Bidgoli, Touraj Varaee

专题命中 代码评测 :repository(abstract);分类 cs.LG

Comments 8 pages. arXiv admin note: substantial text overlap with arXiv:1403.1946; and text overlap with arXiv:1106.1813 by other authors

Journal ref International Journal of Computer Applications,Vol 69,No 17,pp 28-35,2013

详情

展开后加载摘要…

URL PDF HTML 收藏
1403.1946 2014-03-11 cs.LG 57%

Improving Performance of a Group of Classification Algorithms Using Resampling and Feature Selection

Mehdi Naseriparsa, Amir-masoud Bidgoli, Touraj Varaee

专题命中 代码评测 :repository(abstract);分类 cs.LG

Comments 7 pages

Journal ref World of Computer Science and Information Technology Journal,Vol 3, No 4,pp 70-76,2013

详情

展开后加载摘要…

URL PDF HTML 收藏
1401.7727 2014-01-31 cs.LG cs.CR 57%

Security Evaluation of Support Vector Machines in Adversarial Environments

Battista Biggio, Igino Corona, Blaine Nelson, Benjamin I. P. Rubinstein, Davide Maiorca, Giorgio Fumera, Giorgio Giacinto, and Fabio Roli

专题命中 代码评测 :repository(abstract);分类 cs.LG

Comments 47 pages, 9 figures; chapter accepted into book 'Support Vector Machine Applications'

详情

展开后加载摘要…

URL PDF HTML 收藏
1305.4345 2013-05-21 cs.LG 57%

Ensembles of Classifiers based on Dimensionality Reduction

Alon Schclar, Lior Rokach, Amir Amit

专题命中 代码评测 :repository(abstract);分类 cs.LG

Comments 31 pages, 4 figures, 4 tables, Submitted to Pattern Analysis and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0308032 2011-11-09 cs.CL q-bio.OT 57%

Evaluation of text data mining for database curation: lessons learned from the KDD Challenge Cup

Alexander S. Yeh, Lynette Hirschman, Alexander A. Morgan

专题命中 代码评测 :repository(abstract);分类 cs.CL

Comments 9 pages. This is close to how it appears on the publisher's website (http://bioinformatics.oupjournals.org/cgi/reprint/19/suppl_1/i331) The article wording is the same. Uses bioinformatics-altered.cls, bioinformaticsbib.sty, bioinformaticstitle.sty

Journal ref Bioinformatics Vol. 19 Suppl. 1 2003, pages i331-i339

详情

展开后加载摘要…

URL PDF HTML 收藏
1103.4056 2011-03-22 cs.SE 57%

Software is a directed multigraph (and so is software process)

Robert Dabrowski, Krzysztof Stencel, Grzegorz Timoszuk

专题命中 代码评测 :repository(abstract);分类 cs.SE

Comments 4 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0504065 2009-12-01 cs.AI 57%

Estimating Classification Uncertainty of Bayesian Decision Tree Technique on Financial Data

Vitaly Schetinin, Jonathan E. Fieldsend, Derek Partridge, Wojtek J. Krzanowski, Richard M. Everson, Trevor C. Bailey, Adolfo Hernandez

专题命中 代码评测 :repository(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/9810010 2009-11-30 cs.PL cs.PF 57%

C++ Templates as Partial Evaluation

Todd L. Veldhuizen

专题命中 代码评测 :code generation(abstract);分类 cs.PL

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01153 2026-07-30 cs.CL cs.AI cs.SE 版本更新 56%

Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control

面向AI安全评估的对抗语用学:指令冲突、嵌入命令与策略模糊性基准

Brett Reynolds

机构 * Humber Polytechnic(汉博理工学院) University of Toronto(多伦多大学)

专题命中 代码评测 :分类 cs.SE、cs.CL、cs.AI;repository(comments)

AI总结 提出对抗语用学基准和标注协议,通过语言学控制的分类法评估模型在指令冲突、嵌入命令等场景下的行为,为安全评估提供实证和方法论工具。

Comments 32-page main paper plus 13-page supplement; 6 figures and 17 tables total; code and data artifact available at the linked repository

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13104 2026-07-16 cs.AI cs.CL cs.LG 新提交 56%

Self-Improvements in Modern Agentic Systems: A Survey

现代智能系统中的自我改进:一项综述

Zhe Ren, Yimeng Chen, Dandan Guo, Guowei Rong, Tonghui Li, R. B. Xiong, Qingfeng Lan, Wenyi Wang, Li Nanbo, Yibo Yang, Mingchen Zhuge, Jürgen Schmidhuber

机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学) University of Alberta(阿尔伯塔大学) The Swiss AI Lab IDSIA/USI/SUPSI(瑞士人工智能实验室IDSIA/USI/SUPSI)

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

AI总结 综述现代自我改进智能体从研究走向部署,目标是经验驱动的可控进化。提出系统级框架,将智能体视为基础模型与操作支架的耦合配置,自我改进形式化为更新算子,还组织回顾了相关工作、应用、评估等内容并展望未来。

Comments 97 pages, 12 figures. Project page: https://selfimproving-agent.github.io/ Repository: https://github.com/selfimproving-agent/awesome-Self-Improving-Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22473 2026-06-09 cs.CL cs.AI cs.LG 版本更新 56%

Component Ablation for Efficient Hybrid Language Model Architectures: Performance, Resilience, and Compression Implications

组件消融用于高效混合语言模型架构:性能、鲁棒性和压缩影响

Hector Borobia, Elies Seguí-Mas, Guillermina Tormo-Carbó

机构 * Doctoral Program in Computer Science, University of Valencia(瓦伦西亚大学计算机科学博士项目)

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

AI总结 本文通过组件消融研究混合语言模型,发现注意力机制与替代序列处理路径对性能有显著影响,揭示了模型鲁棒性与压缩优化的关键因素。

Comments 25 pages, 7 figures, 6 tables; revised title, abstract, figures, and data/code repository URL

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12850 2026-05-26 cs.CL cs.AI cs.CR cs.LG 56%

Persona-Model Collapse in Emergent Misalignment

涌现性失调中的人格模型崩溃

Davi Bastos Costa, Renato Vicente

机构 * TELUS Digital Research Hub(TELUS数字研究中心) Center for Artificial Intelligence and Machine Learning(人工智能与机器学习中心) Institute of Mathematics, Statistics and Computer Science(数学、统计与计算机科学研究所) University of São Paulo(圣保罗大学)

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

AI总结 提出人格模型崩溃假说,通过道德易感性(S)和道德稳健性(R)两个指标,证明在有害数据上微调大语言模型会导致模型模拟、区分和维持一致角色的内部能力恶化,从而引发涌现性失调。

Comments 23 pages, 7 figures, 7 tables; NeurIPS 2026 submission; Corrected code repository URL

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06652 2026-05-08 cs.LG cs.AI cs.CL 56%

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels

在没有基准的情况下:在无标签的情况下验证比较LLM安全性评分

Sushant Gautam, Finn Schwall, Annika Willoch Olstad, Fernando Vallecillos Ruiz, Birk Torpmann-Hagen, Sunniva Maria Stordal Bjørklund, Leon Moonen, Klas Pettersen, Michael A. Riegler

机构 * Simula Metropolitan Center for Digital Engineering(Simula 数字工程中心) Oslo Metropolitan University(奥斯陆 Metropolitan 大学) University of Oslo(奥斯陆大学) Simula Research Laboratory(Simula 研究实验室) Norwegian Directorate of Health(挪威健康 Directorate)

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

AI总结 本文提出在无标签情况下验证LLM安全性的方法,通过构建仪器有效性链来替代真实标签,通过实验验证其有效性,并展示了在不同场景下的应用和结果。

Comments SimpleAudit Repository: https://github.com/kelkalot/simpleaudit

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17325 2026-03-18 cs.LG cs.AI cs.CL 56%

Generalizable End-to-End Tool-Use RL with Synthetic CodeGym

通用端到端工具使用强化学习与合成CodeGym

Weihua Du, Hailei Gong, Zhan Ling, Kang Liu, Lingfeng Shen, Xuesong Yao, Yufei Xu, Dingyuan Shi, Yiming Yang, Jiecao Chen

机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所)

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

AI总结 本文提出CodeGym框架,通过合成多样化的多轮工具使用环境,提升LLM代理在不同任务配置下的泛化能力,实验显示Qwen2.5-32B-Instruct在OOD基准测试中准确率提升8.7个百分点。

Comments 24 pages. Accepted to ICLR 2026. Project repository: https://github.com/StigLidu/CodeGym

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17058 2026-01-27 cs.DB cs.AI cs.CL cs.LG 56%

Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs

大模型能帮你清理数据吗?基于大模型的应用级数据准备综述

Wei Zhou, Jun Zhou, Haoyu Wang, Zhenghao Li, Qikang He, Shaokun Han, Guoliang Li, Xuanhe Zhou, Yeye He, Chunwei Liu, Zirui Tang, Bin Wang, Shen Tang, Kai Zuo, Yuyu Luo, Zhenzhe Zheng, Conghui He, Jingren Zhou, Fan Wu

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) Microsoft Research(微软研究院) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Shanghai AI Laboratory(上海人工智能实验室) Xiaohongshu Inc.(小红书公司) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Alibaba Group(阿里巴巴集团)

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

AI总结 本文综述了基于大语言模型的应用级数据准备方法,探讨了数据清洗、整合与丰富三大任务的技术、优势与局限,并提出了可扩展的大语言模型-数据系统和稳健评估协议的未来研究方向。

Comments Please refer to our repository for more details: https://github.com/weAIDB/awesome-data-llm

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11579 2025-11-18 cs.LG cs.AI cs.CL 56%

Decoupling Positional and Symbolic Attention Behavior in Transformers

Felipe Urrutia, Jorge Salas, Alexander Kozachinskiy, Cristian Buc Calderon, Hector Pasten, Cristobal Rojas

机构 * University of Chile(智利大学) CENIA Faculty of mathematics UC(数学学院 UC) IMC UC

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

Comments 32 pages, 12 figures, repository available

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23338 2025-09-30 cs.DB cs.AI cs.CL cs.IR cs.LG 56%

PARROT: A Benchmark for Evaluating LLMs in Cross-System SQL Translation

Wei Zhou, Guoliang Li, Haoyu Wang, Yuxing Han, Xufei Wu, Fan Wu, Xuanhe Zhou

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

Comments To appear in NeurIPS 2025. Welcome your submission to challenge our leaderboard at: https://code4db.github.io/parrot-bench/. Also visit our code repository at: https://github.com/weAIDB/PARROT

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19056 2025-04-29 cs.CV cs.AI cs.CL cs.LG cs.MM 56%

Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions

Mohammad Mahdi Abootorabi, Omid Ghahroodi, Pardis Sadat Zahraei, Hossein Behzadasl, Alireza Mirrokni, Mobina Salimipanah, Arash Rasouli, Bahar Behzadipour, Sara Azarnoush, Benyamin Maleki, Erfan Sadraiye, Kiarash Kiani Feriz, Mahdi Teymouri Nahad, Ali Moghadasi, Abolfazl Eshagh Abianeh, Nizi Nazar, Hamid R. Rabiee, Mahdieh Soleymani Baghshah, Meisam Ahmadi, Ehsaneddin Asgari

机构 * Computer Engineering Department, Sharif University of Technology(谢里夫理工大学计算机工程系) Iran University of Science and Technology(伊朗科学技术大学) Qatar Computing Research Institute(卡塔尔计算研究院)

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

Comments 50 main pages, 30 pages appendix, 21 figures, 8 tables, GitHub Repository: https://github.com/llm-lab-org/Generative-AI-for-Character-Animation-Survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03112 2025-01-30 cs.CL cs.AI cs.CY cs.LG 56%

LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases

Dylan Bouchard, Mohit Singh Chauhan, David Skarbrevik, Viren Bajaj, Zeya Ahmad

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

Comments Journal of Open Source Software; LangFair repository: https://github.com/cvs-health/langfair

Journal ref Journal of Open Source Software, 10(105), 7570 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.14200 2024-03-12 eess.AS cs.AI cs.CL cs.CV cs.LG cs.SD 56%

Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation

Khazar Khorrami, Okko Räsänen

专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)

Comments Final manuscript published in Language Development Research under CC BY-NC-SA 4.0. Pre-print redistributed through arXiv with permission. Replaces corrupted PsyArXiv pre-print repository at https://psyarxiv.com/37zna

Journal ref Language Development Research, 1(1), 123-191 (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08164 2026-08-24 cs.CV 版本更新 50%

Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions

计算机视觉中的持续测试时间适应:方法、基准和未来方向

Sarthak Kumar Maharana, Shambhavi Mishra, Yunbei Zhang, Shuaicheng Niu, Taki Hasan Rafi, Jihun Hamm, Marco Pedersoli, Jose Dolz, Yunhui Guo

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) LIVIA ETS Montreal, ILLS International Laboratory on Learning Systems (ILLS)(蒙特利尔LIVIA ETS,学习系统国际实验室(ILLS)) Tulane University(路易斯安那州立大学) Nanyang Technological University(南洋理工大学) Hanyang University(翰阳大学)

专题命中 代码评测 :repository(abstract)

AI总结 本文针对计算机视觉中训练与测试数据分布不同的问题,定义CTTA问题,分析持续域转移模式,提出分层分类法将现有方法分为三类,回顾代表性方法并展示实验结果,讨论局限性与新兴方向,为持续测试时间适应研究提供路线图。

Comments TMLR 2026 (July edition)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17787 2026-08-21 cs.IR 版本更新 50%

Beyond Uniform Token Training: A Multi-Target Framework for Learning Token-Weighted Objectives in Generative Recommenders

基于课程学习的生成推荐系统中的标记加权多目标学习

Wei-Ning Chiu, Song-Duo Ma, Han-Jay Shu, Chuan-Ju Wang, Pu-Jen Cheng

专题命中 代码评测 :repository(abstract)

AI总结 本文提出基于课程学习的生成推荐系统中的标记加权多目标学习方法,通过两种信息增益策略提升推荐性能,实验表明其在不同语义ID构造下表现优异。

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15205 2026-08-18 math.DG 新提交 50%

Geometric turbulence: a geodesic-regression crisis indicator for equity covariance dynamics, with evidence from African markets

几何湍流:用于股权协方差动态的测地回归危机指标——来自非洲市场的证据

E. B. Camara, Y. U. Gaba, I. G. M. Nsiloulou

专题命中 代码评测 :repository(abstract)

AI总结 该研究提出一种基于对称正定锥测地回归的几何湍流指标,可准确识别市场压力事件,其预测性能优于传统欧氏回归,对应的去风险策略能降低最大回撤。

Comments 22 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14660 2026-08-17 quant-ph 版本更新 50%

Heuristic and Optimal Synthesis of CNOT and Clifford Circuits

CNOT与Clifford电路的启发式及最优综合

Mark Webster, Stergios Koutsioumpas, Dan E Browne

专题命中 代码评测 :repository(abstract)

AI总结 本研究针对CNOT与Clifford电路,提出最优、A*及贪心三类综合算法,可最小化两量子比特门数量或电路深度,经基准测试性能优于现有方法,相关算法已在GitHub仓库开源。

Comments Accepted in Quantum, 6 Aug 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13441 2026-08-14 cs.CV 新提交 50%

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ

Edit2TikZ:面向基于TikZ的科学图像编辑的全面且具有挑战性的基准

Zongyun Zhang, Jiacheng Ruan, Xian Gao, Ruizhu Zhou, Lingcheng Meng, Lining Hu, Ting Liu, Yuzhuo Fu

专题命中 代码评测 :code generation(abstract)

AI总结 本文推出科学图像编辑基准Edit2TikZ,评估主流多模态大语言模型发现其性能不足,通过构建混合训练集并采用课程学习,可显著提升紧凑模型的编译成功率。

Comments 9 pages, 6 figures, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21947 2026-08-12 physics.comp-ph physics.atom-ph 版本更新 50%

AMELI: Angular Matrix Elements of Lanthanide Ions

AMELI: 镧系离子的角矩阵元

Reinhard Caspary

专题命中 代码评测 :repository(abstract)

AI总结 提出基于Slater行列式基和Racah分类的通用框架,计算f^N组态中任意球张量算符的角矩阵元,并构建开源数据库AMELI,以精确算术消除浮点误差,替代传统半经验计算表格。

Comments v2: Abstract, introduction and conclusion rewritten to clarify the novelty and intention of this work. Added Wybourne crystal field Hamiltonians in Section IV G. Improved phase synchronization in Section V B. New Section IX "Application Examples" v3: Minor errors and typographic flaws corrected

Journal ref J. Chem. Phys. 165, 064308 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14633 2026-08-11 cs.CV 版本更新 50%

VIGIL: Tackling Hallucination Detection in Image Recontextualization

VIGIL:应对图像再上下文化中的幻觉检测

Joanna Wojciechowicz, Maria Łubniewska, Jakub Antczak, Justyna Baczyńska, Wojciech Gromski, Wojciech Kozłowski, Maciej Zieba

机构 * Wroclaw University of Science and Technology(沃拉茨拉夫大学科学与技术学院)

专题命中 代码评测 :repository(abstract)

AI总结 VIGIL通过细粒度分类填补多模态模型图像再上下文化中幻觉检测的空白,提出多阶段检测流程并公开资源促进透明度。

Comments 19 pages, 8 figures, 7 tables. Code and data are available at: https://github.com/mlubneuskaya/vigil and https://huggingface.co/datasets/joannaww/VIGIL

详情

展开后加载摘要…

URL PDF HTML 收藏