arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

代码大模型 / AI 编程

代码生成、软件工程智能体、程序修复、测试生成和开发者工具。

2026-01-16 至 2026-01-16 共收录 4 信号源:cs.SE, cs.CL, cs.AI, cs.LG, cs.PL

1. 代码评测 4 篇

2601.10496 2026-01-16 cs.SE cs.AI 62%

Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs

模型看见,模型做?基于暴露的bug-修复偏好评估

Ali Al-Kaswan, Claudio Spiess, Prem Devanbu, Arie van Deursen, Maliheh Izadi

机构 * Delft University of Technology(代尔夫特理工大学) University of California at Davis(加州大学戴维斯分校)

专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI

AI总结 研究发现LLM在生成代码时更倾向于重复bug而非修复,暴露于bug的示例加剧了这一倾向,而修复暴露则仅有微小改进,揭示了LLM可能传播记忆错误的风险。

Comments MSR 2026 Technical Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09084 2026-01-16 cs.CL cs.LG 62%

How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation

需要多少人判断?人类偏好评估的可行性极限

Wilson Y. Lee

机构 * Independent Researcher(独立研究者)

专题命中 代码评测 :code generation(abstract);分类 cs.CL、cs.LG

AI总结 研究发现人类偏好评估中,当偏好信号分散时,比例分配是最佳策略,而精心设计的基准测试能通过减少方差提高检测性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00010 2026-01-16 cs.CL 57%

PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization

PlotCraft: 推动大语言模型在复杂和交互式数据可视化中的极限

Jiajun Zhang, Jianke Zhang, Zeyu Cui, Jiaxi Yang, Lei Zhang, Binyuan Hui, Qiang Liu, Zilei Wang, Liang Wang, Junyang Lin

机构 * USTC(University of Science and Technology of China) THU(Tsinghua University) Alibaba Group(阿里巴巴集团) CASIA(Chinese Academy of Sciences Institute of Automation) SIAT(State Key Laboratory of Information Security)

专题命中 代码评测 :code generation(abstract);分类 cs.CL

AI总结 PlotCraft 提出了一种新的基准和数据集,用于评估大语言模型在复杂和交互式数据可视化任务中的性能,展示了 PlotCraftor 在复杂任务中的显著改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10104 2026-01-16 cs.CV cs.AI 57%

MathDoc: Benchmarking Structured Extraction and Active Refusal on Noisy Mathematics Exam Papers

MathDoc: 评估噪声数学试卷上的结构化提取和主动拒绝基准

Chenyue Zhou, Jiayi Tuo, Shitong Qin, Wei Dai, Mingxuan Wang, Ziwei Zhao, Duoyang Li, Shiyang Su, Yanxi Lu, Yanbiao Ma

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Gaotu Techedu Inc.(高途科技公司) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理研究重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程研究中心,教育部) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Science and Technology of China(中国科学技术大学)

专题命中 代码评测 :repository(abstract);分类 cs.AI

AI总结 MathDoc是首个针对噪声数学试卷的结构化提取和主动拒绝评估基准,揭示了当前MLLMs在处理退化文档时的可靠性缺陷。

详情

展开后加载摘要…

URL PDF HTML 收藏