Learning Bug Context for PyTorch-to-JAX Translation with LLMs
学习 PyTorch 到 JAX 翻译中的 Bug 上下文
Hung Phan, Son Vu, Tuan Dinh, Nesreen Ahmed, Ali Payani, Ali Jannesari
机构
*
Department of Computer Science, Iowa State University, Ames, IA, USA(Iowa State大学计算机科学系)
;
Department of Human Science, Iowa State University, Ames, IA, USA(Iowa State大学人类科学系)
;
Cisco AI Research, Outshift, San Jose, CA, USA(Cisco AI研究,Outshift,San Jose,加利福尼亚州,美国)
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
AgentBeats:面向开放性、标准化和可复现性的智能体评估代理化
Xiaoyuan Liu, Jianhong Tu, Yuqi Chen, Siyuan Xie, Sihan Ren, Tianneng Shi, Gal Gantar, Evan Sandoval, Donghyun Lee, Daniel Miao, Peter J. Gilbert, Nick Hynes, Mauro Staver, Warren He, David Marn, Andrew Low, Xi Zhang, Elron Bandel, Michal Shmueli-Scheuer, Siva Reddy, Alexandre Drouin, Alexandre Lacoste, Ramayya Krishnan, Elham Tabassi, Yu Su, Victor Barres, Chenguang Wang, Wenbo Guo, Dawn Song
机构
*
University of California, Berkeley(加州大学伯克利分校)
;
Purdue University(普渡大学)
;
University of Ljubljana(卢布尔雅那大学)
;
University of Washington(华盛顿大学)
;
Oasis Labs
;
University of Maryland(马里兰大学)
;
IBM Research(IBM研究院)
;
Mila
;
McGill University(麦吉尔大学)
;
ServiceNow Research(ServiceNow研究院)
;
Carnegie Mellon University(卡内基梅隆大学)
;
National Institute of Standards and Technology(美国国家标准与技术研究院)
;
The Ohio State University(俄亥俄州立大学)
;
University of Cambridge(剑桥大学)
;
University of California, Santa Barbara(加州大学圣塔芭芭拉分校)
Deja Vu at Scale: Paraphrase-Robust Detection of Duplicate Gherkin Steps in Behaviour-Driven Software Testing with Sentence-Transformer Embeddings and a 1.1M-Step Open Benchmark
Comments28 pages, 2 figures, 4 tables. Submitted to Information and Software Technology (Elsevier). Tool, corpus, labelled benchmark, and rubric released at https://github.com/amughalbscs16/cukereuse-release under Apache-2.0
Towards Automated Kernel Generation in the Era of LLMs
面向LLM时代的自动化内核生成
Yang Yu, Peiyu Zang, Chi Hsu Tsai, Haiming Wu, Yixin Shen, Jialing Zhang, Haoyu Wang, Zhiyou Xiao, Jingze Shi, Yuyu Luo, Wentao Zhang, Chunlei Men, Guang Liu, Yonghua Lin
机构
*
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Beijing Normal University(北京师范大学)
;
Peking University(北京大学)
;
Beijing Institute of Technology(北京理工大学)
;
Cornell University(康奈尔大学)
;
Beijing Jiaotong University(北京交通大学)
;
Renmin University of China(中国人民大学)
;
Hong Kong University of Science and Technology (Guangzhou)(广州科技大学)
SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors
SURGE: 大型语言模型作为通用代理代码执行器的潜力
Bohan Lyu, Siqiao Huang, Zichen Liang
机构
*
Department of Computer Science and Technology, Tsinghua(清华大学计算机科学与技术系)
;
Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua(清华大学交叉信息研究院)
An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods
LLM生成代码安全性的提示方法实证评估
Mohammed Kharma, Ahmed Sabbah, Mohammad Alkhanafseh, Mohammad Hammoudeh, David Mohaisen
机构
*
Department of Computer Science, Birzeit University(计算机科学系,巴勒斯坦比泽大学)
;
King Fahd University of Petroleum and Minerals(国王法赫德石油和矿物大学)
;
University of Central Florida(中央佛罗里达大学)
机构
*
Gradient
;
Soochow University(苏州大学)
;
Independent Researcher(独立研究者)
;
University of Southern California(南加州大学)
;
Rice University(Rice大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
University of the Chinese Academy of Sciences(中国科学院大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
GT-HarmBench:通过博弈论视角评估AI安全风险
Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang, Zhijing Jin
机构
*
ETH Zürich(苏黎世联邦理工学院)
;
Berea College(贝雷学院)
;
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
Max Planck Institute for Intelligent Systems, Tübingen, Germany(图宾根德国智能系统马克斯·普朗克研究所)
机构
*
New York University Abu Dhabi(纽约大学阿布扎克校区)
;
Nanyang Technological University(南洋理工大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Harvard University(哈佛大学)
;
Zhejiang University(浙江大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Beijing University of Technology(北京理工大学)
;
The Hong Kong Polytechnic University(香港理工大学)
Comments11 pages, 5 figures, 4 tables. Submitted to arXiv in Computation and Language
Journal refMachine Learning and Principles and Practice of Knowledge Discovery in Databases, ECML PKDD 2025, Communications in Computer and Information Science, vol. 2843, pp. 353-367, Springer, Cham (2026)
CommentsThis paper is withdrawn due to issues in attribution to related work and the fair attribution of benchmark results, which were not adequately addressed at the time of submission. These issues affect the experimental analysis and require substantial revision