Learning Bug Context for PyTorch-to-JAX Translation with LLMs
学习 PyTorch 到 JAX 翻译中的 Bug 上下文
Hung Phan, Son Vu, Tuan Dinh, Nesreen Ahmed, Ali Payani, Ali Jannesari
机构
*
Department of Computer Science, Iowa State University, Ames, IA, USA(Iowa State大学计算机科学系)
;
Department of Human Science, Iowa State University, Ames, IA, USA(Iowa State大学人类科学系)
;
Cisco AI Research, Outshift, San Jose, CA, USA(Cisco AI研究,Outshift,San Jose,加利福尼亚州,美国)
Deja Vu at Scale: Paraphrase-Robust Detection of Duplicate Gherkin Steps in Behaviour-Driven Software Testing with Sentence-Transformer Embeddings and a 1.1M-Step Open Benchmark
Comments28 pages, 2 figures, 4 tables. Submitted to Information and Software Technology (Elsevier). Tool, corpus, labelled benchmark, and rubric released at https://github.com/amughalbscs16/cukereuse-release under Apache-2.0
Towards Automated Kernel Generation in the Era of LLMs
面向LLM时代的自动化内核生成
Yang Yu, Peiyu Zang, Chi Hsu Tsai, Haiming Wu, Yixin Shen, Jialing Zhang, Haoyu Wang, Zhiyou Xiao, Jingze Shi, Yuyu Luo, Wentao Zhang, Chunlei Men, Guang Liu, Yonghua Lin
机构
*
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Beijing Normal University(北京师范大学)
;
Peking University(北京大学)
;
Beijing Institute of Technology(北京理工大学)
;
Cornell University(康奈尔大学)
;
Beijing Jiaotong University(北京交通大学)
;
Renmin University of China(中国人民大学)
;
Hong Kong University of Science and Technology (Guangzhou)(广州科技大学)
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes
LLM推理的周期表:推理范式、方法与失败模式的结构化综述
Avinash Anand, Mahisha Ramesh, Avni Mittal, Ashutosh Kumar, Rishitej Reddy Vyalla, Erik Cambria, Zhengkui Wang, Timothy Liu, Aik Beng Ng, Simon See, Rajiv Ratn Shah
机构
*
Singapore Institute of Technology(新加坡理工大学)
;
Nvidia AI Center (SNAIC)(英伟达人工智能中心(SNAIC))
;
MIDAS Lab, IIIT Delhi(IIIT德里MIDAS实验室)
;
MIDAS Lab, IIT Mandi(IIT曼迪MIDAS实验室)
;
Owl Autonomous Imaging, Inc.(Owl自主成像公司)
;
College of Computing & Data Science, NTU Singapore(新加坡南洋理工大学计算与数据科学学院)
;
NVIDIA AI Technology Centre, Singapore(英伟达新加坡人工智能技术中心)
;
Department of Computer Science and Engineering, IIT Kanpur(IIT坎普尔计算机科学与工程系)
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
SciVisAgentBench:用于评估科学数据分析和可视化代理的基准测试
Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski, Jianxin Sun, Guoxi Liu, Helgi I. Ingolfsson, David Lenz, Hanqi Guo, Hongfeng Yu, Teja Leburu, Michael Molash, Bei Wang, Tom Peterka, Chaoli Wang, Shusen Liu
机构
*
University of Notre Dame(圣母大学)
;
Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)
;
University of Utah(犹他大学)
;
University of Nebraska–Lincoln(内布拉斯加大学林肯分校)
;
The Ohio State University(俄亥俄州立大学)
;
Argonne National Laboratory(阿贡国家实验室)
;
Anthropic PBC
Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs
业务即规则:面向LLM的业务规则流建模基准与框架
Chen Yang, Ruping Xu, Ruizhe Li, Bin Cao, Jing Fan
机构
*
Zhejiang University of Technology(浙江工业大学)
;
Zhejiang Key Laboratory of Visual Information Intelligent Processing(浙江省视觉信息智能处理重点实验室)
;
University of Aberdeen(阿伯丁大学)
;
University of Birmingham(伯明翰大学)