Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment
超越正确性:通过可扩展的智能体判断标注增强代码大模型的架构推理能力
Kirill Vasilevski, Ximing Dong, Benjamin Rombaut, Milad Soltany, Ruochen Deng, Jiahuei Lin, Arthur Leung, Dayi Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan
机构
*
Centre for Software Excellence, Huawei Canada(华为加拿大软件卓越中心)
;
Department of Computer Science, University of Manitoba, Canada(曼尼托巴大学计算机科学系)
;
School of Computing, Queen’s University, Canada(皇后大学计算科学学院)
CommentsAccepted to ISSRE 2026. Major revision and retitling of arXiv:2510.20692v1. Refocuses the paper on reliable neurosymbolic access-control policy analysis; updates the PolicySummarizer method, multi-cloud evaluation, and user-study results. 13 pages, 6 figures. Corrected arXiv title metadata to match the accepted ISSRE version. No substantive content changes from the previous version
Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments
自然语言解释中的判断质量测量:来自预测锦标赛的证据
Christopher W. Karvetski, Sheldon S. Huang, Simas Kučinskas, Nadja Flechner, Jingyu Hu, Philip Tetlock, Ezra Karger
机构
*
Forecasting Research Institute(预测研究所)
;
Good Judgment Inc(Good Judgment公司)
;
University of Toronto(多伦多大学)
;
Vector Institute for Artificial Intelligence(向量人工智能研究所)
;
Stanford University(斯坦福大学)
;
School of Arts and Sciences & Wharton, University of Pennsylvania(宾夕法尼亚大学文理学院与沃顿商学院)
;
Federal Reserve Bank of Chicago(芝加哥联邦储备银行)
;
Federal Reserve System(联邦储备系统)
专题命中
推理与问题求解
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL
机构
*
Purdue University(普渡大学)
;
Indian Institute of Technology, Kharagpur(印度理工学院卡哈拉格普尔分校)
;
DEVCOM Army Research Lab(DEVCOM陆军研究实验室)
;
National University of Singapore(新加坡国立大学)
机构
*
School of Computer Science, University of Auckland(奥克兰大学计算机科学学院)
;
Department of Electronics and Electrical Engineering, National Yang Ming Chiao Tung University(国立阳明交通大学电子与电机工程学系)
;
School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算与信息系统学院)
;
Shenzhen Key Laboratory of Internet Information Collaboration, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)深圳市互联网信息协同重点实验室)
专题命中
推理与问题求解
:LLM(summary_cn);large language model(abstract);language model(abstract);分类 cs.AI
The CRISTAL Method: Neurosymbolic analysis from AI-synthesized world models
CRISTAL方法:来自AI合成世界模型的神经符号分析
Rafael Kaufmann, Felix Neubürger, Michael Walters, Thomas Kopinski, Dimitrije Marković
机构
*
GAIA Lab(GAIA实验室)
;
South Westphalia University of Applied Sciences(西南弗里西亚应用科学大学)
;
Technical University Dresden(德累斯顿技术大学)
;
Primordia Co.(Primordia公司)
When are likely answers right? On Sequence Probability and Correctness in LLMs
何时可能答案正确?关于LLM中的序列概率与正确性
Johannes Zenn, Jonas Geiping
机构
*
Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
;
AI Center Tübingen(图宾根人工智能中心)
;
University of Tübingen(图宾根大学)
;
IMPRS-IS(图宾根马克斯·普朗克研究所)
专题命中
推理与问题求解
:LLM(title_cn);large language model(abstract);language model(abstract);分类 cs.LG