Comments15 pages, 6 tables. Introduces the Reward-Shaped Failure Hypothesis and AIRA, a deterministic inspection framework for detecting failure-untruthful patterns in AI-generated code. Includes three empirical studies and a strict matched-control replication
Stability-Weighted Decoding for Diffusion Language Models
扩散语言模型的稳定性加权解码
Yue Wu, Jian Huang
机构
*
Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,Location,Country)
;
School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,Location,Country)
;
Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Hong Kong, China(数据科学与人工智能系,香港理工大学,香港,中国)
LLMs for Text-Based Exploration and Navigation Under Partial Observability
基于部分可观测性的文本基于探索与导航中的大型语言模型
Stephan Sandfuchs, Maximilian Melchert, Jörg Frochte
机构
*
AKIS -- Interdisciplinary Institute for Applied AI and Data Science Ruhr(AKIS——鲁尔跨学科应用人工智能与数据科学研究所)
;
Bochum University of Applied Sciences(波鸿应用科学大学)
Comments15 pages, (to be published Springer Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering [LNICST] )
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Renmin University of China(中国人民大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构
*
SKLP, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所智能计算机研究中心)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Beijing Institute of Technology(北京理工大学)
Sketch2Simulation: Automating Flowsheet Generation via Multi Agent Large Language Models
Sketch2Simulation: 通过多智能体大语言模型自动化流程图生成
Abdullah Bahamdan, Emma Pajak, John D. Hedengren, Antonio del Rio Chanona
机构
*
Sargent Centre for Process Systems Engineering(塞格伦过程系统工程中心)
;
Imperial College London(帝国理工学院伦敦分校)
;
Department of Chemical Engineering(化学工程系)
;
Brigham Young University(BYU( Brigham Young University ))