arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

2026-01-21 至 2026-01-21 共收录 14
2601.14209 2026-01-21 cs.LG cs.AI cs.CL

InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning

InT:自我提出干预使LLM推理中的信用分配成为可能

Matthew Y. R. Yang, Hao Bai, Ian Wu, Gene Yang, Amrith Setlur, Aviral Kumar

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 InT通过自我提出干预实现LLM推理中的细粒度信用分配,提升模型在数学推理任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08778 2026-01-21 cs.AI cs.DB

Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards

广泛注释错误破坏文本到SQL基准测试和排行榜

Tengjun Jin, Yoojin Choi, Yuxuan Zhu, Daniel Kang

机构 * University of Illinois (UIUC)(伊利诺伊大学香槟分校)

AI总结 本研究发现文本到SQL基准测试中广泛存在的注释错误显著影响了代理性能和排行榜排名,可能误导研究方向和部署选择。

Comments 18 pages, 14 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03515 2026-01-21 cs.RO cs.AI cs.LG cs.SY eess.SY stat.AP

Can the Waymo Open Motion Dataset Support Realistic Behavioral Modeling? A Validation Study with Naturalistic Trajectories

Waymo开放运动数据集能否支持真实的行为建模?一项与自然轨迹相结合的验证研究

Yanlin Zhang, Sungyong Chung, Nachuan Li, Dana Monzer, Hani S. Mahmassani, Samer H. Hamdar, Alireza Talebpour

机构 * Department of Civil and Environmental Engineering, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校土木与环境工程系) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Northwestern University Transportation Center(西北大学交通中心) George Washington University(乔治·华盛顿大学)

AI总结 本研究通过对比自然主义数据与Waymo数据集,发现其无法准确反映真实自动驾驶行为,需谨慎使用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13183 2026-01-21 cs.CL

OpenExempt: A Diagnostic Benchmark for Legal Reasoning and a Framework for Creating Custom Benchmarks on Demand

OpenExempt:法律推理的诊断基准及自定义基准框架

Sergio Servantez, Sarah B. Lawsky, Rajiv Jain, Daniel W. Linna, Kristian Hammond

机构 * Northwestern University(西北大学) Adobe Research(Adobe研究) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 OpenExempt通过动态生成法律推理任务和解决方案,提供一个用于诊断评估的基准和框架,揭示模型在复杂推理中的性能差异。

Comments 25 pages, 9 Figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12758 2026-01-21 cs.CL cs.AI cs.LG

VISPA: Pluralistic Alignment via Automatic Value Selection and Activation

VISPA:通过自动价值选择和激活实现多元对齐

Shenyan Zheng, Jiayou Zhong, Anudeex Shetty, Heng Ji, Preslav Nakov, Usman Naseem

机构 * University of Waterloo(滑铁卢大学) University of Melbourne(墨尔本大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) MBZUAI(马克斯·普朗克智能系统研究所) Macquarie University(麦考瑞大学)

AI总结 VISPA通过自动价值选择和激活实现多元对齐,适用于多种模型和场景,提供了一种可扩展的语言模型对齐方法。

Comments WIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13855 2026-01-21 cs.CL cs.AI

Harnessing Consistency for Robust Test-Time LLM Ensemble

利用一致性提升鲁棒性测试时LLM集成

Zhichen Zeng, Qi Yu, Xiao Lin, Ruizhong Qiu, Xuying Ning, Tianxin Wei, Yuchen Yan, Jingrui He, Hanghang Tong

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 CoRE通过利用模型一致性提升LLM集成的鲁棒性,通过token和model级别的一致性改进集成性能。

Comments 18 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02106 2026-01-21 physics.flu-dyn cs.AI cs.LG gr-qc physics.comp-ph

Resolving Turbulent Magnetohydrodynamics: A Hybrid Operator-Diffusion Framework

解析湍流磁流体动力学:一种混合运算-扩散框架

Semih Kacmaz, E. A. Huerta, Roland Haas

机构 * National Center for Supercomputing Applications, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校国家超级计算中心) Department of Physics, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校物理系) Data Science and Learning Division, Argonne National Laboratory(阿贡国家实验室数据科学与学习部门) Department of Computer Science, The University of Chicago(芝加哥大学计算机科学系) Department of Astronomy, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校天文学系) Department of Physics and Astronomy, University of British Columbia(不列颠哥伦比亚大学物理与天文学系)

AI总结 该研究提出混合运算-扩散框架,结合PINOs与生成扩散模型,实现对高雷诺数MHD湍流的高精度模拟与预测。

Comments 16 pages, 6 figures, 1 table. Content synced with the published version

Journal ref Mach. Learn.: Sci. Technol. 6 (2025) 035057

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03833 2026-01-21 gr-qc astro-ph.IM cs.AI

Sequence modeling of higher-order wave modes of binary black hole mergers

二体黑洞并合高阶波模式的序列建模

Victoria Tiki, Kiet Pham, Eliu Huerta

机构 * Learning Division, Argonne National Laboratory(Argonne国家实验室学习部) Department of Physics, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校物理系) NCSA, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校NCSA) School of Physics and Astronomy, University of Minnesota(明尼苏达大学物理与天文学学院) Department of Computer Science, The University of Chicago(芝加哥大学计算机科学系) Department of Astronomy, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校天文学系)

AI总结 本文提出基于transformer的模型,用于高精度建模二体黑洞并合的高阶引力波模式,实现非线性动力学的快速准确预测。

Comments 32 pages, 2 appendices, 17 figures

Journal ref Class. Quantum Grav. 43 (2026) 015009

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12208 2026-01-21 cs.CL

CoReflect: Conversational Evaluation via Co-Evolutionary Simulation and Reflective Rubric Refinement

CoReflect:通过共进化模拟与反思性评分表细化进行对话评估

Yunzhe Li, Richie Yueqi Feng, Tianxin Wei, Chin-Chia Hsu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 CoReflect通过共进化模拟与反思性评分表细化,实现对话系统的自适应评估,提升测试用例复杂性和评分表精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11560 2026-01-21 cs.IR cs.AI cs.LG

DeepEvidence: Empowering Biomedical Discovery with Deep Knowledge Graph Research

DeepEvidence: 通过深度知识图谱研究赋能生物医学发现

Zifeng Wang, Zheng Chen, Ziwei Yang, Xuan Wang, Qiao Jin, Yifan Peng, Zhiyong Lu, Jimeng Sun

机构 * Keiji AI Institute of Scientific and Industrial Research, Osaka University(大阪大学科学工业研究所) Bioinformatics Center, Institute for Chemical Research, Kyoto University(京都大学化学研究所生物信息中心) Division of Intramural Research, National Library of Medicine, National Institutes of Health(国家卫生研究院生物医学图书馆内部研究部) Department of Population Health Sciences, Weill Cornell Medicine(韦尔·科恩医学中心流行病学与健康科学系) School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机与数据科学学院)

AI总结 DeepEvidence通过深度知识图谱研究框架,系统化地连接异构生物医学资源,提升科学发现的效率和证据综合能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11559 2026-01-21 cs.AI cs.CL cs.LG

MIMIC-RD: Can LLMs differentially diagnose rare diseases in real-world clinical settings?

MIMIC-RD: 能否在真实临床环境中让大语言模型对罕见病进行差异性诊断?

Zilal Eiz AlDin, John Wu, Jeffrey Paul Fung, Jennifer King, Mya Watts, Lauren ONeill, Adam Richard Cross, Jimeng Sun

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Illinois College of Medicine(伊利诺伊大学医学学院)

AI总结 MIMIC-RD通过直接映射临床文本实体到Orphanet,评估LLM在真实临床环境下的罕见病差异性诊断能力,发现现有模型表现不佳,揭示了临床需求与现有能力之间的差距。

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08626 2026-01-21 cs.CL

How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction

大语言模型对顺序敏感性如何?OrderProbe用于确定性结构重建

Yingjie He, Zhaolu Kang, Kehan Jiang, Qianyuan Zhang, Jiachen Qian, Chunlei Meng, Yujie Feng, Yuan Wang, Jiabao Dou, Aming Wu, Leqi Zheng, Pengxiang Zhao, Jiaxin Liu, Zeyu Zhang, Lei Wang, Guansu Wang, Qishi Zhan, Xiaomin He, Meisheng Zhang, Jianyuan Ni

机构 * Peking University(北京大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) City University of Hong Kong(香港城市大学) Fudan University(复旦大学) The Hong Kong Polytechnic University(香港理工大学) Tsinghua University(清华大学) Zhejiang University(浙江大学) University of Illinois Urbana-Champaign(伊利诺伊大学香槟分校) Marquette University(马凯特大学) Juniata College(朱尼阿特学院)

AI总结 研究通过OrderProbe基准评估大语言模型对输入顺序的敏感性,发现即使在前沿模型上,结构重建仍面临挑战,且语义能力与结构鲁棒性存在脱节。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21046 2026-01-21 cs.AI

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

自我进化代理的综述:何时、何地、如何进化以实现人工超级智能

Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, Hongru Wang, Han Xiao, Yuhang Zhou, Shaokun Zhang, Jiayi Zhang, Jinyu Xiang, Yixiong Fang, Qiwen Zhao, Dongrui Liu, Qihan Ren, Cheng Qian, Zhenhailong Wang, Minda Hu, Huazheng Wang, Qingyun Wu, Heng Ji, Mengdi Wang

机构 * Princeton University(普林斯顿大学) Princeton AI Lab(普林斯顿人工智能实验室) Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学) University of Sydney(悉尼大学) Shanghai Jiao Tong University(上海交通大学) Pennsylvania State University(宾夕法尼亚州立大学) University of Michigan(密歇根大学) Oregon State University(俄勒冈州立大学) The Chinese University of Hong Kong(香港中文大学) Fudan University(复旦大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The University of Hong Kong(香港大学) University of California, Santa Barbara(加州大学圣芭芭拉分校) University of California San Diego(加州大学圣地亚哥分校) University of Edinburgh(爱丁堡大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文综述了自我进化代理的现状,探讨了进化机制、适应方法及挑战,为实现人工超级智能提供路线图。

Comments 77 pages, 9 figures, Transactions on Machine Learning Research (01/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05938 2026-01-21 cs.LG cs.AI hep-ex hep-ph hep-th

Uncertainty Quantification From Scaling Laws in Deep Neural Networks

深度神经网络中从缩放定律量化不确定性

Ibrahim Elsharkawy, Yonatan Kahn, Benjamin Hooberman

机构 * Department of Physics, University of Illinois Urbana-Champaign, Urbana, IL, USA(伊利诺伊大学厄巴纳-香槟分校物理系) Department of Physics, University of Toronto, Toronto, ON, Canada(多伦多大学物理系) Vector Institute, Toronto, ON, Canada(向量研究所)

AI总结 本文研究了深度神经网络中通过缩放定律量化不确定性的方法,发现测试损失的方差与均值比值在足够大的训练集下与网络宽度无关。

Comments 18+3 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏