Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
Duel-Evolve:通过LLM自我偏好实现无奖励的测试时间扩展
机构 * Columbia University(哥伦比亚大学) ; Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)
AI总结 Duel-Evolve通过LLM自我偏好实现无奖励的测试时间优化,提升MathBench和LiveCodeBench的准确性。
高校专区
Duel-Evolve:通过LLM自我偏好实现无奖励的测试时间扩展
机构 * Columbia University(哥伦比亚大学) ; Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)
AI总结 Duel-Evolve通过LLM自我偏好实现无奖励的测试时间优化,提升MathBench和LiveCodeBench的准确性。
PoSh:利用场景图引导LLM-as-a-Judge进行详细图像描述
机构 * Columbia University(哥伦比亚大学) ; The University of Texas at Austin(德克萨斯大学奥斯汀分校) ; The National Gallery of Art(国家艺术馆) ; UCLA(加州大学洛杉矶分校) ; UNC Chapel Hill(北卡罗来纳大学教堂山分校)
AI总结 PoSh通过利用场景图引导LLM-as-a-Judge,提供了一种更准确的详细图像描述评估方法,并展示了其在新数据集DOCENT中的优越表现。
Comments Accepted at ICLR 2026. 26 pages, 9 figures. Metric/benchmark available at https://github.com/amith-ananthram/posh
认知模型和AI算法为设计语言代理提供模板
机构 * Department of Computer Science, Princeton University(普林斯顿大学计算机科学系) ; Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology(麻省理工学院脑科学与认知科学系) ; Zuckerman Mind Brain Behavior Institute(祖克曼心智大脑行为研究所) ; Department of Psychiatry, Columbia University(哥伦比亚大学精神医学系) ; Neuroscience Institute(神经科学研究所) ; Department of Machine Learning, Carnegie Mellon University(卡内基梅隆大学机器学习系) ; Department of Psychology, Princeton University(普林斯顿大学心理学系)
AI总结 本文提出利用认知模型和AI算法设计语言代理的模板,旨在开发更有效且可解释的语言代理。