arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Edinburgh(爱丁堡大学)

2026-01-30 至 2026-01-30 共收录 3
2510.13036 2026-01-30 cs.AI cs.LG

Repairing Reward Functions with Feedback to Mitigate Reward Hacking

通过反馈修复奖励函数以缓解奖励黑客

Stephane Hatgis-Kessell, Logan Mondal Bhamidipaty, Emma Brunskill

机构 * Computer Science Department, Stanford University(计算机科学系, 斯坦福大学) School of Informatics, The University of Edinburgh(信息学院, 埃迪索恩大学)

AI总结 通过反馈修复奖励函数以缓解奖励黑客,提出PBRR方法,通过学习过渡依赖的修正项来改进代理奖励函数,从而在较少偏好下实现高性能策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10163 2026-01-30 cs.CL

Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions

Compound-QA:一个评估大语言模型在复合问题上的基准

Yutao Hou, Yajing Luo, Zhiwen Ruan, Hongru Wang, Weifeng Ge, Yun Chen, Guanhua Chen

机构 * Shanghai University of Finance and Economics(上海财经大学) University of Edinburgh(爱丁堡大学) Southern University of Science and Technology(南方科技大学) Fudan University(复旦大学)

AI总结 Compound-QA基准通过评估LLM在复合问题上的表现,揭示其在理解和推理方面的不足,并提出改进策略。

Comments Accepted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11232 2026-01-30 eess.IV cs.CV

Scale-Equivariant Imaging: Self-Supervised Learning for Image Super-Resolution and Deblurring

尺度等变成像:用于图像超分辨率和去模糊的自监督学习

Jérémy Scanvic, Mike Davies, Patrice Abry, Julián Tachella

机构 * ENSL, CNRS, Laboratoire de Physique(ENSL、CNRS、物理实验室) School of Engineering, University of Edinburgh(工程学院、爱丁堡大学)

AI总结 本文提出尺度等变成像方法,通过利用图像分布的尺度不变性,提升图像超分辨率和去模糊任务的性能,实现与全监督学习相当的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏