arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Oxford(牛津大学)

2025-12-19 至 2025-12-19 共收录 4
2510.08697 2025-12-19 cs.SE cs.AI cs.CL

BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution

BigCodeArena: 通过执行揭示更多可靠的代码生成人类偏好

Terry Yue Zhuo, Xiaolong Jin, Hange Liu, Juyong Jiang, Tianyang Liu, Chen Gong, Bhupesh Bishnoi, Vaisakhi Mishra, Marek Suppa, Noah Ziems, Saiteja Utpala, Ming Xu, Guangyu Song, Kaixin Li, Yuhan Cao, Bo Liu, Zheng Liu, Sabina Abdurakhmanova, Wenhao Yu, Mengzhao Jia, Jihan Yao, Kenneth Hamilton, Kumar Shridhar, Minh Chien Vu, Dingmin Wang, Jiawei Liu, Zijian Wang, Qian Liu, Binyuan Hui, Meg Risdal, Ahsen Khaliq, Atin Sood, Zhenchang Xing, Wasi Uddin Ahmad, John Grundy, David Lo, Banghua Zhu, Xiaoning Du, Torsten Scholak, Leandro von Werra

机构 * Monash University(墨尔本大学) CSIRO’s Data61(CSIRO的数据61) Purdue University(普渡大学) Independent(独立) HKUST (Guangzhou)(香港科技大学(广州)) UCSD(加州大学圣地亚哥分校) UVA(弗吉尼亚大学) CNRS, France(法国国家科学研究中心) IBM Cisco(思科) Comenius University in Bratislava(布拉提斯拉瓦康门纽斯大学) University of Notre Dame(Notre Dame大学) Uber Tano Labs(Tano实验室) NUS(国立大学新加坡) Institute of Automation, CAS(中国科学院自动化研究所) Tencent AI Lab(腾讯AI实验室) University of Washington(华盛顿大学) Nevsky Collective(Nevsky集体) ETH Zurich(苏黎世联邦理工学院) Detomo Inc(Detomo公司) University of Oxford(牛津大学) UIUC(伊利诺伊大学香槟分校) Google(谷歌) NVIDIA Singapore Management University(新加坡管理学院) ServiceNow Research(ServiceNow研究) Hugging Face

AI总结 BigCodeArena通过执行揭示更多可靠的代码生成人类偏好,提出自动评分基准评估LLM编码质量。

Comments Built with love by the BigCode community :)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15892 2025-12-19 cs.CR cs.AI

VET Your Agent: Towards Host-Independent Autonomy via Verifiable Execution Traces

验证你的代理:通过可验证的执行轨迹实现主机无关的自主性

Artem Grigor, Christian Schroeder de Witt, Simon Birnbach, Ivan Martinovic

机构 * University of Oxford(牛津大学)

AI总结 VET通过可验证的执行轨迹实现主机无关的自主性,为未来完全自主的代理系统奠定基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07741 2025-12-19 cs.LG cs.SD

A multimodal Bayesian Network for symptom-level depression and anxiety prediction from voice and speech data

一种多模态贝叶斯网络用于从语音和语音数据中预测症状层面的抑郁和焦虑

Agnes Norbury, George Fairs, Alexandra L. Georgescu, Matthew M. Nour, Emilia Molimpakis, Stefano Goria

机构 * thymia Limited(thymia有限公司) Institute of Psychiatry, Psychology & Neuroscience, King’s College London(心理学与神经科学研究院,伦敦国王学院) Department of Psychiatry, University of Oxford(牛津大学精神病学系) Max Planck UCL Centre for Computational Psychiatry and Ageing, University College London(Max Planck大学学院计算精神病学与衰老中心,伦敦大学学院)

AI总结 本文提出了一种多模态贝叶斯网络模型,用于从语音和语音数据中预测抑郁和焦虑症状,通过评估模型性能和公平性,展示了其在临床应用中的潜力。

Journal ref Scientific Reports (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08282 2025-12-19 cs.LG cs.AI

Individualised Treatment Effects Estimation with Composite Treatments and Composite Outcomes

基于复合治疗和复合结局的个体化治疗效应估计

Vinod Kumar Chauhan, Lei Clifton, Gaurav Nigam, David A. Clifton

机构 * Institute of Biomedical Engineering at the University of Oxford UK(牛津大学生物医学工程研究所) Nuffield Department of Primary Care Health Sciences at the University of Oxford UK(牛津大学初级卫生保健科学系) Nuffield Department of Medicine at the University of Oxford(牛津大学医学系) Oxford-Suzhou Institute of Advanced Research (OSCAR)(牛津-苏州先进研究所)

AI总结 本文提出H-Learner方法,通过动态共享信息解决复合治疗和复合结局下的个体化治疗效应估计问题。

Comments Accepted to The 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (7 pages (double column), 4 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏