arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Cambridge(剑桥大学)

2026-08-27 至 2026-08-27 共收录 7
2608.26009 2026-08-27 cs.AI cs.LG cs.LO 新提交

Imitation Learning for Connection-Tableau Construction

用于连接表构造的模仿学习

Fredrik Rømming, Mantas Bakšys, Martin S. Fixman, Sean B. Holden

机构 * University of Cambridge(剑桥大学)

AI总结 该研究将连接表构造建模为迁移系统中的策略,结合模仿学习与图神经网络训练策略,在多个基准上较leanCoP提升了定理证明的问题解决率并大幅减少步数。

Comments 9 pages. Code: this https URL (https://github.com/fredrrom/connections)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25990 2026-08-27 cs.LG 新提交

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

频谱分配:为何Muon优于Adam,以及如何改进Muon

Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland

机构 * University of Cambridge(剑桥大学) Tsinghua University(清华大学)

AI总结 本文通过频谱分析揭示Muon优于Adam的机制,提出SAMuon及其简化版SAMuon-lite,在多规模modded-nanogpt模型上,二者均优于AdamW和Muon,且SAMuon可减少13.3%-24.0%训练令牌。

Comments 34 pages, 13 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25924 2026-08-27 cs.CV 新提交

Visual General Intelligence: A White Paper

视觉通用智能:白皮书

Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian, Shangzhe Wu, Oishi Deb, Ryousuke Yamada, Christian Rupprecht, Jianyuan Wang, Kohsuke Ide, Koichi Namekata, Xianzheng Ma, Yiming Chen, Robert Geirhos, Aditi Raghunathan, Yuki M. Asano, Deva Ramanan, David Fouhey, Andrew J. Davison, Yilun Du, Jiajun Wu, Zhuang Liu

机构 * National Institute of Advanced Industrial Science and Technology (AIST)(日本国立先进工业科学技术研究所(AIST)) Visual Geometry Group (VGG), University of Oxford(牛津大学视觉几何组(VGG)) CADDi(CADDi公司) OpenAI(OpenAI公司) University of Cambridge(剑桥大学) University of Technology Nuremberg(纽伦堡工业大学) University of Tsukuba(筑波大学) Google DeepMind(谷歌DeepMind) Carnegie Mellon University(卡内基梅隆大学) New York University(纽约大学) Imperial College London(伦敦帝国学院) Harvard University(哈佛大学) Stanford University(斯坦福大学) Princeton University(普林斯顿大学)

AI总结 本文以视觉为中心探讨通用人工智能(AGI)路径,提出视觉通用智能(VGI)概念,明确AGI时代计算机视觉的发展原则、模态、基准等关键方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25832 2026-08-27 cs.CL cs.AI cs.GT cs.LG 新提交

Skill Issue: Are Skills Language-Invariant in LLMs?

技能问题:大型语言模型的技能是否具有语言不变性?

Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen

机构 * A*STAR(新加坡科技研究局) Weizmann Institute of Science(魏茨曼科学研究所) MIT-IBM Watson AI Lab(麻省理工学院-IBM沃森人工智能实验室) University of Cambridge(剑桥大学) EleutherAI

AI总结 该研究以多语言自博弈方法量化大型语言模型的跨语言技能不一致性,发现同一模型在不同语言下博弈实力差异显著,语言可影响决策阶段,技能差异是多语言模型开发的主要障碍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25707 2026-08-27 cs.LG 新提交

Fairness-Aware Test-Time Prompt Tuning

感知公平性的测试时提示调优

Yoann Launay, Parameswaran Kamalaruban, Tom Kempton, Stuart Burrell, David Sutton

机构 * University of Cambridge(剑桥大学) Visa Inc.(维萨公司) University of Manchester(曼彻斯特大学)

AI总结 本文针对分布偏移下视觉-语言模型的公平性问题,提出了感知公平性的情节式测试时自适应方法FairTPT,通过软提示调优联合优化熵,实现公平性提升并优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25623 2026-08-27 cs.AI cs.CY cs.HC 新提交

Using profiles of cognitive capability to assess AI suitability for workplace tasks

利用认知能力画像评估AI对职场任务的适用性

Jonathan Prunty, Marko Tešić, Patrick Quinn, José Hernández-Orallo, Lucy Cheke

机构 * University of Cambridge(剑桥大学) Department for Science, Innovation and Technology(科学、创新与技术部) Universitat Politècnica de València(瓦伦西亚理工大学)

AI总结 本研究提出基于共享核心认知能力的画像流程,通过AI与职场任务的认知维度匹配,为AI职场任务范围界定提供工具,还探讨了扩展至人机协同画像的方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25114 2026-08-27 cs.LG cs.AI 新提交

Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment

Flower Hub:用于联邦学习仿真与部署的可复现基准测试平台

Yan Gao, Mohammad Naseri, Javier Fernandez-Marques, Dimitris Stripelis, Lorenzo Sani, Davide Eynard, Fan Zhang, Hong Jia, Ting Dang, D. B. Emerson, Fatemeh Tavakoli, Ole Werger, Lars Wulfert, Petros Demetrakopoulos, Sofia Tsekeridou, InSeo Song, KangYoon Lee, Honghao Li, Lingjuan Lyu, John P Dickerson, Daniel Janes Beutel, Nicholas D. Lane

机构 * Flower Labs University of Cambridge(剑桥大学) University of Auckland(奥克兰大学) University of Melbourne(墨尔本大学) Vector Institute Fraunhofer IMS(弗劳恩霍夫应用固体物理与材料力学研究所) NetCompany Gachon University(嘉泉大学) Owkin Sony AI(索尼人工智能)

AI总结 Flower Hub是一款联邦学习基准测试平台,可将基准打包为可执行应用,支持跨仿真与部署运行,涵盖多领域任务,推动联邦学习基准测试向可复用方向发展。

详情

展开后加载摘要…

URL PDF HTML 收藏