arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

Google(谷歌)

2026-06-30 至 2026-06-30 共收录 12
2606.30414 2026-06-30 cs.LG

Diffusion Fine-tuning with Rewarded Moment Matching Distillation

基于奖励矩匹配蒸馏的扩散模型微调

Alexis Jacq, Guillaume Couairon, Valentin De Bortoli, Quentin Berthet, Arnaud Doucet, Romuald Elie

机构 * Google DeepMind(谷歌DeepMind)

AI总结 提出奖励矩匹配蒸馏(RMMD)框架,同时蒸馏扩散模型并最大化奖励函数,通过在线策略训练和KL正则化实现高质量生成,在ImageNet上取得更优的FID-奖励权衡,并在气象预报模型GenCast上实现7.5倍加速且性能超越教师模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30058 2026-06-30 cs.CV

Emergence of a Shared Canonical Object Frame from In-the-Wild Videos

从野外视频中涌现共享规范物体框架

Tom Fischer, Martin Sundermeyer, Adam Kortylewski, Eddy Ilg

机构 * University of Technology Nuremberg(图腾技术大学努尔堡分校) Google(谷歌) CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍茨中心CISPA)

AI总结 提出一种自监督方法,利用野外物体视频和SfM噪声相机位姿,通过共享粗规范网格学习密集对应,无需标注即可涌现规范框架,在类别级位姿估计上达到竞争性精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29490 2026-06-30 cs.LG cs.AI

Reported Confidence in LLMs Tracks Commitment More Than Correctness

LLM中的报告置信度追踪承诺而非正确性

Dharshan Kumaran

机构 * Google DeepMind(谷歌DeepMind)

AI总结 本研究通过两阶段弃权范式,发现LLM的言语置信度预测弃权决策远优于预测答案正确性,而令牌对数概率则相反,表明言语置信度是内部承诺准备状态的行为输出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29032 2026-06-30 cs.LG stat.ML

Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

策略感知模拟器学习的理论基础与有效算法

Christoph Dann, Yishay Mansour, Mehryar Mohri

机构 * Google Research(谷歌研究) Tel Aviv University(特拉维夫大学) Courant Institute of Mathematical Sciences(数学科学学院)

AI总结 针对模型强化学习中模拟器利用问题,提出以策略鲁棒性为目标,通过零和极小极大博弈学习模拟器,并给出理论保证与有效算法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08729 2026-06-30 cs.CV cs.GR cs.MM cs.SD

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

Unison:和谐运动、语音和声音以实现以人为中心的音频视频生成

Shihao Cheng, Jiaxu Zhang, Quanyue Song, Shansong Liu, Zhizhi Guo, Xiaolei Zhang, Chi Zhang, Xuelong Li, Zhigang Tu

机构 * DeepMind OpenAI

AI总结 Unison通过显式促进运动、语音和声音模态的协调,解决音频视频生成中模态不一致的问题,提升感知质量和跨模态同步性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10109 2026-06-30 cs.AI cs.HC cs.LG

LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

基于自我报告的LLM代理能够实现通用个体模拟

Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein

机构 * Computer Science Department, Stanford University(斯坦福大学计算机科学系) Department of Communication Studies, Northwestern University(西北大学传播学系) Department of Communication, University of Washington(华盛顿大学传播学系) Google DeepMind(谷歌DeepMind) Department of Sociology, Stanford University(斯坦福大学社会学系) Sciences Po(巴黎政治学院)

AI总结 本文研究了基于自我报告数据的LLM代理在模拟个体行为方面的有效性,通过不同数据源构建代理并验证其在多种任务中的准确性和跨群体公平性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12089 2026-06-30 cs.GT cs.AI cs.HC

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation

选择你的代理:在多方谈判中采用AI顾问、教练和代理的权衡

Kehang Zhu, Nithum Thain, Vivian Tsai, James Wexler, Crystal Qian

机构 * Harvard University(哈佛大学) Google DeepMind Canada(谷歌深Mind加拿大) Google DeepMind United States(谷歌深Mind美国) Google DeepMind New York United States(谷歌深Mind纽约美国)

AI总结 研究探讨了在多方谈判中采用AI顾问、教练和代理的权衡,发现尽管参与者更偏好具有更高控制权的顾问,但代理模式能显著提升集体收益,揭示了AI生成提议与人类行为之间的过滤效应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20606 2026-06-30 cs.CV

Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking

探查并利用视频扩散变换器特征以实现稳健的点跟踪

Soowon Son, Honggyu An, Jisu Nam, Hyunah Ko, Chaehyun Kim, Dahyun Chung, Siyoon Jin, Jung Yi, Junhwa Hur, Seungryong Kim

机构 * KAIST AI(韩国科学技术院人工智能) Google DeepMind(谷歌DeepMind)

AI总结 本文探讨了视频扩散变换器在点跟踪中的优势,提出DiTracker框架,通过整合视频DiT特征提升跟踪鲁棒性,实验证明其在挑战性场景下表现优异。

Comments Project Page: https://cvlab-kaist.github.io/DiTracker/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06090 2026-06-30 cs.SE cs.AI cs.PF

SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?

SWE-fficiency:语言模型能否在真实工作负载上优化现实世界仓库?

Jeffrey Jian Ma, Milad Hashemi, Amir Yazdanbakhsh, Kevin Swersky, Ofir Press, Enhui Li, Vijay Janapa Reddi, Parthasarathy Ranganathan

机构 * Harvard University(哈佛大学) Google DeepMind(谷歌DeepMind) Princeton University(普林斯顿大学) Google(谷歌)

AI总结 本文提出SWE-fficiency基准,评估语言模型在真实工作负载上优化软件仓库性能的能力,发现现有模型在定位优化机会和保持正确性方面表现不佳。

Comments Appearing at ICML 2026. Data, code, and leaderboard are available at https://swefficiency.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15864 2026-06-30 cs.LG

Improving Rectified Flow with Boundary Conditions

通过边界条件改进校正流

Xixi Hu, Runlong Liao, Keyang Xu, Bo Liu, Yeqing Li, Eugene Ie, Hongliang Fei, Qiang Liu

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Google(谷歌)

AI总结 本文提出边界约束校正流模型,通过强制边界条件提升生成模型性能,在ImageNet上使用ODE和SDE采样分别提升FID分数8.01%和8.98%。

Comments ICCV 2025

Journal ref Proc. IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 18177-18186

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00599 2026-06-30 cs.CV

XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity

XYZ-IBD:在现实工业复杂性下评估鲁棒的6D物体姿态估计的基准测试

Junwen Huang, Jiaqi Hu, Peter KT Yu, Slobodan Ilic, Martin Sundermeyer, Benjamin Busam

机构 * Technical University of Munich(慕尼黑工业大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) XYZ Robotics ROBOX Google(谷歌)

AI总结 XYZ-IBD是一个专门针对工业分拣的高精度基准,包含75个多视角真实场景,通过高密度随机堆叠和多实例模糊性反映真实机器人操作挑战,验证了工业视觉中6D姿态估计的性能退化问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18864 2026-06-30 cs.AI cs.CL cs.HC cs.LG physics.soc-ph q-bio.OT

Accelerating scientific discovery with Co-Scientist

用Co-Scientist加速科学发现

Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad Tomasev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Rimanic, Marina Boia, Ivan Budiselic, Ben Feinstein, Mathias Bellaiche, Tom Sheffer, Jan Freyberg, Jeremy Ratcliff, Ottavia Bertolli, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R D Costa, José R Penadés, Gary Peltz, Yossi Matias, James Manyika, Demis Hassabis, Yunhan Xu, Pushmeet Kohli, Annalisa Pawlosky, Alan Karthikesalingam, Vivek Natarajan

机构 * Google Cloud AI Research(谷歌云人工智能研究) Google DeepMind(谷歌DeepMind) Google Research(谷歌研究) Stanford University School of Medicine(斯坦福大学医学院) Houston Methodist(休斯顿卫理公会医院) Sequome Fleming Initiative and Imperial College London(弗莱明倡议与伦敦帝国理工学院)

AI总结 Co-Scientist是一种基于Gemini的多智能体AI系统,通过异步任务框架和竞赛进化过程,提升科学假设生成质量,应用于药物再利用、新靶点发现和抗菌机制解释,验证了其加速科学发现的能力。

Comments 157 pages in total (main 42 pages, supplementary information 115 pages), 4 main figures, 1 main table, 6 extended data figures, 2 extended data tables, 9 supplementary figures, 4 supplementary tables, 37 main references, 117 supplementary references. Nature (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏