arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Georgia Institute of Technology(佐治亚理工学院)

2025-12-15 至 2025-12-15 共收录 4
2512.11792 2025-12-15 cs.CV

Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation

从跟踪中获取结构:从自回归视频跟踪模型中蒸馏保持结构的运动用于视频生成

Yang Fei, George Stoica, Jingyuan Liu, Qifeng Chen, Ranjay Krishna, Xiaojuan Wang, Benlin Liu

机构 * HKUST(香港科技大学) University of Washington(华盛顿大学) Georgia Tech(佐治亚理工学院) Adobe(Adobe公司)

AI总结 通过蒸馏自回归视频跟踪模型中的结构保持运动先验,SAM2VideoX在视频生成任务中实现了更高的保真度和一致性。

Comments Project Website: https://sam2videox.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11315 2025-12-15 cs.LG

Benchmarking the Generality of Vision-Language-Action Models

对视觉-语言-动作模型通用性的基准测试

Pranav Guruprasad, Sudipta Chowdhury, Harsh Sikka, Mridul Sharma, Helen Lu, Sean Rivera, Aryan Khurana, Hangliang Ren, Yangyue Wang

机构 * Manifold Research Metarch AI Georgia Tech(佐治亚理工学院) Tufts University(塔夫茨大学) Northeastern University(东北大学) Birla Institute of Technology and Science, Pilani(比拉理工学院,帕利尼) Institute for Research and Innovation in Intelligent Systems (IRIIS)(智能系统研究与创新研究所)

AI总结 本文提出MultiNet v1.0基准,评估视觉-语言-动作模型在六个基础能力领域的跨领域泛化能力,发现现有模型在未见领域和模态转移时表现显著退化。

Comments 23 pages, 7 figures, and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11183 2025-12-15 cs.LG

Progress over Points: Reframing LM Benchmarks Around Scientific Objectives

基于点的进展:围绕科学目标重新构建语言模型基准测试

Alwin Jin, Sean M. Hendryx, Vaskar Nath

机构 * Georgia Institute of Technology(佐治亚理工学院) Scale AI

AI总结 本文提出以科学目标为导向的语言模型基准测试环境,通过标准化数据集和训练框架,推动语言模型领域的科学进步。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11090 2025-12-15 stat.ML cs.LG cs.NA math.NA

Data-Driven Model Reduction using WeldNet: Windowed Encoders for Learning Dynamics

基于WeldNet的数据驱动模型降维:用于学习动态的窗口编码

Biraj Dahal, Jiahui Cheng, Hao Liu, Rongjie Lai, Wenjing Liao

机构 * Georgia Institute of Technology(佐治亚理工学院) Meta Hong Kong Baptist University(香港 Baptist大学) Purdue University(普渡大学)

AI总结 WeldNet通过窗口编码和自编码器实现数据驱动的非线性模型降维,有效捕捉复杂系统的动态特性,优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏