arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作者

Alexei A. Efros

Computer Vision

至 收录 115
2607.18237 2026-07-21 cs.CV cs.LG 新提交

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

视觉相似性的多种意义:一种文本提示的图像感知度量

Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei A. Efros, Richard Zhang

AI总结 研究人类视觉相似性判断的上下文依赖问题,通过引入大规模数据集微调VLM,产生能捕捉多种视觉相似性意义的TPIPS度量,该度量与人类感知更契合,还在多种任务中开启新功能。

Comments Project Webpage: https://peterwang512.github.io/TPIPS

URL PDF HTML 收藏
2607.14645 2026-07-21 cs.CV 版本更新

Autoregressive Modeling of Film with Applications in Video Montage

电影的自回归建模及其在视频蒙太奇中的应用

Marcelo Sandoval-Castañeda, Fabian Caba Heilbron, Shiry Ginosar, Bryan Russell, Josef Sivic, Alexei A. Efros, Greg Shakhnarovich

机构 * TTI-Chicago(芝加哥丰田理工学院) Adobe(奥多比公司) Czech Institute of Informatics, Robotics and Cybernetics, Czech Technical University(捷克技术大学捷克信息学、机器人学与控制论研究所) UC Berkeley(加州大学伯克利分校)

AI总结 研究针对视频蒙太奇挑战,提出FilmGPT自回归Transformer,通过在电影语料库训练捕捉电影“语法”,推理时用镜头约束解码算法选最佳镜头,在镜头预测和电影编辑任务中表现出色,还适用于多种视频蒙太奇应用。

Journal ref SIGGRAPH Conference Papers '26, Article 129, 1-11, ACM, 2026

URL PDF HTML 收藏
2606.20536 2026-06-19 cs.CV 新提交

The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation

FID 彩票:量化生成模型评估中的隐藏随机性

Nicolas Dufour, Alexei A. Efros, Patrick Pérez

机构 * Kyutai UC Berkeley(加州大学伯克利分校)

AI总结 研究FID作为随机变量在训练和生成种子上的方差,发现重训练比重采样导致更大FID波动,提出新评估协议:使用每类最优引导、报告多个训练种子的误差条。

Comments Website: https://kyutai.org/fid-lottery

URL PDF HTML 收藏
2606.03990 2026-06-03 cs.LG cs.CL cs.CV

Neuron Populations Exhibit Divergent Selectivity with Scale

神经元群体随规模表现出分化的选择性

Amil Dravid, Yasaman Bahri, Alexei A. Efros, Yossi Gandelsman

机构 * UC Berkeley(加州大学伯克利分校) TTIC

AI总结 通过分析Rosetta神经元在不同规模模型中的分布与特性,发现其数量遵循次线性幂律增长,且选择性随规模增强,而非Rosetta神经元则保持低选择性,提出一个平衡特征效用与神经元容量的分析模型解释这一极化现象。

Comments Project page and code: https://avdravid.github.io/rosetta-neuron-scaling/

URL PDF HTML 收藏
2604.18572 2026-06-03 cs.CV cs.AI cs.LG

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

回到柏拉图的洞穴:大规模检验跨模态表示收敛性

A. Sophia Koepke, Daniil Zverev, Shiry Ginosar, Alexei A. Efros

机构 * UC Berkeley(伯克利大学) Technical University Munich, MCML(慕尼黑技术大学) University of Tübingen, Tübingen AI Center(图宾根大学) Toyota Technical Institute at Chicago(芝加哥丰田技术研究所)

AI总结 本文通过大规模数据集实验,质疑了柏拉图表示假说中跨模态表示收敛的证据,发现对齐度随数据规模增大而显著下降,且仅反映粗粒度语义重叠。

Comments Project page: http://akoepke.github.io/cave_umwelten/

URL PDF HTML 收藏
2601.00090 2026-05-04 cs.CV cs.LG

It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models

迟到也不晚:通过噪声优化在训练好的扩散模型中恢复崩溃

Anne Harrington, A. Sophia Koepke, Shyamgopal Karthik, Trevor Darrell, Alexei A. Efros

机构 * UC Berkeley(加州大学伯克利分校) University of Tübingen, Tübingen AI Center(图宾根大学,图宾根人工智能中心) TU Munich, MCML(慕尼黑工业大学,MCML)

AI总结 本文通过噪声优化提升扩散模型生成多样性,减少模式崩溃,同时保持基础模型的保真度。

Comments CVPR 2026. Project page at https://akoepke.github.io/divgen/index.html

URL PDF HTML 收藏
2511.10721 2025-11-17 cs.CV cs.LG

Fast Data Attribution for Text-to-Image Models

Sheng-Yu Wang, Aaron Hertzmann, Alexei A Efros, Richard Zhang, Jun-Yan Zhu

机构 * Carnegie Mellon University(卡内基梅隆大学) Adobe Research(Adobe研究) UC Berkeley(伯克利大学)

Comments NeurIPS 2025 camera ready. Project page: https://peterwang512.github.io/FastGDA

URL PDF HTML 收藏
2506.08010 2025-10-28 cs.CV cs.AI

Vision Transformers Don't Need Trained Registers

Nick Jiang, Amil Dravid, Alexei Efros, Yossi Gandelsman

机构 * UC Berkeley(伯克利大学)

Comments Project page and code: https://avdravid.github.io/test-time-registers. Accepted to NeurIPS '25 (spotlight)

URL PDF HTML 收藏
2507.02864 2025-09-23 cs.RO cs.CV

The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio

Renhao Wang, Haoran Geng, Tingle Li, Feishi Wang, Gopala Anumanchipalli, Trevor Darrell, Boyi Li, Pieter Abbeel, Jitendra Malik, Alexei A. Efros

机构 * University of California, Berkeley(加州大学伯克利分校)

Comments Conference on Robot Learning 2025

URL PDF HTML 收藏
2410.04201 2025-05-27 cs.CV

IT$^3$: Idempotent Test-Time Training

Nikita Durasov, Assaf Shocher, Doruk Oner, Gal Chechik, Alexei A. Efros, Pascal Fua

Comments Accepted at ICML 2025

URL PDF HTML 收藏
2410.18082 2025-05-12 cs.LG

Prioritized Generative Replay

Renhao Wang, Kevin Frans, Pieter Abbeel, Sergey Levine, Alexei A. Efros

机构 * Department of Electrical Engineering and Computer Science(电气工程与计算机科学系) University of California, Berkeley(加州大学伯克利分校)

Comments Project page available at: https://pgenreplay.github.io

URL PDF HTML 收藏
2401.14391 2025-04-11 cs.CV

Rethinking Patch Dependence for Masked Autoencoders

Letian Fu, Long Lian, Renhao Wang, Baifeng Shi, Xudong Wang, Adam Yala, Trevor Darrell, Alexei A. Efros, Ken Goldberg

Comments Transactions on Machine Learning Research (TMLR) 2025

URL PDF HTML 收藏
2503.21770 2025-03-28 cs.CV

Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting

Anand Bhattad, Konpat Preechakul, Alexei A. Efros

Comments project page: https://visualjenga.github.io/

URL PDF HTML 收藏
2406.09408 2025-02-21 cs.CV cs.LG

Data Attribution for Text-to-Image Models by Unlearning Synthesized Images

Sheng-Yu Wang, Aaron Hertzmann, Alexei A. Efros, Jun-Yan Zhu, Richard Zhang

Comments NeurIPS 2024 camera ready version. Project page: https://peterwang512.github.io/AttributeByUnlearning Code: https://github.com/PeterWang512/AttributeByUnlearning

URL PDF HTML 收藏
2406.04341 2025-02-14 cs.CV

Interpreting the Second-Order Effects of Neurons in CLIP

Yossi Gandelsman, Alexei A. Efros, Jacob Steinhardt

Comments project page: https://yossigandelsman.github.io/clip_neurons/index.html

URL PDF HTML 收藏
2501.12390 2025-01-23 cs.CV

GPS as a Control Signal for Image Generation

Chao Feng, Ziyang Chen, Aleksander Holynski, Alexei A. Efros, Andrew Owens

Comments Project page: https://cfeng16.github.io/gps-gen/

URL PDF HTML 收藏
2501.12387 2025-01-22 cs.CV

Continuous 3D Perception Model with Persistent State

Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A. Efros, Angjoo Kanazawa

URL PDF HTML 收藏
2307.05014 2025-01-07 cs.CV cs.LG

Test-Time Training on Video Streams

Renhao Wang, Yu Sun, Arnuv Tandon, Yossi Gandelsman, Xinlei Chen, Alexei A. Efros, Xiaolong Wang

Comments Project website with videos, dataset and code: https://test-time-training.github.io/video

URL PDF HTML 收藏
2401.10889 2024-12-23 cs.CV cs.AI

Synthesizing Moving People with 3D Control

Boyi Li, Junming Chen, Jathushan Rajasegaran, Yossi Gandelsman, Alexei A. Efros, Jitendra Malik

URL PDF HTML 收藏
2406.09417 2024-12-12 cs.CV cs.GR cs.LG

Rethinking Score Distillation as a Bridge Between Image Distributions

David McAllister, Songwei Ge, Jia-Bin Huang, David W. Jacobs, Alexei A. Efros, Aleksander Holynski, Angjoo Kanazawa

Comments NeurIPS 2024. Project webpage: https://sds-bridge.github.io/

URL PDF HTML 收藏
2405.10320 2024-12-11 cs.CV

Toon3D: Seeing Cartoons from New Perspectives

Ethan Weber, Riley Peterlinz, Rohan Mathur, Frederik Warburg, Alexei A. Efros, Angjoo Kanazawa

Comments Please see our project page: https://toon3d.studio

URL PDF HTML 收藏
2406.09413 2024-11-25 cs.CV cs.GR cs.LG

Interpreting the Weight Space of Customized Diffusion Models

Amil Dravid, Yossi Gandelsman, Kuan-Chieh Wang, Rameen Abdal, Gordon Wetzstein, Alexei A. Efros, Kfir Aberman

Comments Project Page: https://snap-research.github.io/weights2weights

URL PDF HTML 收藏
2409.05862 2024-09-11 cs.CV

Evaluating Multiview Object Consistency in Humans and Image Models

Tyler Bonnen, Stephanie Fu, Yutong Bai, Thomas O'Connell, Yoni Friedman, Nancy Kanwisher, Joshua B. Tenenbaum, Alexei A. Efros

Comments Project page: https://tzler.github.io/MOCHI/ Code: https://github.com/tzler/mochi_code Huggingface dataset: https://huggingface.co/datasets/tzler/MOCHI

URL PDF HTML 收藏
2408.02752 2024-08-07 cs.CV cs.AI

Diffusion Models as Data Mining Tools

Ioannis Siglidis, Aleksander Holynski, Alexei A. Efros, Mathieu Aubry, Shiry Ginosar

Comments Project Page: https://diff-mining.github.io/ Accepted in ECCV 2024

URL PDF HTML 收藏
2312.07504 2024-07-31 cs.CV

COLMAP-Free 3D Gaussian Splatting

Yang Fu, Sifei Liu, Amey Kulkarni, Jan Kautz, Alexei A. Efros, Xiaolong Wang

Comments Project Page: https://oasisyang.github.io/colmap-free-3dgs

URL PDF HTML 收藏
2310.05916 2024-04-01 cs.CV cs.AI

Interpreting CLIP's Image Representation via Text-Based Decomposition

Yossi Gandelsman, Alexei A. Efros, Jacob Steinhardt

Comments Project page and code: https://yossigandelsman.github.io/clip_decomposition/

URL PDF HTML 收藏
2402.16936 2024-02-28 cs.CV cs.LG

Disentangled 3D Scene Generation with Layout Learning

Dave Epstein, Ben Poole, Ben Mildenhall, Alexei A. Efros, Aleksander Holynski

URL PDF HTML 收藏
2307.05473 2023-12-27 cs.CV

Differentiable Blocks World: Qualitative 3D Decomposition by Rendering Primitives

Tom Monnier, Jake Austin, Angjoo Kanazawa, Alexei A. Efros, Mathieu Aubry

Comments Project webpage with code and videos: https://www.tmonnier.com/DBW. V2 update includes comparisons based on NeuS, hyperparameter analysis and failure cases

URL PDF HTML 收藏
2312.00785 2023-12-04 cs.CV

Sequential Modeling Enables Scalable Learning for Large Vision Models

Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan Yuille, Trevor Darrell, Jitendra Malik, Alexei A Efros

Comments Website: https://yutongbai.com/lvm.html

URL PDF HTML 收藏
2311.01462 2023-11-03 cs.CV cs.LG

Idempotent Generative Network

Assaf Shocher, Amil Dravid, Yossi Gandelsman, Inbar Mosseri, Michael Rubinstein, Alexei A. Efros

URL PDF HTML 收藏