arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2025-12-18 至 2025-12-18 共收录 55 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 2 篇

2508.06032 2025-12-18 cs.CV 70%

Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body Parts

学习3D纹理感知表示以解析多样化的人类服装和身体部位

Kiran Chhatre, Christopher Peters, Srikrishna Karanam

专题命中 文生图 :text-to-image(abstract);diffusion(abstract);分类 cs.CV

AI总结 Spectrum通过改进的3D纹理生成模型,实现了对多样化人类服装和身体部位的精细解析与分割。

Comments Association for the Advancement of Artificial Intelligence (AAAI) 2026, 14 pages, 11 figures. Webpage: https://s-pectrum.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15442 2025-12-18 cs.LG 67%

Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting

通过链式推理和任务指令提示降低版权侵权风险

Neeraj Sarna, Yuanyuan Li, Michael von Gablenz

机构 * Munich RE(慕尼黑RE)

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

AI总结 本文通过链式推理和任务指令提示结合负提示和提示重写,降低生成图像的版权侵权风险,并评估不同模型复杂度下的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 图像编辑 3 篇

2512.14423 2025-12-18 cs.CV 83%

The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy

注意力共享中的魔鬼:通过注意力协同提升复杂非刚性图像编辑的忠实性

Zhuo Chen, Fanyue Wei, Runze Xu, Jingjing Li, Lixin Duan, Angela Yao, Wen Li

机构 * University of Electronic Science and Technology of China(电子科技大学) National University of Singapore(新加坡国立大学)

专题命中 图像编辑 :image editing(title,abstract);diffusion(abstract);分类 cs.CV

AI总结 SynPS通过协同利用位置嵌入和语义信息,解决现有注意力共享机制中的注意力崩溃问题,提升复杂非刚性图像编辑的忠实性。

Comments Project page:https://synps26.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15603 2025-12-18 cs.CV 77%

Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition

Qwen-Image-Layered: 通过层分解实现内在可编辑性

Shengming Yin, Zekai Zhang, Zecheng Tang, Kaiyuan Gao, Xiao Xu, Kun Yan, Jiahao Li, Yilei Chen, Yuxiang Chen, Heung-Yeung Shum, Lionel M. Ni, Jingren Zhou, Junyang Lin, Chenfei Wu

机构 * HKUST(GZ)(香港科技大学(广州)) Alibaba(阿里巴巴)

专题命中 图像编辑 :image generation(abstract);diffusion(abstract);image editing(abstract);分类 cs.CV

AI总结 Qwen-Image-Layered通过层分解实现图像的内在可编辑性,提出端到端扩散模型及多阶段训练策略,提升图像分解质量与编辑一致性。

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15512 2025-12-18 cs.CV cs.MM 62%

VAAS: Vision-Attention Anomaly Scoring for Image Manipulation Detection in Digital Forensics

VAAS:面向数字取证中图像篡改检测的视觉-注意力异常评分

Opeyemi Bamigbade, Mark Scanlon, John Sheppard

机构 * Forensics and Security Research Group, South East Technological University(法医与安全研究组,南东技术大学) Security Research Group, School of Computer Science, University College Dublin(安全研究组,计算机科学学院,都柏林大学学院)

专题命中 图像编辑 :image generation(abstract);分类 cs.CV、cs.MM

AI总结 VAAS通过整合视觉注意力和片段一致性评分,提出一种双模块框架用于图像篡改检测,提升检测精度和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 扩散模型 42 篇

2411.05005 2025-12-18 cs.CV cs.LG cs.RO 83%

Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models

Diff-2-in-1:通过扩散模型弥合生成与密集感知的鸿沟

Shuhong Zheng, Zhipeng Bao, Ruoyu Zhao, Martial Hebert, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Carnegie Mellon University(卡内基梅隆大学) Tsinghua University(清华大学)

专题命中 扩散模型 :diffusion(title,abstract);image synthesis(abstract);分类 cs.CV

AI总结 Diff-2-in-1通过结合扩散模型的生成与感知能力,实现多模态数据生成与密集视觉感知的统一框架,提升视觉感知的判别能力。

Comments 26 pages, 14 figures

Journal ref ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15270 2025-12-18 eess.IV cs.CV cs.MM 81%

Generative Preprocessing for Image Compression with Pre-trained Diffusion Models

生成式预处理用于基于预训练扩散模型的图像压缩

Mengxi Guo, Shijie Zhao, Junlin Li, Li Zhang

机构 * Bytedance Inc.(字节跳动公司)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV、cs.MM

AI总结 本文提出利用预训练扩散模型进行图像压缩预处理,通过两阶段框架提升压缩效率与视觉质量。

Comments Accepted as a PAPER and for publication in the DCC 2026 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04050 2025-12-18 cs.GR cs.CV 81%

TerraFusion: Joint Generation of Terrain Geometry and Texture Using Latent Diffusion Models

TerraFusion:利用潜在扩散模型联合生成地形几何与纹理

Kazuki Higo, Toshiki Kanai, Yuki Endo, Yoshihiro Kanamori

机构 * University of Tsukuba(茨口大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV、cs.GR

AI总结 本文提出利用潜在扩散模型联合生成地形高度图和纹理,通过无监督训练和手绘草图控制实现真实感地形生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15022 2025-12-18 cs.CV 79%

LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion Models

LoRAverse:一种子模框架用于检索扩散模型的多样化适配器

Mert Sonmezer, Matthew Zheng, Pinar Yanardag

机构 * Middle East Technical University(美索不达米亚技术大学) Virginia Tech(弗吉尼亚理工大学)

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 LoRAverse通过子模框架解决从大量LoRA适配器中检索多样化模型的问题,提升扩散模型的个性化应用效果。

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 17879-17888

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14760 2025-12-18 cs.CV 79%

AquaDiff: Diffusion-Based Underwater Image Enhancement for Addressing Color Distortion

AquaDiff:基于扩散的水下图像增强以解决颜色失真

Afrah Shaahid, Muzammil Behzad

专题命中 扩散模型 :diffusion(title,abstract);分类 cs.CV

AI总结 AquaDiff通过融合色先验和条件扩散过程,有效纠正水下图像颜色失真,提升整体图像质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16090 2025-12-18 cond-mat.mtrl-sci 78%

External magnetic field suppression of carbon diffusion in iron

外部磁场抑制铁中碳扩散

Luke J. Wirth, Dallas R. Trinkle

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本研究通过DFT计算揭示外部磁场通过改变电子密度态结构抑制铁中碳扩散的机制。

Comments 6 pages, 5 figures, 10 pages supplemental material with 3 figures

Journal ref Phys. Rev. Lett. 135, 256302 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15346 2025-12-18 physics.bio-ph cond-mat.soft cond-mat.stat-mech 78%

Reconstruction of the Bacterial Flagellar Motor's Energy Landscape, Viscous Load, and Torque Generation Across Diffusion Regimes

重构细菌鞭毛马达在不同扩散 regime 中的能量景观、粘性负载和扭矩生成

N. J. Lopez-Alamilla, A. L. Nord, F. Pedaci, J. Palmeri, N. -O. Walliser

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本研究通过单分子数据重构细菌鞭毛马达的能量景观,揭示其扭矩-速度关系在临界倾斜处的非线性特性,并提出适用于其他循环步进马达的通用方法。

Comments 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15176 2025-12-18 cs.LG cs.AI 78%

DEER: Draft with Diffusion, Verify with Autoregressive Models

DEER:通过扩散模型草稿,通过自回归模型验证

Zicong Cheng, Guo-Wei Yang, Jia Li, Zhijie Deng, Meng-Hao Guo, Shi-Min Hu

机构 * Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 DEER通过扩散模型草稿和自回归模型验证,实现高效推测解码,大幅提升解码速度和草稿长度。

Comments Homepage : https://czc726.github.io/DEER/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15071 2025-12-18 q-fin.MF 78%

Arbitrage-Free Pricing with Diffusion-Dependent Jumps

无套利定价与扩散依赖跳跃

Hamza Virk, Yihren Wu, Majnu John

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出一种多类型跳跃扩散模型,通过无套利条件实现扩散依赖跳跃的定价

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14906 2025-12-18 astro-ph.GA astro-ph.HE 78%

The diffusion coefficient in the Large Magellanic Cloud

Large Magellanic Cloud 中的扩散系数

Javier Reynoso-Cordova, Daniele Gaggero, Marco Regis, Marco Taoso

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文通过数值模拟研究大麦哲伦星云中宇宙射线的扩散系数,利用射电观测数据推断出扩散系数D0,并为解释非热信号提供了工具。

Comments 13 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14765 2025-12-18 cs.LG cs.AI 78%

Guided Discrete Diffusion for Constraint Satisfaction Problems

引导离散扩散用于约束满足问题

Justin Jung

机构 * Justin Jung

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出了一种离散扩散引导方法,用于无监督解决数独谜题等约束满足问题。

Comments Originally published in Jan 2025 on the SpringtailAI Blog

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14725 2025-12-18 cs.LG cs.AI 78%

Generative Urban Flow Modeling: From Geometry to Airflow with Graph Diffusion

生成性城市风场建模:从几何到气流的图扩散

Francisco Giral, Álvaro Manzano, Ignacio Gómez, Petros Koumoutsakos, Soledad Le Clainche

机构 * Applied Mathematics Department, ETSIAE-School of Aeronautics, Universidad Politécnica de Madrid(应用数学系、ETSIAE航空学院、马德里理工大学) Microflown Technologies(Microflown技术公司) Computational Science and Engineering Laboratory, Harvard University(计算科学与工程实验室、哈佛大学)

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出了一种基于图扩散的生成模型,用于在非结构化网格上合成城市风场,通过结合层次图神经网络和分数扩散建模,实现对复杂几何形状的高效风场模拟与预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20681 2025-12-18 cs.DB 78%

Downsizing Diffusion Models for Cardinality Estimation

缩小扩散模型用于基数估计

Xinhe Mu, Zhaoqi Zhou, Zaijiu Shang, Chuan Zhou, Gang Fu, Guiying Yan, Guoliang Li, Zhiming Ma

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文提出ADC框架,通过混合架构实现轻量高效的扩散模型,用于高精度基数估计,在处理多向依赖数据时表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14671 2025-12-18 astro-ph.SR astro-ph.EP physics.flu-dyn 78%

Diffusion-Free Dynamics in Rotating Spherical Shell Convection Driven By Internal Heating and Cooling

无扩散动力学在由内部加热和冷却驱动的旋转球壳对流中

Neil T. Lewis, Tom Joshi-Hartley, Steven M. Tobias, Laura K. Currie, Matthew K. Browning

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 研究通过内部加热和冷却驱动旋转球壳对流,发现部分对流性质无扩散依赖,为恒星和巨行星对流模型发展提供新思路。

Comments Accepted for publication in ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.14960 2025-12-18 math.AP 78%

On $p$-Laplacian reaction-diffusion problems with dynamical boundary conditions in perforated media

关于带有动力边界条件的渗流问题的p-拉普拉斯反应扩散问题

María Anguiano

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 本文研究了在周期性孔洞领域中p-拉普拉斯反应扩散问题的均质化,考虑了动力边界条件的影响,并推导了非线性p-拉普拉斯方程的收敛结果。

Comments 20 pages. arXiv admin note: text overlap with arXiv:1912.02445, arXiv:1712.01183, arXiv:2004.06513

Journal ref Mediterranean Journal of Mathematics, Volume 20, 2023, article number 124

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.06513 2025-12-18 math.AP 78%

Reaction-diffusion equation on thin porous media

薄多孔介质上的反应扩散方程

María Anguiano

专题命中 扩散模型 :diffusion(title,abstract)

AI总结 研究薄多孔介质中反应扩散方程在ε趋近于0时的二维极限行为,揭示动力边界条件对源项的影响。

Comments 17 pages. arXiv admin note: text overlap with arXiv:1712.01183, arXiv:1912.02445

Journal ref Bulletin of the Malaysian Mathematical Sciences Society, Volume 44, 2021, pages 3089-3110

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15167 2025-12-18 math.OC 71%

Long-Run Average Reward Maximization of A Regulated Regime-Switching Diffusion Model

受监管的 regime-switching 扩散模型的长期平均奖励最大化

Lingjia Zeng, Manman Li

专题命中 扩散模型 :diffusion(title)

AI总结 本文提出在受监管的 regime-switching 环境中,通过综合控制框架和数值近似方法,解决保险公司的长期平均奖励最大化问题。

Comments 35 pages,1 table, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03338 2025-12-18 cs.CV cs.AI cs.CY 70%

Safer Prompts: Reducing Risks from Memorization in Visual Generative AI

更安全的提示:减少视觉生成AI中记忆化风险

Lena Reissinger, Yuanyuan Li, Anna-Carolina Haensch, Neeraj Sarna

机构 * Ludwig Maximilian University of Munich(慕尼黑路易斯·马克西米利安大学) Munich RE(慕尼黑RE)

专题命中 扩散模型 :image generation(abstract);diffusion(abstract);分类 cs.CV

AI总结 本文提出通过提示工程技术减少视觉生成AI中因记忆训练数据导致的安全风险,提升生成图像与训练数据的相似性降低,同时保持输出的相关性和美观性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15261 2025-12-18 cs.CV 70%

MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement

MMMamba: 一种多功能跨模态在上下文融合框架用于全色锐化和零样本图像增强

Yingying Wang, Xuanhua He, Chen Wu, Jialing Huang, Suiyun Zhang, Rui Liu, Xinghao Ding, Haoxuan Che

专题命中 扩散模型 :image generation(abstract);diffusion(abstract);分类 cs.CV

AI总结 MMMamba通过跨模态在上下文融合框架实现全色锐化和零样本图像增强,结合Mamba架构和多模态交错扫描机制,提升跨模态交互能力与计算效率。

Comments \link{Code}{https://github.com/Gracewangyy/MMMamba}

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15701 2025-12-18 cs.CV 57%

VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression

VLIC: 基于视觉-语言模型的人类对齐图像压缩感知判官

Kyle Sargent, Ruiqi Gao, Philipp Henzler, Charles Herrmann, Aleksander Holynski, Li Fei-Fei, Jiajun Wu, Jason Zhang

机构 * Stanford University(斯坦福大学) Google Research(谷歌研究) Google DeepMind(谷歌DeepMind)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 VLIC利用视觉-语言模型进行图像压缩,通过零样本推理实现人类感知对齐,提升压缩性能。

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15347 2025-12-18 cs.CV cs.LG 57%

Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models

扩展与剪枝:为生成模型中的有效GRPO最大化轨迹多样性

Shiran Ge, Chenyi Huang, Yuang Ai, Qihang Fan, Huaibo Huang, Ran He

机构 * MAIS & NLPR, Institute of Automation, CAS(自动化研究所信息与智能系统研究所与神经语言处理研究所) School of Artificial Intelligence, UCAS(人工智能学院)

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 本文提出Pro-GRPO框架,通过扩展与剪枝策略在生成模型中提升GRPO效果,减少计算开销并提高轨迹多样性。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15262 2025-12-18 eess.IV cs.MM 57%

Audio-Visual Cross-Modal Compression for Generative Face Video Coding

音频-视觉跨模态压缩用于生成式面部视频编码

Youmin Xu, Mengxi Guo, Shijie Zhao, Weiqi Li, Junlin Li, Li Zhang, Jian Zhang

专题命中 扩散模型 :diffusion(abstract);分类 cs.MM

AI总结 本文提出AVCC框架,通过联合压缩音频和视频流,提升生成式面部视频编码的率-失真性能。

Comments Accepted as a PAPER and for publication in the DCC 2026 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13421 2025-12-18 cs.CV 57%

RecTok: Reconstruction Distillation along Rectified Flow

RecTok: 修复流中的重建蒸馏

Qingyu Shi, Size Wu, Jinbin Bai, Kaidong Yu, Yujing Wang, Yunhai Tong, Xiangtai Li, Xuelong Li

机构 * Peking University(北京大学) Nanyang Technological University(南洋理工大学) TeleAI

专题命中 扩散模型 :diffusion(abstract);分类 cs.CV

AI总结 RecTok通过流语义蒸馏和重建-对齐蒸馏提升高维视觉分词器的重建和生成性能,实现更丰富的潜在空间结构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15392 2025-12-18 math.OC 50%

Asymptotic behaviour of stochastic inertial dynamics incorporating a Tikhonov regularization term

包含Tikhonov正则化项的随机惯性动力学的渐进行为

Chiara Schindler

专题命中 扩散模型 :diffusion(abstract)

AI总结 本文研究了包含Tikhonov正则化项的随机惯性动力学在噪声下的渐进行为,推导了能量函数的上界及收敛速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15371 2025-12-18 physics.soc-ph physics.app-ph 50%

Network localization governs social contagion dynamics with macro-level reinforcement

网络定位调控社会传染动力学中的宏观层面强化

Leyang Xue, Kai-Cheng Yang, Peng-Bi Cui, Zengru Di

专题命中 扩散模型 :diffusion(abstract)

AI总结 研究揭示网络定位通过调控临界点和强化阈值影响社会传染动力学,发现促进弱传染的网络导致较慢扩散,而抑制弱传染的网络促进更广泛传播。

Comments 12 pages, 4 figures, 67 references

详情

展开后加载摘要…

URL PDF HTML 收藏