arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 3490 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3490 篇

2511.17358 2025-11-24 cs.CL 50%

Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding

不要学习,而是依托:自然语言推理与视觉依托的案例

Daniil Ignatev, Ayman Santeer, Albert Gatt, Denis Paperno

机构 * Utrecht University(乌特勒支大学)

专题命中 文生图 :text-to-image(abstract)

AI总结 本文提出了一种基于视觉依托的零样本自然语言推理方法,通过生成视觉表示并比较与假设的相似度,实现高精度推理,展示了对文本偏见的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16038 2025-11-21 cs.HC 50%

Panel-by-Panel Souls: A Performative Workflow for Expressive Faces in AI-Assisted Manga Creation

面板式灵魂:一种用于AI辅助漫画创作中表现力面部的表演性工作流程

Qing Zhang, Jing Huang, Yifei Huang, Jun Rekimoto

专题命中 文生图 :text-to-image(abstract)

AI总结 本文提出了一种交互式工作流程,用于AI辅助漫画创作中表现力面部的生成,通过双混合流程实现艺术家意图与AI执行的高效衔接。

Comments NeurIPS 2025 Creative AI Track, The Thirty-Ninth Annual Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10268 2025-11-14 cs.AI 50%

Causal-HalBench: Uncovering LVLMs Object Hallucinations Through Causal Intervention

Zhe Xu, Zhicai Wang, Junkang Wu, Jinda Lu, Xiang Wang

专题命中 文生图 :text-to-image(abstract)

Comments accepted for publication in the Association for the Advancement of Artificial Intelligence (AAAI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12170 2025-11-03 cs.NE 50%

Language Model Crossover: Variation through Few-Shot Prompting

Elliot Meyerson, Mark J. Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K. Hoover, Joel Lehman

专题命中 文生图 :text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02631 2025-10-31 cs.HC cs.AI 50%

Reflection on Data Storytelling Tools in the Generative AI Era from the Human-AI Collaboration Perspective

Haotian Li, Yun Wang, Huamin Qu

机构 * Microsoft Research Asia(微软亚洲研究院) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 文生图 :text-to-image(abstract)

Comments This paper is a sequel to the CHI 24 paper "Where Are We So Far? Understanding Data Storytelling Tools from the Perspective of Human-AI Collaboration (https://doi.org/10.1145/3613904.3642726), aiming to refresh our understanding with the latest advancements. It is accepted at IEEE VIS 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15551 2025-10-27 cs.LG 50%

PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image Detectors

Sepehr Dehdashtian, Mashrur M. Morshed, Jacob H. Seidman, Gaurav Bharaj, Vishnu Naresh Boddeti

机构 * Michigan State University(密歇根州立大学) Reality Defender

专题命中 文生图 :text-to-image(abstract)

Comments Accepted as NeurIPS 2025 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18450 2025-09-30 cs.CL 50%

BRIT: Bidirectional Retrieval over Unified Image-Text Graph

Ainulla Khan, Yamada Moyuru, Srinidhi Akella

机构 * Fujitsu Research India(富士通印度研究)

专题命中 文生图 :text-to-image(abstract)

Comments Accepted in EMNLP-2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21874 2025-09-29 cs.LG 50%

Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models

Yifei Peng, Yaoli Liu, Enbo Xia, Yu Jin, Wang-Zhou Dai, Zhong Ren, Yao-Xiang Ding, Kun Zhou

机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) National Key Laboratory for Novel Software Technology(新型软件技术国家实验室)

专题命中 文生图 :text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15432 2025-09-22 cs.IR 50%

SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models

Thong Nguyen, Yibin Lei, Jia-Huei Ju, Andrew Yates

专题命中 文生图 :text-to-image(abstract)

Comments Accepted

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15154 2025-09-09 cs.AI cs.LG 50%

Online Prompt Pricing based on Combinatorial Multi-Armed Bandit and Hierarchical Stackelberg Game

Meiling Li, Hongrun Ren, Haixu Xiong, Zhenxing Qian, Xinpeng Zhang

专题命中 文生图 :text-to-image(abstract)

Comments The paper has been withdrawn by the authors because the current experimental results are not sufficiently reliable. Further optimization and refinement of the methodology are required before the work can be disseminated

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13287 2025-09-05 cs.LG 50%

PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMs

Xiaoyan Hu, Ho-fung Leung, Farzan Farnia

机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系) Independent Researcher(独立研究者)

专题命中 文生图 :text-to-image(abstract)

Comments accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05476 2025-08-08 eess.IV 50%

MM2CT: MR-to-CT translation for multi-modal image fusion with mamba

Chaohui Gong, Zhiying Wu, Zisheng Huang, Gaofeng Meng, Zhen Lei, Hongbin Liu

专题命中 文生图 :image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05432 2025-08-08 cs.AI cs.CY 50%

Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI

Krzysztof Janowicz, Zilong Liu, Gengchen Mai, Zhangyu Wang, Ivan Majic, Alexandra Fortacz, Grant McKenzie, Song Gao

机构 * University of Vienna(维也纳大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Maine(缅因大学) McGill University(麦吉尔大学) University of Wisconsin(威斯康星大学)

专题命中 文生图 :text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09509 2025-08-05 cs.LG 50%

EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling

Theodoros Kouzelis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis

机构 * Archimedes,Athena Reaserch Center, Greece(阿基米德、阿泰纳研究中心) National Technical University of Athens, Greece(雅典技术大学) University of Crete, Greece(克里特大学)

专题命中 文生图 :image synthesis(abstract)

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08983 2025-07-15 cs.LG cs.CR 50%

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models

Anshuman Suri, Harsh Chaudhari, Yuefeng Peng, Ali Naseh, Amir Houmansadr, Alina Oprea

机构 * Northeastern University(东北大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 文生图 :text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03015 2025-07-11 cs.CL cs.CY cs.LG 50%

Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench

Felix Friedrich, Thiemo Ganesha Welsch, Manuel Brack, Patrick Schramowski, Kristian Kersting

机构 * TU Darmstadt(图宾根大学) DFKI(德意志联邦人工智能研究中心) CERTAIN(CERTAIN公司) Centre for Cognitive Science, Darmstadt(达姆施塔特认知科学中心)

专题命中 文生图 :text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18002 2025-06-24 quant-ph 50%

A Survey of Quantum Generative Adversarial Networks: Architectures, Use Cases, and Real-World Implementations

Mujahidul Islam, Serkan Turkeli, Fatih Ozaydin

专题命中 文生图 :image synthesis(abstract)

Comments 30 pages, comments and suggestions are welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04599 2025-06-23 cs.HC 50%

Fuzzy Linkography: Automatic Graphical Summarization of Creative Activity Traces

Amy Smith, Barrett R. Anderson, Jasmine Tan Otto, Isaac Karth, Yuqian Sun, John Joon Young Chung, Melissa Roemmele, Max Kreminski

专题命中 文生图 :text-to-image(abstract)

Comments ACM C&C 2025. Code available at https://github.com/mkremins/fuzzy-linkography

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15140 2025-06-19 astro-ph.GA astro-ph.SR 50%

ALMASOP. A Rotating Feature Rich in Complex Organic Molecules in a Protostellar Core

Shih-Ying Hsu, Chin-Fei Lee, Doug Johnstone, Sheng-Yuan Liu, Tie Liu, Leonardo Bronfman, Huei-Ru Vivien Chen, Somnath Dutta, David J. Eden, Naomi Hirano, Mika Juvela, Kee-Tae Kim, Yi-Jehng Kuan, Woojin Kwon, Chang Won Lee, Jeong-Eun Lee, Shanghuo Li, Sheng-Jun Lin, Chun-Fan Liu, Xunchuan Liu, J. A. López-Vázquez, Qiuyi Luo, Mark G. Rawlings, Dipen Sahu, Patricio Sanhueza, Hsien Shang, Kenichi Tatematsu, Yao-Lun Yang

专题命中 文生图 :image synthesis(abstract)

Comments 19 pages, 10+1 figures, accepted by ApJ

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10949 2025-06-17 cs.CR cs.AI 50%

Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors

Chen Yueh-Han, Nitish Joshi, Yulin Chen, Maksym Andriushchenko, Rico Angell, He He

机构 * New York University(纽约大学) École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院)

专题命中 文生图 :text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06506 2025-06-10 cs.CL 50%

Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes

Kshitish Ghate, Tessa Charlesworth, Mona Diab, Aylin Caliskan

机构 * Carnegie Mellon University(卡内基梅隆大学) Northwestern University(西北大学) University of Washington(华盛顿大学)

专题命中 文生图 :text-to-image(abstract)

Comments Accepted to ACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08720 2025-06-02 cs.CY cs.AI 50%

AI for Just Work: Constructing Diverse Imaginations of AI beyond "Replacing Humans"

Weina Jin, Nicholas Vincent, Ghassan Hamarneh

机构 * Simon Fraser University(西蒙弗雷泽大学)

专题命中 文生图 :image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14758 2025-05-22 cs.CY cs.AI 50%

Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art

Alayt Issak, Uttkarsh Narayan, Ramya Srinivasan, Erica Kleinman, Casper Harteveld

机构 * Northeastern University(东北大学) Fujitsu Research of America(富士通美国研究)

专题命中 文生图 :text-to-image(abstract)

Journal ref Proceedings of the 17th Conference on Creativity \& Cognition (C\&C), June 23-25, 2025, Virtual, United Kingdom

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04600 2025-05-21 cs.CY 50%

Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm

Laura Wagner, Eva Cetinic

专题命中 文生图 :text-to-image(abstract)

Comments 28 pages, 9 figures, 2 interactive figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13791 2025-04-21 cs.SD cs.AI eess.AS 50%

Collective Learning Mechanism based Optimal Transport Generative Adversarial Network for Non-parallel Voice Conversion

Sandipan Dhar, Md. Tousin Akhter, Nanda Dulal Jana, Swagatam Das

机构 * Department of Computer Science and Engineering, National Institute of Technology Durgapur, India(印度德瓦格普国家理工学院计算机科学与工程系) Department of Computer Science and Engineering, Indian Institute of Technology Bombay, India(印度孟买印度理工学院计算机科学与工程系) Electronics and Communication Sciences Unit, Indian Statistical Institute, Kolkata, India(印度统计研究所加尔各答电子与通信科学单元)

专题命中 文生图 :image synthesis(abstract)

Comments 7 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16376 2025-04-03 cs.AI cs.HC 50%

Beyond Text-to-Text: An Overview of Multimodal and Generative Artificial Intelligence for Education Using Topic Modeling

Ville Heilala, Roberto Araya, Raija Hämäläinen

专题命中 文生图 :text-to-image(abstract)

Journal ref Proceedings of the 40th ACM/SIGAPP Symposium on Applied Computing (SAC'25), March 31--April 4, 2025, Catania, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14541 2025-03-17 cs.LG 50%

Why LLMs Are Bad at Synthetic Table Generation (and what to do about it)

Shengzhe Xu, Cho-Ting Lee, Mandar Sharma, Raquib Bin Yousuf, Nikhil Muralidhar, Naren Ramakrishnan

专题命中 文生图 :image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02874 2025-03-05 cs.HC 50%

Prompting Generative AI with Interaction-Augmented Instructions

Leixian Shen, Haotian Li, Yifang Wang, Xing Xie, Huamin Qu

专题命中 文生图 :text-to-image(abstract)

Comments accepted to CHI LBW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20459 2025-03-03 cs.CY 50%

Broken Letters, Broken Narratives: A Case Study on Arabic Script in DALL-E 3

Arshia Sobhan, Philippe Pasquier, Gabriela Aceves Sepulveda

专题命中 文生图 :text-to-image(abstract)

Comments 8 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18576 2025-02-27 cs.HC cs.CY 50%

Investigating Youth AI Auditing

Jaemarie Solyst, Cindy Peng, Wesley Hanwen Deng, Praneetha Pratapa, Jessica Hammer, Amy Ogan, Jason Hong, Motahhare Eslami

专题命中 文生图 :text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏