Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
机构 * Google DeepMind(谷歌DeepMind)
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI
Journal ref International Conference on Machine Learning, 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Google DeepMind(谷歌DeepMind)
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI
Journal ref International Conference on Machine Learning, 2025
专题命中 多模态生成 :multimodal(abstract);分类 cs.AI
机构 * Korea University(韩国大学) ; Gauss Labs Inc.(Gauss实验室)
专题命中 多模态生成 :multi-modal(abstract)