Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments ECCV 2024; Project page at https://idea2img.github.io/
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments ECCV 2024; Project page at https://idea2img.github.io/
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
Comments ECCV24
专题命中 多模态生成 :image-text(abstract);any-to-any(abstract);分类 cs.CV
Comments 14 pages, 9 figures
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
专题命中 多模态生成 :multi-modal(abstract);MLLM(abstract);分类 cs.CV
Comments CVPR 2024
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments We open-source our data, model, and code at: https://github.com/linzhiqiu/t2v_metrics ; Project page: https://linzhiqiu.github.io/papers/vqascore
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CL
Comments Taiyi-Diffusion-XL Tech Report
专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV
Comments Codes and models: \url{https://github.com/FoundationVision/LlamaGen}
专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Technical Report; Dataset released in https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments The First review of State Space Model (SSM)/Mamba and their applications in artificial intelligence, 33 pages
专题命中 多模态生成 :cross-modal(abstract);image-text(abstract);分类 cs.CV
Comments Accepted by AAAI2024
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 31 pages, 22 figures, 30M PDF file size; Project Page: https://llava-vl.github.io/llava-interactive/
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments 10 pages
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Work in progress. Code is available at https://github.com/ZiyuGuo99/Point-Bind_Point-LLM
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments Technical Report; Project released at: https://github.com/AILab-CVC/SEED
专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments 13 pages
专题命中 多模态生成 :multi-modal(abstract);image-text(abstract);分类 cs.CV
专题命中 多模态生成 :multi-modal(abstract);image-text(abstract);分类 cs.CV
Comments Accepted by CVPR 2022, https://github.com/drboog/Lafite
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments WACV-2020 (Accepted)
专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV
专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
专题命中 多模态生成 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI
Comments Multimodal, Visual Question Answering, Vision and Language
专题命中 多模态生成 :multi-modal(abstract,comments);分类 cs.CV、cs.AI、cs.MM
Comments CVPR 2021. Code: https://github.com/weihaox/TediGAN Data: https://github.com/weihaox/Multi-Modal-CelebA-HQ Video: https://youtu.be/L8Na2f5viAM
G-CARL:面向患者的医学报告解读的基于 grounded 核对清单对齐的奖励学习
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本文针对现有医学视觉-语言任务无法兼顾医学事实性与患者语境沟通的问题,提出PMRI任务及G-CARL框架,构建MMedReport基准,实验证实其解读更贴合患者需求。
面向纳米光子学的大语言模型综合综述:从代理建模到自主设计
专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract)
AI总结 该综述探讨大语言模型(LLMs)如何通过语义接口、代码生成及工具编排改进纳米光子学工作流,梳理相关方法的两类模式及跨学科应用,展望具备物理感知的多模态基础模型,推动AI从被动工具向主动科研合作者转变。
Comments Accepted for publication in Advanced Photonics
AutoDesign:面向长视距智能体设计的元工具优化
机构 * Meituan(美团) ; MBZUAI(Mohamed bin Zayed University of Artificial Intelligence) ; Huazhong University of Science and Technology(华中科技大学) ; Peking University(北京大学) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong(香港中文大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 AutoDesign是符合人类设计先验的元工具优化框架,以论文转海报生成任务为实例,在PosterBench上性能优于Claude Design,集成其学习的DesignHarness可提升代码智能体性能,且获人类最高偏好。
Comments Tech Report. Code at: https://github.com/Yaxin9Luo/AutoDesign
CURV:通过课程可视化接地推理增强图表理解
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 针对多模态大语言模型视觉接地与推理不足的问题,提出CURV课程学习框架,结合CCQA数据集,在图表问答任务中实现显著性能提升并具备良好泛化性。