Boosting Multi-Modal E-commerce Attribute Value Extraction via Unified Learning Scheme and Dynamic Range Minimization
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM
Comments accepted in AAAI-2023
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments AAAI-23 DeFactify 2 Workshop (1st Prize)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Information Systems Frontiers, 2022
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments 13 pages, 6 figures, published to AAAI
专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM
Comments 7 pages, 1 figure, to appear in MuSe 2022 (ACM MM2022 co-located workshop)
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL、cs.AI、cs.MM
Comments Accepted by COLING 2022
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、eess.AS
Comments 8 pages, 2 figures, to appear in MuSe 2022 (ACM MM2022 co-located workshop)
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted by AAAI2022
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、eess.AS
Comments In INTERSPEECH 2021
Journal ref Proc. Interspeech 2021, 2381-2385
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to ACM MM 2021
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract)
Comments Under review
专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)
Comments Accepted
Journal ref IEEE Transactions on Multimedia, 2021
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、eess.AS
Comments Camera-ready version for EACL 2021
专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Submitted to the 17th International Conference on Principles of Knowledge Representation and Reasoning (2020)
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)
Comments Code available at https://github.com/idearibosome/embracenet
Journal ref Information Fusion 51 (2019) 259-270
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments CVPR2019 accepted paper
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract)
Comments To appear at NIPS 2018; 9 pages with supplement
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract)
Comments Accepted in "2018 International Conference on Pattern Recognition"
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Published at AAAI-18, 7 pages
专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
VAD:为多模态在线策略蒸馏中的目标重建归因视觉证据
机构 * Shanghai Jiao Tong University(上海交通大学) ; Xiaohongshu Inc.(小红书公司) ; The Chinese University of Hong Kong(香港中文大学) ; Zhejiang University(浙江大学) ; Southeast University(东南大学)
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL
AI总结 该研究提出视觉归因蒸馏(VAD)算法,通过反事实目标重建分离教师校正中的视觉证据分量,在6个4B/9B规模的细粒度视觉基准上,性能优于现有蒸馏方法。
Comments The project is accessible at https://github.com/DeepExperience/VAD_Multimodal_OPD
停止思考,开始观察:通过无推理对齐实现多模态文档问答的高效训练后优化
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI
AI总结 研究多模态文档问答中高效训练后优化问题,提出感知 - RFT 框架,用组相对策略优化绕过推理令牌直接对齐视觉与基础输出,通过构建变体评估推理必要性,发现启用推理模型有优势,还识别基础差异,表明早期转换可减少训练数据并保持精度。
Comments Accepted at ICML 2026, Workshop on Efficient Multimodal Question Answering (EMM-QA)