Visual Agreement Regularized Training for Multi-Modal Machine Translation
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL
Comments Accepted by AAAI 2020
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL
Comments Accepted by AAAI 2020
专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、eess.AS
Comments 8 pages
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments This paper also appears at ACL 2017
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、eess.AS
Comments ACL 2019
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Accepted at ACL 2019
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL
Comments Accepted to CVPR 2019
专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.CL
Comments Accepted to CVPR2019
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments IEEE Intelligence Systems. arXiv admin note: substantial text overlap with arXiv:1707.09538
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments Machine Learning for Health (ML4H) Workshop at NeurIPS 2018 arXiv:1811.07216
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI
Journal ref ICCV17: second workshop on Closing the Loop Between Vision and Language. Venice, Italy. 2017
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments EMNLP 2018
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL
Comments Published as a conference paper at ICLR 2018. 12 pages
专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV、cs.AI
Comments To appear in the 7th IEEE Symposium Series on Computational Intelligence (IEEE SSCI 2016), 8 pages, 6 figures. Minor revisions, in response to reviewers' comments
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Clarified main contributions, minor correction to Equation 8, additional comparisons in Table 2, added more related work
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 10 pages, ISERC 2013, IIE Annual Conference. Proceedings. Institute of Industrial Engineers
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Appears in NIPS 2016. The datasets introduced in this work will be gradually released on the project page
可复现的多模态可供性预测
机构 * Istituto Italiano di Tecnologia(意大利技术研究院) ; EPFL(洛桑联邦理工学院)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
AI总结 针对可供性预测方法评估难的问题,本文提出Affordance Sheet以规范任务表述等信息,实现可供性模型的可复现基准测试与现实场景可靠评估。
Comments Paper accepted to Workshop on Human-Centered Multimodal Intelligence in the Wild (HCMIW) in European Conference on Computer Vision (ECCV) 2026; 18 pages, 3 figures, 7 tables. Project webpage at https://apicis.github.io/aff-sheet
对话状态的多模态信号有多可靠?来自远程二元协作任务的证据
机构 * Colby College(科尔比学院)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL
AI总结 研究探讨从多模态行为测量对话状态的特征可靠性,提出三维评估框架用于视频会议二元对话特征评估,发现语言特征预测佳但跨任务通用性差,声学可靠性受说话者身份影响,交互特征是唯一可靠信号,强调相关评估对对话系统特征选择的重要性。
Comments Accepted, to appear in Proceedings of ACM International Conference on Multimodal Interaction 2026, 13 pages, 6 figures, 4 tables
GDP.pdf:针对专业PDF文档的基础多模态推理基准测试
机构 * Surge AI
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
AI总结 该研究针对专业PDF文档构建多模态推理基准测试GDP.pdf,由专业人员编写问题-文档对,通过严格筛选保留问题,有详细评分标准和能力分类。评估七个前沿模型,发现多数错误源于特定模式,公开了完整基准测试。
Comments 9 pages. v2: results updated to July 2026 leaderboard (17 models). Accepted at the 2nd Workshop on Knowledge-Intensive Multimodal Reasoning (KnowledgeMR) at CVPR 2026 (non-archival), under the former title "PDFParse: A Benchmark for Grounded Multimodal Reasoning over Professional PDF Documents". Dataset: https://huggingface.co/datasets/surgeai/GDP.pdf ; Code: https://github.com/surge-ai/gdp-pdf
MM-tau-p$^2$: 人格自适应提示用于双控制设置中多模态代理的鲁棒性评估
机构 * Sprinklr AI
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI;multimodal(comments)
AI总结 本文提出MM-tau-p$^2$基准,通过12个新指标评估多模态代理在双控制环境中的鲁棒性,结合用户输入解决查询,并展示即使使用前沿LLM,多模态鲁棒性等指标仍需考虑。
Comments A benchmark for evaluating multimodal both voice and text LLM agents in dualcontrol settings. We introduce persona adaptive prompting and 12 new metrics to assess robustness safety efficiency and recovery in customer support scenarios
BanglaMM-Disaster: 一种基于Transformer的多模态深度学习框架,用于孟加拉语多类灾害分类
机构 * Department of Computer Science and Engineering(计算机科学与工程系) ; Chittagong University of Engineering and Technology(奇特格隆工程与技术大学) ; Department of Electronics and Telecommunication Engineering(电子与电信工程系) ; Wilmington University(维明顿大学) ; College of Graduate and Professional Studies(研究生与专业研究学院) ; Trine University(特林大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
AI总结 BanglaMM-Disaster通过结合文本和视觉数据,提出了一种多模态深度学习框架,用于孟加拉语多类灾害分类,提升了灾害响应效率。
Comments Presented at the 2025 IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON), November 21-22, 2025, University of Rajshahi, Bangladesh. 6 pages, 9 disaster classes, multimodal dataset with 5,037 samples
机构 * Harvard University(哈佛大学) ; MIT Media Lab(麻省理工学院媒体实验室)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI
Comments Accepted at the Multimodal Algorithmic Reasoning (MAR) Workshop, NeurIPS 2025
机构 * Purdue University(普渡大学) ; Indiana University(印第安纳大学) ; Curtin University(Curtin大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments The extended full version of the accepted paper in 2025 IEEE BHI conference with title: Evaluating Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata. Dataset is available at: https://skynet.ecn.purdue.edu/~coburn6/ACETADA/
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV
Comments Accepted by ICPR Multi-Modal Visual Pattern Recognition Workshop
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI
Comments 29 pages, submited to conference, code available at: https://github.com/rxn4chemistry/multimodal-spectroscopic-dataset
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI
Journal ref S. González, A. K.-C. Yi, W.-T. Hsieh, W.-C. Chen, C.-L. Wang, V. C.-C. Wu, S.-H. Chang, Multi-modal heart failure risk estimation based on short ECG and sampled long-term HRV, Information Fusion 107 (2024) 102337