arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10610cs.CVcs.AI

WasteAssistant:用于智能垃圾分类与可持续管理的监管引导视觉问答框架

WasteAssistant: Regulation-Guided Visual Question Answering Framework for Intelligent Waste Segregation and Sustainable Managemen

Khush Kataruka, Harshit Maurya, Anuja Vats, Murari Mandal, Kiran Raja, Praveen Kumar Chandaliya

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对现有智能垃圾分类系统的不足,提出集成视觉语言模型和多模态大语言模型的框架,构建新数据集,经实验验证基于BLIP的模型效果更佳,提升了分类准确率,确保监管合规,推动多模态AI在垃圾管理中的应用。

中文摘要 AI 辅助

高效的垃圾分类对可持续城市管理和环境治理至关重要。现有自动化系统受单模态视觉处理、上下文理解不足和监管一致性弱的限制。为解决这些问题,我们提出一个语言引导的视觉人工智能框架,集成视觉语言模型和多模态大语言模型进行联合视觉语言推理。该框架实现了符合印度2016年固体废物管理规则的视觉问答范式。我们构建了一个新的WasteVQA数据集,包含21个垃圾类别的13500个问答对。实验表明,基于BLIP的模型BLEU分数为0.8291,BERTScore为0.9273,优于传统基于CNN的方法。这项工作提高了源头分类准确率,确保监管合规,并支持面向市政和市民的垃圾管理的可扩展部署,促进多模态人工智能在可持续城市基础设施中的应用。

英文摘要

Efficient waste segregation is critical for sustainable urban management and environmental governance. Existing automated systems are limited by single-modality visual processing, insufficient contextual understanding, and weak regulatory alignment. To address these issues, we propose a language-guided vision-AI framework that integrates vision-language models and multimodal large language models for joint visual-linguistic reasoning. This framework implements a visual question answering paradigm aligned with India's Solid Waste Management Rules 2016. We construct a new WasteVQA dataset with 13,500 question-answer pairs across 21 waste categories. Experiments show that the BLIP-based model achieves a BLEU score of 0.8291 and a BERTScore of 0.9273, outperforming traditional CNN-based methods. This work improves source-level segregation accuracy, ensures regulatory compliance, and supports scalable deployment for municipal and citizen-facing waste management, promoting multimodal AI in sustainable urban infrastructure. The source code and dataset are available at: https://github.com/Khushkataruka/WasteAssistant

发表机构

  • SVNIT(萨达尔·瓦拉巴伊·国家理工学院)
  • NTNU(挪威科技大学)
  • KIIT(卡林加工业技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑