用于多模态大语言模型的几何增强部分估计
Geometry-Enhanced Portion Estimation for Multimodal LLMs
浏览论文内容
中文总结 AI 辅助
研究针对多模态大语言模型部分估计弱的问题,提出用精确部分头部增强冻结商业MLLM的方法,该头部基于DINOv2主干构建,无需深度传感器和微调,在三个基准测试中降低部分误差,优于旗舰MLLM及原图像模型。
中文摘要 AI 辅助
基于图像的饮食评估有望取代成本高昂且容易产生偏差的人工回忆,但部分估计仍然是一个主要障碍。多模态大语言模型(MLLM)能够在无约束照片中零样本识别多种食物,但在部分估计方面表现较弱。我们通过一个精确的部分头部增强一个冻结的商业MLLM,该头部是基于冻结的DINOv2主干构建的小型几何增强网络,具有结构化的softmax所有权体积,使用MLLM的每种食物名称、边界框和密度范围,无需深度传感器,无需微调MLLM。在三个真实世界基准上进行全开放词汇评估,该头部相对于单独的MLLM将每种食物的部分误差降低了33%-41%,优于每个旗舰MLLM的直接估计,并在各自报告的指标上超过了每个基准最初发布的仅图像模型。
英文摘要
Image-based dietary assessment promises to replace costly, bias-prone manual recalls, but portion estimation remains a major blocker. Multimodal LLMs (MLLMs) recognize a wide range of foods zero-shot in uncontrolled photos, yet they are weak at portion estimation -- a gap we measure across the current frontier (Gemini, GPT, and Claude flagships alike). We present a method that enhances a frozen, commercial MLLM with an accurate portion head: a small geometry-enhanced network on a frozen DINOv2 backbone with a structured softmax-ownership volume, consuming the MLLM's per-food name, bounding box, and density range -- no depth sensor, no MLLM fine-tuning. Evaluated fully open-vocabulary on three real-world benchmarks, the head cuts per-food portion error by 33-41% relative to the MLLM alone, outperforms every flagship MLLM's direct estimates, and surpasses each benchmark's originally published image-only model at its own reported metric.