arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

THPL:面向RAS中虹鳟投喂管理的视觉到语言决策支持框架

THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS

Meng Liang, Guanbo Feng, Haozhuang Chi, Shilong Zhao, Zhixin Xiong, Yuhang He, Wenfeng Han, Tianhao Zhao, Zhihong Ma, Ying Liu

arXiv 2610.02378首次发表:更新:

发表机构

Zhejiang University; Nanyang Technological University; Dalian Ocean University; Nanjing Agricultural University(浙江大学; 南洋理工大学; 大连海洋大学; 南京农业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出THPL框架,结合轨迹提取、分层行为编码和反事实多模态优化,将视觉信息转化为可解释的投喂决策,显著提升准确率与因果一致性。

AI 中文摘要

在循环水养殖系统(RAS)中,精准投喂对于降低成本和提高鱼类福利至关重要。然而,现有方法缺乏鱼类行为与管理知识之间的认知对齐,阻碍了其转化为可执行、可解释的投喂决策。为解决这一问题,我们提出了THPL,一个专为RAS中虹鳟(Oncorhynchus mykiss)设计的生成式投喂决策框架。首先,Fishsort提取轨迹以建立量化投喂强度的活动系数(AC)。其次,分层行为编码器(HBE)利用时间Transformer和集合Transformer建模个体时间进程和群体动态,将轨迹张量转化为显式物理和隐式软令牌的双重证据表示。最后,这些令牌与环境参数、元数据和专家规则相结合,通过LoRA对大型语言模型(LLM)进行微调,随后采用反事实多模态直接偏好优化(mDPO)以强化因果推理。结果表明,AC与专家标注的投喂强度呈统计学显著的单调正相关(Spearman ρ = 0.925,p < 0.001)。消融实验显示,决策准确率从33.33%(仅文本基线)提升至具有双重证据令牌时的93.33%,证实了连续时空令牌为LLM提供了必要的物理基础。与标准LoRA相比,反事实mDPO将决策准确率从93.33%提升至96.67%,将METEOR从58.10%提升至85.30%,将Self-BLEU-2从58.79%降低至52.88%,并将Distinct-3从6.68%提升至7.81%,抑制了模板化和执行偏差,同时增强了因果一致性和操作安全性。总体而言,通过将连续运动学与LLM推理相结合,本研究为精准水产养殖提供了一种新颖的决策支持范式。

英文摘要

In Recirculating Aquaculture Systems (RAS), precision feeding is critical for minimizing costs and improving fish welfare. However, existing methods lack cognitive alignment between fish behaviors and management knowledge, impeding translation into executable, interpretable feeding decisions. To address this, we propose THPL, a generative feeding decision framework tailored for rainbow trout (Oncorhynchus mykiss) in RAS. First, Fishsort extracts trajectories to establish an Activity Coefficient (AC) quantifying feeding intensity. Second, a Hierarchical Behavior Encoder (HBE) models individual temporal progression and collective dynamics using Temporal and Set Transformers, transforming trajectory tensors into dual-evidence representations of explicit physical and implicit soft tokens. Finally, these tokens are integrated with environmental parameters, metadata, and expert rules to fine-tune an LLM via LoRA, followed by counterfactual multimodal Direct Preference Optimization (mDPO) to reinforce causal reasoning. Results show that AC exhibits a statistically significant monotonic positive correlation with expert-annotated feeding intensity (Spearman $ρ= 0.925$, $p < 0.001$). Ablations indicate that decision accuracy improves from 33.33% (text-only baseline) to 93.33% with dual-evidence tokens, confirming that continuous spatiotemporal tokens provide necessary physical grounding for LLMs. Compared with standard LoRA, counterfactual mDPO elevates decision accuracy from 93.33% to 96.67%, advances METEOR from 58.10% to 85.30%, reduces Self-BLEU-2 from 58.79% to 52.88%, and increases Distinct-3 from 6.68% to 7.81%, suppressing templating and actuation biases while reinforcing causal consistency and operational safety. Overall, by integrating continuous kinematics with LLM reasoning, this study provides a novel decision support paradigm for precision aquaculture.

CommentsMeng Liang and Guanbo Feng contributed equally. Corresponding authors: Zhihong Ma and Ying Liu. 50 pages, 10 figures, 3 tables. Supplementary video: https://youtu.be/Tg2Qk7m46-A

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑