发表机构
Pusan National University(釜山国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CueRator提出智能体搜索符号决策规则以适配冻结多模态编码器,在开放词汇音视频事件感知上超越现有方法,显著缩小已见未见差距。
AI 中文摘要
大型语言模型智能体已被用于搜索符号结构,如程序和方程。我们提出CueRator,一个用于策略感知决策规则发现的智能体框架,它通过搜索将跨模态相似性转换为预测的决策规则,来适配冻结的对比式多模态编码器。我们在开放词汇的音频-视觉事件感知任务上验证了该方法,现有方法在适应性和对未见类别的泛化之间存在权衡:训练模块以泛化为代价获得适应性,而固定规则则相反。该框架将用于泛化的符号公式与一个轻量级策略配对,该策略为每个视频预测其参数以实现适应性。一个报告引导的多智能体循环离线发现公式,评估每个候选公式的表达上限以及训练策略能否实现它。在OV-AVEBench上,CueRator将总平均值从57.8提升至60.2,未见类别性能从55.8提升至59.9,优于现有最佳方法,并将已见-未见差距从7.1降至1.2。消融实验将增益归因于公式和策略,并表明两种反馈信号对有效搜索都是必要的。CueRator在两个额外的音频-视觉事件感知任务上也优于各自的基线,且仅重训练策略时,发现的规则在不同编码器上仍保持竞争力。代码可在该https URL获取。
英文摘要
Large language model agents have been used to search over symbolic structures such as programs and equations. We propose CueRator, an agentic framework for policy-aware decision-rule discovery, which adapts frozen contrastive multimodal encoders by searching for the decision rule that converts their cross-modal similarities into predictions. We validate it on open-vocabulary audio-visual event perception, where existing methods involve a trade-off between adaptivity and generalization to unseen categories: trained modules adapt at the cost of generalization, and fixed rules the reverse. The framework pairs a symbolic formulation for generalization with a lightweight policy that predicts its parameters per video for adaptivity. A report-guided multi-agent loop discovers the formulation offline, evaluating each candidate on its expressive ceiling and on whether a trained policy can realize it. On OV-AVEBench, CueRator raises the total average from 57.8 to 60.2 and unseen-category performance from 55.8 to 59.9 over the best existing method, reducing the seen-unseen gap from 7.1 to 1.2. Ablations attribute the gains to both the formulation and the policy and show that both feedback signals are necessary for effective search. CueRator also improves over the respective baselines on two further audio-visual event perception tasks, and the discovered rule remains competitive across encoders with only the policy retrained. Code is available at https://github.com/cvsp-lab/cuerator.
Comments40 pages, 18 figures