arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IROH:基于多阶段混合检索与理由蒸馏LLM裁判的幽默洞察排名——面向JOKER 2026任务1英语

IROH: Insightful Ranking Of Humor using Multi-Stage Hybrid Retrieval with Rationale-Distilled LLM Judges for JOKER 2026 Track Task 1 English

Ana-Maria Luisa Mocanu, Sebastian Mocanu, Ciprian-Octavian Truica, Elena-Simona Apostol

arXiv 2609.15618首次发表:更新:

发表机构

National University of Science and Technology POLITEHNICA Bucharest; Academy of Romanian Scientists(布加勒斯特国立科技大学POLITEHNICA; 罗马尼亚科学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VANGUARD团队提出IROH三阶段检索系统,结合混合检索、交叉编码器重排序和理由蒸馏LLM裁判,在JOKER 2026任务1中以0.6347 MAP夺冠,并揭示理由蒸馏裁判是排名质量的关键。

AI 中文摘要

我们的团队VANGUARD提出了IROH(幽默洞察排名),一个用于CLEF 2026 JOKER任务1英语的三阶段检索系统,在排行榜上以0.6347 MAP的成绩获得第一名。我们的流程结合了混合稀疏-稠密检索、交叉编码器重排序,以及LoRA适配的大语言模型裁判集成。我们采用Gemma 4在两种提示策略(通用型和类型化)下生成查询感知的理由,并生成多达四种类型的结构化硬负样本用于训练数据构建。通过对三种交叉编码器架构、四种稠密嵌入器和八种裁判配置的消融实验,我们的主要发现有三点:(1)理由蒸馏裁判是排名质量的主要驱动因素,而将理由附加到第一阶段索引中的贡献可忽略不计;(2)结构化硬负样本在几乎所有配置中都会损害泛化能力,尽管它们抬高了局部验证分数;(3)在我们消融的组件中,更轻、校准更好的模型与更大的对应模型相比具有竞争力或更强,其中通用理由的Qwen2.5-7B裁判(0.6055 MAP)优于所有Gemma-4-31B配置,并且通用理由相对于类型化理由的优势几乎完全集中在较小的模型中。

英文摘要

Our team, VANGUARD, presents IROH (Insightful Ranking of Humor), a three-stage retrieval system for JOKER Task 1 English at CLEF 2026, achieving first place on the leaderboard with 0.6347 MAP. Our pipeline combines hybrid sparse-dense retrieval, cross-encoder reranking, and a LoRA-adapted Large Language Model judge ensemble. We employ Gemma 4 to generate query-aware rationales under two prompt strategies, generic and typed, and produce up to four types of structured hard negatives for training data construction. Through an ablation across three cross-encoder architectures, four dense embedders, and eight judge configurations, our key findings are threefold: (1) the rationale-distilled judge is the primary driver of ranking quality, whereas appending rationales to the first-stage index contributes negligibly; (2) structured hard negatives degrade generalisation in nearly all configurations despite inflating local validation scores; and (3) across the components we ablate, the lighter, better-calibrated model is competitive with or stronger than its larger counterpart, with the generic-rationale Qwen2.5-7B judge (0.6055 MAP) outperforming every Gemma-4-31B configuration, and the advantage of generic over typed rationales is concentrated almost entirely in the smaller model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑