arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BITEM在NTCIR-19 R2C2任务中的表现:从智能体RAG流水线信号预测置信度

BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals

Julien Knafou, Luc Mottin, Alexandre Flament, Paul van Rijen, Esteban Gaillac, Patrick Ruch

arXiv 2609.37993首次发表:更新:

发表机构

HES-SO; SIB Swiss Institute of Bioinformatics(瑞士西部应用科学与艺术大学; 瑞士生物信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BITEM团队用单一智能体RAG流水线参加NTCIR-19 R2C2任务,通过编排器计算置信度,提出accHMR指标,以准确率乘HMR更公平地评估置信度,并在检索和答案生成上取得领先排名。

AI 中文摘要

BITEM团队使用单一智能体流水线参加了NTCIR-19 R2C2任务的两个子任务。在该流水线中,一个模型在电影语料库上搜索、阅读并记录证据,同时一个编排器持有记录并裁定可以提交的内容。一条主张只有在蕴含级联根据其引用的段落进行检查后才会被接纳,而一个答案只有在足够多的已检查证据支持它时才会被发布。每个问题运行三到四次,每次检索都从语料库中剔除先前运行已见过的内容。随每个答案提交的置信度由编排器根据运行留下的信息计算,从不询问模型,模型也没有任何自我评分的方式。两次检索运行在22个队伍中分别排名第4和第5,合并各次运行的结果带来了0.0709的nDCG@20增益,且在多跳和重度后处理问题上增益最大,组织者将合并运行评为该领域的顶尖。25个答案运行中有16个基于这两次运行提供的段落构建,其中12个由其他团队提交。HMR奖励那些在回答正确时置信度高、回答错误时置信度低的系统。该流水线达到了0.9219的准确率,在25个队伍中排名第6,而随这些答案提交的置信度给出了0.4915的HMR,排名第13。基于这些相同的记录信号手工制定的几条规则,无需进一步调用模型或进一步检索,将准确率提升至0.9375(排名第5),HMR提升至0.6985(排名第9)。仅按HMR排名可能会奖励那些以低置信度错误回答的系统,因此我们提出了accHMR,即准确率乘以HMR,它按准确率比例报告奖励,修订后的规则本可得分0.6549(排名第5)。对于未来工作,在流水线已产生的数字上拟合模型,而不是手工编写此类规则,将是一个真正的进步。

英文摘要

The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an entailment cascade has checked it against the passage it cites, and an answer is released only once enough checked evidence stands behind it. Each question is run three or four times, every pass retrieving from a corpus stripped of what the earlier passes have already seen. The confidence filed with each answer is computed by the orchestrator from what the run leaves behind and is never asked of the model, which is offered no way to rate itself. The two retrieval runs placed 4th and 5th of 22, pooling the passes was worth 0.0709 nDCG@20, and the gain was largest on the multi-hop and post-processing-heavy questions, where the organisers rank the pooled run top of the field. Sixteen of the 25 answer runs were built on passages these two runs supplied, 12 of them filed by other teams. HMR rewards a system whose confidence is high where it answers right and low where it answers wrong. The pipeline reached an accuracy of 0.9219, 6th of 25, while the confidence filed with those answers gave an HMR of 0.4915, 13th. A few rules crafted over those same recorded signals, with no further model call and no further retrieval, raise that to an accuracy of 0.9375, 5th, and an HMR of 0.6985, 9th. Ranking on HMR alone can reward a system for answering wrongly with low confidence, so we propose accHMR, the accuracy multiplied by HMR, which reports the reward in proportion to the accuracy, and on which the revised rules would have scored 0.6549, 5th. For future work, fitting a model on the numbers the pipeline already produces, rather than writing such rules by hand, would be a real step forward.

Comments8 pages. Participant paper for the NTCIR-19 R2C2 task

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑