发言者很重要:针对意大利议会议事录的感知权威多视图检索增强生成
Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings
浏览论文内容
中文总结 AI 辅助
该研究针对意大利议会议事录的多视角访问难题,提出 ParliamentRAG 系统,通过主题依赖的权威模型解决 RAG 应用于议会文本的三类风险,在 15 个政策主题评估中表现优于 Google NotebookLM 在来源相关维度的性能。
中文摘要 AI 辅助
议会议事录是民主审议的主要记录,但其庞大的体量和碎片化的特点使公民、记者和研究者难以进行多视角访问。将检索增强生成(RAG)应用于议会 transcript 会带来三个特定风险:最频繁发言者的主导地位、无法根据主题专业知识对发言者进行加权、以及在政治敏感文本中出现引用归属错误。我们提出了 ParliamentRAG,这是一个针对意大利众议院的 RAG 系统,可共同解决这些风险。其核心贡献是一个依赖主题的权威模型,该模型将发言者的职业、教育背景和过往发言等可解释成分相结合,根据当前查询估计每位发言者的权威程度。针对用户查询,该系统会检索相关的发言片段,识别议会团体中与主题相关的专家,并生成综合他们观点的摘要,同时附上支撑性引语。我们通过结合自动指标和六位领域专家的盲法 A/B 人工评估的两级协议,针对 15 个政策主题将 ParliamentRAG 与 Google NotebookLM 进行了评估。该系统在各政治团体间的覆盖度更高(0.97 对 0.95),引语忠实度达到 1.00(对 0.95),在与来源相关的维度上获得了更强的专家偏好,而 NotebookLM 在散文导向的维度上仍表现更优。
英文摘要
Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make multi-perspective access difficult for citizens, journalists, and researchers. Applying Retrieval-Augmented Generation (RAG) to parliamentary transcripts introduces three specific risks: dominance of the most frequent speakers, inability to weight speakers according to topical expertise, and citation misattribution in politically sensitive text. We present ParliamentRAG, a RAG system for the Italian Chamber of Deputies that addresses these risks jointly. Its core contribution is a topic-dependent authority model that estimates each speaker's authority as a function of the current query, combining interpretable components such as profession, education, and previous interventions. Given a user query, the system retrieves relevant speech chunks, identifies topic-relevant experts across parliamentary groups, and generates a summary synthesizing their perspectives, accompanied by supporting quotations. ParliamentRAG is evaluated against Google NotebookLM on 15 policy topics via a two-level protocol combining automated metrics and blind A/B human evaluation by six domain experts. The system achieves higher coverage across political groups (0.97 vs. 0.95), perfect quotation faithfulness (1.00 vs. 0.95), and stronger expert preferences on source-related dimensions, while NotebookLM remains stronger on prose-oriented dimensions.
发表机构
- University of Milano-Bicocca(米兰比可卡大学)
机构由 AI 辅助整理,请以论文原文为准。