发表机构
Sungkyunkwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究复现了RPC与LCF两种提升LLM推理可靠性的方法,在多任务域与多模型上压力测试,发现RPC优势未达显著且样本量增大后效果逆转,LCF效果弱且无统计显著性。
AI 中文摘要
我们独立复现了两种近期用于提升大语言模型(LLM)推理可靠性的方法,并在不同领域和模型上对其进行压力测试(RPC针对四个新任务域采用Qwen3-8B,LCF针对四个7-8B模型)。第一种方法RPC在推理阶段聚合token概率与自一致性;第二种方法LCF训练投影器,将隐藏状态拆分为“内容”与“逻辑”两部分,并编辑逻辑部分使其趋向有效区域。验证此类可靠性主张十分重要,因为原始评估由各方法的作者完成,从未被独立复现或在不同模型与域间进行压力测试,且LCF未提供公开代码。我们重新运行RPC已发表的路径聚合流程,重新实现LCF的投影器、对比与干预流程,随后将两种方法扩展至文本转SQL、法律抽取、谬误识别与先例分级任务,并直接探测LCF的表示。RPC在作者发布的推理路径上完全复现了原始网格结果;在四个新域中,其相对于自一致性的优势从未达到显著水平(持平或微小混合差异,配对p≥0.28);在我们调整预算的唯一域BIRD上,其优势随K增大而按预期增长,但当我们将样本量扩大至n=200时,最大差距(K=32时准确率提升2.5,p=0.16)逆转为-0.25。LCF的逻辑有效性方向真实但较弱(最佳子层的可分性为0.82,而语义属性对照的可分性为0.95);其唯一正向效果(Qwen3的ΔProb)不显著(p=0.56),同时会显著降低其余三个模型中两个的ΔProb。
英文摘要
We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B models). The first, RPC, aggregates token probabilities and self-consistency at inference; the second, LCF, trains projectors that split hidden states into "content" and "logic" and edits the logic part toward a valid region. Validating such reliability claims matters because the original evaluations are run by each method's own authors and were never independently reproduced or stress-tested across models and domains, and LCF shipped no public code. We re-run RPC's published-path aggregation and re-implement LCF's projector, contrastive, and intervention pipeline, then extend both to text-to-SQL, legal extraction, fallacy identification, and precedent grading, and probe LCF's representation directly. RPC reproduces the original grid exactly on the authors' released reasoning paths; on four new domains its edge over self-consistency is never significant (ties or small mixed differences, paired p >= 0.28), and on BIRD, the one domain where we vary the budget, the edge grows with K as predicted but its largest gap (+2.5 accuracy at K=32, p=0.16) reverses to -0.25 when we enlarge the sample to n=200. LCF's logic-validity direction is real but weak (0.82 separability at the single best sub-layer versus 0.95 for a semantic-attribute control); its one positive effect (Qwen3 $Δ$Prob) is not significant (p=0.56), while it significantly reduces $Δ$Prob on two of the other three models.
Comments16 pages, 3 figures, 9 tables. Code, data, and experiment logs: https://github.com/rabqatab/llm-reasoning-reliability-reproduction