arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Pennsylvania(宾夕法尼亚大学)

2026-02-11 至 2026-02-11 共收录 3
2602.10017 2026-02-11 cs.CL

SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation

SCORE:特定性、上下文利用、鲁棒性与相关性用于无参考LLM评估

Homaira Huda Shomee, Rochana Chaturvedi, Yangxinyu Xie, Tanwi Mallick

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Argonne National Laboratory(阿贡国家实验室) University of Pennsylvania(宾夕法尼亚大学)

AI总结 本文提出了一种无参考评估框架,用于评估LLM在高风险领域任务中的特定性、鲁棒性、相关性和上下文利用,通过精心编纂的数据集和人工评估验证了多指标评估的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09158 2026-02-11 cs.LG cs.AI

What do Geometric Hallucination Detection Metrics Actually Measure?

几何幻觉检测度量实际上测量什么?

Eric Yeats, John Buckheit, Sarah Scullen, Brendan Kennedy, Loc Truong, Davis Brown, Bill Kay, Cliff Joslyn, Tegan Emerson, Michael J. Henry, John Emanuello, Henry Kvinge

机构 * Pacific Northwest National Laboratory(太平洋西北国家实验室) University of Washington(华盛顿大学) University of Pennsylvania(宾夕法尼亚大学) Colorado State University(科罗拉多州立大学) University of Texas, El Paso(德克萨斯大学埃尔帕索分校) Laboratory for Advanced Cybersecurity Research, National Security Agency(国家安全局高级网络安全研究实验室)

AI总结 本文研究几何统计在检测幻觉中的作用,通过合成数据集分析不同属性对幻觉检测的影响,并提出归一化方法提升多领域检测性能。

Comments Published at the 2025 ICML Workshop on Reliable and Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19492 2026-02-11 cs.CL

Machine Text Detectors are Membership Inference Attacks

机器文本检测器是成员推断攻击

Ryuto Koike, Liam Dugan, Masahiro Kaneko, Chris Callison-Burch, Naoaki Okazaki

机构 * Institute of Science Tokyo(东京科学研究院) University of Pennsylvania(宾夕法尼亚大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

AI总结 本文研究了成员推断攻击与机器文本检测之间的可转移性,证明了两者在渐近最优性能度量标准上的相同性,并通过实验证明了跨任务性能的强相关性。

详情

展开后加载摘要…

URL PDF HTML 收藏