arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Empirical Methods in Natural Language Processing · 会议 · Natural Language Processing

2026-06-09 至 2026-06-09 共收录 4
2606.09435 2026-06-09 cs.CL 新提交

MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

MUDIDI:一种基于语言模型的多语言词典数字化两阶段框架

David Setiawan, Temuulen Khishigsuren, Milind Agarwal, Pagnarith Pit, Aso Mahmudi, Ekaterina Vylomova

机构 * School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院) Melbourne School of Psychological Sciences, The University of Melbourne(墨尔本大学墨尔本心理科学学院) LILT

AI总结 提出MUDIDI两阶段框架,结合语言模型实现多语言词典数字化,在字符识别、标记保留和词条分割上优于现有OCR和视觉语言模型,并发布30本公共领域词典的标注数据集。

Comments 9 pages, preprint, submitted to EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08562 2026-06-09 cs.CL 新提交

Inside the LLM Word Factory

LLM单词工厂内部

Benzi Busigin, Yuval Pinter

机构 * Stein Faculty of Computer and Information Science(Stein计算机与信息科学学院)

AI总结 通过激活修补实验,定位Llama2-7B中英语去分词化过程为第1层的两阶段机制:注意力传递非最终子词的令牌特定信号,MLP将其与局部嵌入组合。该结构泛化至八族十二模型,但深度取决于位置编码类型。

Comments 17 pages, 12 figures. Under review at EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01060 2026-06-09 cs.CL cs.AI cs.LG 版本更新

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models

MENTIS: 对齐改变了什么信念?语言模型中多尺度潜在扭转的测量

Partha Pratim Saha, Samarth Raina, Mayur Parvatikar, Amit Dhanda, Vinija Jain, Aman Chadha, Amitava Das

机构 * Pragya Lab, BITS Pilani Goa, India(BITS Pilani 去掉 Goa 的机构名,因为该机构名中包含 'Goa',但根据规则,如果机构已有常见中文名,使用常见中文名。'Pragya Lab, BITS Pilani' 是 BITS Pilani 的一个实验室,因此翻译为 'BITS Pilani 实验室') IIIT Delhi, India(德里印度理工学院) Amazon, USA(美国亚马逊) Meta, USA(美国Meta) Apple, USA(美国苹果)

AI总结 提出MENTIS框架,通过层间协方差扭转范数、谱扭转诊断和能量-辐射-激活度量,测量偏好对齐在语言模型内部计算中引起的选择性、深度局部的几何结构变化。

Comments Submitted to EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14883 2026-06-09 cs.CL cs.CY

OATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM Assistants

OATH-Frames: 利用大语言模型助手分析在线对无家可归者的态度

Jaspreet Ranjit, Brihi Joshi, Rebecca Dorn, Laura Petry, Olga Koumoundouros, Jayne Bottarini, Peichen Liu, Eric Rice, Swabha Swayamdipta

机构 * Dept. of Computer Science, University of Southern California(计算机科学系,南加州大学) Suzanne-Dwork School of Social Work, University of Southern California(苏兹曼-道克社会工作学院,南加州大学)

AI总结 本文提出OATH-Frames框架,通过大语言模型分析社交媒体上的无家可归者态度,提升大规模分析效率并揭示态度趋势。

Comments Project website: https://dill-lab.github.io/oath-frames/, EMNLP Main 2024

Journal ref In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏