Taming the Beast: Learning to Control Neural Conversational Models
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments PhD thesis
AI 大模型
大语言模型、预训练、指令微调、后训练和语言模型应用。
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments PhD thesis
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted at SemEval 2021 Task 4, 8 Pages (7 Pages main content + 1 pages for references)
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to VarDial 2021 @ EACL 2021
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 5 Pages + 1 Page references + 3 Pages Appendix, Accepted at EACL 2021
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to AAAI 2021
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Preprint. Work in progress
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 8 pages, 6 figures, Published in: 2020 International Joint Conference on Neural Networks (IJCNN)
Journal ref 2020 International Joint Conference on Neural Networks (IJCNN), Glasgow, United Kingdom, 2020, pp. 1-8
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments To appear at EMNLP 2020
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments EMNLP 2020 (15 pages)
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments EMNLP 2020 Deep Learning Inside Out (DeeLIO) Workshop; Code available at https://github.com/styfeng/GenAug
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments To appear at AACL 2020; 9 pages, 12 figures, 2 tables
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments to appear at ACL 2020
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ACL 2020 short paper. 5 pages
Journal ref ACL 2020
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Published in NAACL 2019; The first two authors contribute equally; Code: https://github.com/haofuml/cyclical_annealing
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ACL 2019 Demo Track
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 9 pages, 1 figure
专题命中 其他LLM :language model(abstract);prompting(abstract)
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Submitted to conference track at ICLR 2017
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.AI、cs.LG
ANX:面向AI代理交互的协议优先设计及其支持的3EX解耦架构
机构 * Hangzhou Ziyou Data Technology Co., Ltd.(杭州自由数据科技有限公司)
专题命中 其他LLM :LLM(abstract,comments);分类 cs.CL、cs.AI
AI总结 本文提出ANX协议,通过协议创新、架构优化和工具补充,解决AI代理交互中高消耗、碎片化、安全性不足等问题,其核心创新包括代理原生设计、人机交互、轻量应用和可执行SOP。
Comments This open-source AI agent interaction protocol (ANX) is benchmarked against existing protocols (MCP, A2A, ANP, OpenCLI, SkillWeaver, CHEQ, COLLAB-LLM) across four dimensions: tooling, discovery, security, and multi-agent SOP collaboration. Code: https://github.com/mountorc/anx-protocol
机构 * ZAS(莱布尼茨语言信息中心) ; HU Berlin(柏林洪堡大学)
专题命中 其他LLM :LLM(abstract);分类 cs.CL、cs.AI;language model(comments)
Comments ISCA/ITG Workshop on Diversity in Large Speech and Language Models
专题命中 其他LLM :LLM(abstract);分类 cs.CL、cs.AI;foundation model(comments)
Comments CVPR 2024 Workshop on What is Next in Multimodal Foundation Models
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.LG;LLM(comments)
Comments Best Paper Candidate at ISSRE 2023. Replaced "PLM" with "LLM" for better visibility
机构 * Walmart Global Tech(沃尔玛全球技术)
专题命中 其他LLM :LLM(abstract);分类 cs.AI;large language model(comments);language model(comments)
Comments Accepted at RecSys 2025 EARL Workshop on Evaluating and Applying Recommender Systems with Large Language Models
用递归Transformer从有限数据中挖掘更多价值
机构 * Technical University of Munich(慕尼黑工业大学)
专题命中 其他LLM :language model(abstract);分类 cs.CL、cs.LG
AI总结 该研究针对有限数据预训练场景,提出递归Transformer结合因式分解词嵌入的方法,在10M和100M词预算下性能优于标准Transformer,且与2025年BabyLM挑战赛优胜者表现相当。
一个后缀突破所有:针对合并模型家族的感知 Basin 越狱攻击
机构 * RIKEN AIP(理化学研究所先进智能项目) ; Institute of Science Tokyo(东京科学大学) ; Zhejiang University(浙江大学)
专题命中 其他LLM :foundation model(abstract);分类 cs.CL、cs.LG
AI总结 该研究针对合并模型家族提出BAJ方法,利用预训练基础模型的越狱风险,通过最小-最大优化生成可迁移的对抗性后缀,在多种设置下均实现高迁移成功率且能抵御现有防御。
Comments Accepted by EMNLP findings 2026
智能体何时可以停止?带证据的工具使用大语言模型终止机制
机构 * University of California San Diego(加利福尼亚大学圣迭戈分校)
专题命中 其他LLM :LLM(abstract);分类 cs.AI、cs.LG
AI总结 该研究针对工具使用大语言模型的终止问题,提出带证据的终止(ECT)机制,经实验验证其能显著减少不安全完成与过早无支持终止,满足非劣效性要求,可实现成功恢复。
Enrich-Retrieve-Rank:将能力发现扩展至上下文路由之外
机构 * Amazon AGI(亚马逊AGI)
专题命中 其他LLM :LLM(abstract);分类 cs.CL、cs.AI
AI总结 该研究提出Enrich-Retrieve-Rank流程,将能力发现从上下文路由扩展,经实验验证其在大规模MATS组件场景下,比Full-Ctx、Search&Pick基线性能更优且成本更低,已作为多智能体平台的默认能力发现层投入生产。
Comments 11 pages, 4 figures, and 12 tables
FedPref:用于结构化放射报告抽取的联邦偏好学习
机构 * ETH Zurich(苏黎世联邦理工学院) ; Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局)
专题命中 其他LLM :language model(abstract);分类 cs.AI、cs.LG
AI总结 FedPref 是用于结构化放射报告抽取的联邦偏好学习方法,通过冻结公共语言模型、本地排序与共享模型更新训练适配器,在不均数据场景下提升抽取性能且无需共享敏感数据。
Comments Accepted at ELAMI 2026, held in conjunction with MICCAI 2026. To appear in the Springer proceedings