arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-01-23 至 2026-01-23 共收录 170 信号源:cs.CL, cs.AI, cs.LG

1. 长上下文与记忆 8 篇

2601.06037 2026-01-23 cs.CL cs.AI cs.CV 73%

TeleMem: Building Long-Term and Multimodal Memory for Agentic AI

TeleMem: 构建面向代理AI的长期和多模态记忆

Chunliang Chen, Ming Guan, Xiao Lin, Jiaxu Li, Luxi Lin, Qiyi Wang, Xiangyu Chen, Jixiang Luo, Changzhi Sun, Dell Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 TeleMem通过统一的长期和多模态记忆系统,提升代理AI在长对话和多模态任务中的表现,实现更高的准确率、更低的令牌使用和更快的操作速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15300 2026-01-23 cs.CL 70%

Intelligence Degradation in Long-Context LLMs: Critical Threshold Determination via Natural Length Distribution Analysis

长上下文大语言模型中的智能退化:通过自然长度分布分析确定关键阈值

Weiwei Wang, Jiyong Min, Weijie Zou

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究通过自然长度分布分析确定长上下文LLM的关键阈值,揭示智能退化机制并提出统一框架以指导缓解策略。

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13753 2026-01-23 cs.CL 70%

SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning

SCALAR: 基于科学引用的长上下文学术推理实时评估

Renxi Wang, Honglin Mu, Liqun Ma, Lizhi Lin, Yunlong Feng, Timothy Baldwin, Xudong Han, Haonan Li

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 SCALAR通过基于学术引用的长上下文推理评估框架,评估模型在学术写作中的推理能力,发现多项选择任务能有效区分模型性能,而填空式引用预测任务更具挑战性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04259 2026-01-23 cs.HC cs.CY 67%

Cognitive AI framework 2.0: advances in the simulation of human thought

认知AI框架2.0:人类思维模拟的进展

Rommel Salas-Guerra

专题命中 长上下文与记忆 :large language model(abstract);language model(abstract)

AI总结 认知AI框架2.0通过统一的记忆架构和受控的知识更新,提升人机交互的个性化与适应性,解决可扩展性、偏见缓解和伦理合规等挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15709 2026-01-23 cs.AI cs.DB cs.LG 62%

AgentSM: Semantic Memory for Agentic Text-to-SQL

AgentSM: 语义记忆用于代理文本到SQL

Asim Biswal, Chuan Lei, Xiao Qin, Aodong Li, Balakrishnan Narayanaswamy, Tim Kraska

机构 * Amazon Web Services(亚马逊网络服务) University of California, Berkeley(加州大学伯克利分校) Oracle Corporation(甲骨文公司) Snowflake Inc.(Snowflake公司)

专题命中 长上下文与记忆 :LLM(abstract);分类 cs.AI、cs.LG

AI总结 AgentSM通过构建可解释的语义记忆,提升文本到SQL任务的效率和准确性,减少令牌使用和轨迹长度,达到更高执行精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23761 2026-01-23 cs.SE cs.AI cs.MA 57%

TDFlow: Agentic Workflows for Test Driven Development

TDFlow: 为测试驱动开发设计的代理工作流

Kevin Han, Siddharth Maddikayala, Tim Knappe, Om Patel, Austen Liao, Amir Barati Farimani

机构 * Carnegie Mellon University(卡内基梅隆大学) UC San Diego(南加州大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 长上下文与记忆 :LLM(abstract);分类 cs.AI

AI总结 TDFlow通过测试驱动的工作流实现人类水平的测试解析,提升软件修复性能。

Comments Published in the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026 Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 推理与问题求解 17 篇

2601.16038 2026-01-23 cs.AI 90%

Grounding Large Language Models in Reaction Knowledge Graphs for Synthesis Retrieval

将反应知识图谱接地于大语言模型以实现合成检索

Olga Bunkova, Lorenzo Di Fruscia, Sophia Rupprecht, Artur M. Schweidtmann, Marcel J. T. Reinders, Jana M. Weber

机构 * Department of Intelligent Systems, Delft University of Technology(智能系统系,代尔夫特理工大学) Department of Chemical Engineering, Delft University of Technology(化学工程系,代尔夫特理工大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

AI总结 本研究通过将反应路径检索转化为图查询生成问题,利用对齐示例的单样本提示提升大语言模型在合成规划中的检索准确性。

Comments Accepted at ML4Molecules 2025 (ELLIS UnConference workshop), Copenhagen, Denmark, December 2, 2025. Workshop page: https://moleculediscovery.github.io/workshop2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15816 2026-01-23 eess.SY cs.AI cs.SY 89%

Virtual Traffic Police: Large Language Model-Augmented Traffic Signal Control for Unforeseen Incidents

虚拟交通警察:基于大语言模型的交通信号控制用于突发事件

Shiqi Wei, Qiqing Wang, Kaidi Yang

机构 * Department of Civil and Environmental Engineering, National University of Singapore(环境与土木工程系,新加坡国立大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出一种基于大语言模型的交通信号控制框架,通过虚拟交通警察代理实时调整信号控制器参数,提升应对突发事件的效率和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05882 2026-01-23 cs.CL cs.AI cs.LG 87%

Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes

协作、审议、评估:LLM对齐如何影响协调的多智能体结果

Abhijnan Nath, Carine Graff, Nikhil Krishnaswamy

机构 * Natural Language (SIGNAL) Lab Colorado State University Fort Collins, CO USA Natural Language (SIGNAL) Lab Colorado State University

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究了LLM对齐方法如何影响多智能体协作效果,通过干预代理促进审议式决策,发现鲁棒性方法在支持正确任务结果方面表现更优。

Comments This submission is a new version of arXiv:2509.05882v1. with a substantially revised experimental pipeline and new metrics. In particular, collaborator agents are now instantiated independently via separate API calls, rather than generated autoregressively by a single agent. All experimental results are new. Accepted as an extended abstract at AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04157 2026-01-23 cs.CL cs.LG 84%

FLEx: Language Modeling with Few-shot Language Explanations

FLEx:基于少样本语言解释的语言建模

Adar Avsian, Christopher Richardson, Anirudh Sundar, Larry Heck

机构 * Georgia Institute of Technology(佐治亚理工学院) Microsoft(微软)

专题命中 推理与问题求解 :language model(title,abstract);prompting(abstract);分类 cs.CL、cs.LG

AI总结 FLEx通过少量解释性示例改进语言模型,有效减少错误并优于链式推理方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06299 2026-01-23 cs.CY cs.AI cs.CL cs.LG 83%

How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare

恶意AI群如何威胁民主:代理AI与大语言模型的融合标志着信息战争的新前沿

Daniel Thilo Schroeder, Meeyoung Cha, Andrea Baronchelli, Nick Bostrom, Nicholas A. Christakis, David Garcia, Amit Goldenberg, Yara Kyrychenko, Kevin Leyton-Brown, Nina Lutz, Gary Marcus, Filippo Menczer, Gordon Pennycook, David G. Rand, Maria Ressa, Frank Schweitzer, Dawn Song, Christopher Summerfield, Audrey Tang, Jay J. Van Bavel, Sander van der Linden, Jonas R. Kunst

机构 * Department of Sustainable Communication Technologies, SINTEF Digital(可持续通信技术系,SINTEF数字) Max Planck Institute for Security and Privacy(安全与隐私研究所) Department of Mathematics, City St George’s University of London(数学系,圣乔治大学) Macrostrategy Research Initiative(战略研究计划) Human Nature Lab, Yale University(人性实验室,耶鲁大学) Department of Politics and Public Administration, University of Konstanz(政治与公共管理系,康斯坦茨大学) Harvard Business School, Harvard University(哈佛商学院,哈佛大学) Department of Psychology, University of Cambridge(心理学系,剑桥大学) Department of Computer Science, University of British Columbia(计算机科学系,不列颠哥伦比亚大学) Department of Human Centered Design & Engineering, University of Washington(以人为本设计与工程系,华盛顿大学) Department of Psychology, New York University(心理学系,纽约大学) Observatory on Social Media and Luddy School of Informatics, Computing, and Engineering, Indiana University(社交媒体观察所和信息、计算与工程学院,印第安纳大学)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文探讨了恶意AI群通过融合代理AI与大语言模型对民主构成的威胁,并提出多方面的干预措施。

Comments 5 Pages, This is the author's version of the work. It is posted here by permission of the AAAS for personal use, not for redistribution. The definitive version was published in Science on January 22, 2026, DOI: 10.1126/science.adz1697

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20513 2026-01-23 cs.CL 83%

Self-correction is Not An Innate Capability in Language Models

语言模型中的自我纠正并非一种固有能力

Guangliang Liu, Zimo Qi, Xitong Zhang, Lu Cheng, Kristen Marie Johnson

机构 * Michigan State University(密歇根州立大学) Johns Hopkins University(约翰霍普金斯大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 推理与问题求解 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 本文研究语言模型是否具备固有道德自我纠正能力,通过行为和机制分析发现其缺乏道德敏感性和有效整合外部反馈的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10883 2026-01-23 cs.AI cs.CL 82%

Chat-TS: Enhancing Multi-Modal Reasoning Over Time-Series and Natural Language Data

Chat-TS: 提升时间序列与自然语言数据的多模态推理能力

Paul Quinlan, Qingguo Li, Xiaodan Zhu

机构 * Electrical and Computer Engineering, Queen’s University(皇后大学电气与计算机工程学院) Mechanical and Materials Engineering, Queen’s University(皇后大学机械与材料工程学院) Ingenuity Labs Research Institute, Queen’s University(皇后大学创新实验室研究 institute)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);instruction tuning(abstract)

AI总结 Chat-TS通过整合时间序列标记提升多模态推理能力,提供新数据集和训练策略,在保持自然语言能力的同时增强时间序列推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15495 2026-01-23 cs.AI cs.CL 73%

Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge

追踪知识传播的极限:LLMs在存在冲突知识的多步骤推理中的失败

Yiyang Feng, Zeming Chen, Haotian Wu, Jiawei Zhou, Antoine Bosselut

机构 * EPFL(苏黎世联邦理工学院) Stony Brook University(石溪大学)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 TRACK基准测试揭示了LLMs在存在冲突知识的多步骤推理中因无法有效整合更新信息而性能下降的问题。

Comments Accepted to EACL 2026 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15296 2026-01-23 cs.CL cs.AI 73%

Entropy-Tree: Tree-Based Decoding with Entropy-Guided Exploration

熵树:基于熵引导的树状解码

Longxuan Wei, Yubo Zhang, Zijiao Zhang, Zhihu Wang, Shiwan Zhao, Tianyu Huang, Huiting Zhao, Chenfei Liu, Shenao Zhang, Junchi Yan

机构 * Shanghai Jiaotong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Huawei Technologies Ltd.(华为技术有限公司) Nankai University(南开大学) Nanyang Technological University(南洋理工大学)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 熵树通过利用熵信号指导分支决策,提升解码效率和不确定性估计的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16007 2026-01-23 cs.CV cs.AI 70%

PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models

PhysicsMind: 为基础多模态大语言模型和世界模型中的物理推理和预测进行仿真与现实力学基准测试

Chak-Wing Mak, Guanyu Zhu, Boyi Zhang, Hongji Li, Xiaowei Chi, Kevin Zhang, Yichen Wu, Yangfan He, Chun-Kai Fan, Wentao Lu, Kuangzhi Ge, Xinyu Fang, Hongyang He, Kuan Lu, Tianxiang Xu, Li Zhang, Yongxin Ni, Youhua Li, Shanghang Zhang

机构 * Peking University(北京大学) Mohamed bin Zayed University of Artificial Intelligence(莫扎伊德大学人工智能学院) National University of Singapore(新加坡国立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Science and Technology of China(中国科学技术大学) Cornell University(康奈尔大学) Hong Kong Polytechnic University(香港理工大学) City University of Hong Kong(香港城市大学)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 PhysicsMind是一个结合现实和仿真环境的统一基准,用于评估基础多模态大语言模型和世界模型在物理推理和预测中的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15708 2026-01-23 cs.CL 70%

Persona Switch: Mixing Distinct Perspectives in Decoding Time

Persona Switch: 在解码时间混合不同的视角

Junseok Kim, Nakyeong Yang, Kyomin Jung

机构 * Seoul National University(首尔国立大学)

专题命中 推理与问题求解 :language model(abstract);prompting(abstract);分类 cs.CL

AI总结 Persona Switch是一种在解码过程中动态结合零样本提示和角色扮演提示优势的新方法,通过比较输出置信度来提升模型性能。

Comments EACL'26 Findings, Code is available at https://github.com/junseokkim00/PersonaSwitch

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01910 2026-01-23 cs.AI 70%

MMP-A*: Multimodal Perception Enhanced Incremental Heuristic Search on Path Planning

MMP-A*:多模态感知增强的路径规划增量启发式搜索

Minh Hieu Ha, Khanh Ly Ta, Hung Phan, Tung Doan, Tung Dao, Dao Tran, Huynh Thi Thanh Binh

机构 * Hanoi University of Science and Technology(河内科学技术大学) Vingroup Big Data Research Center(VinGroup大数据研究中心) FPT Software AI Center(FPT软件AI中心)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 MMP-A*通过融合视觉语言模型的空间感知与自适应衰减机制,提升路径规划的几何精度与计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09049 2026-01-23 cs.AI cs.CV cs.RO 70%

VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning

VIKI-R: 通过强化学习协调具身多智能体合作

Li Kang, Xiufeng Song, Heng Zhou, Yiran Qin, Jie Yang, Xiaohong Liu, Philip Torr, Lei Bai, Zhenfei Yin

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 VIKI-R通过强化学习协调多智能体合作,提出分层基准VIKI-Bench,显著提升多智能体视觉驱动合作性能。

Comments Accepted by NeurIPS 2025 Track on Datasets and Benchmarks. Project page: https://faceong.github.io/VIKI-R/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15624 2026-01-23 cs.CV 67%

Explainable Deepfake Detection with RL Enhanced Self-Blended Images

可解释的深度伪造检测与强化学习增强的自混合图像

Ning Jiang, Dingheng Zeng, Yanhong Liu, Haiyang Yi, Shijie Yu, Minghe Weng, Haifeng Shen, Ying Li

机构 * 1School of Software \& Microelectronics, Peking University, Beijing, China 2Mashang Consumer Finance Co., Ltd., Chongqing, China -0.15in

专题命中 推理与问题求解 :large language model(abstract);language model(abstract)

AI总结 本文提出了一种基于强化学习增强的自混合图像方法,用于可解释的深度伪造检测,通过自动化数据生成和定制奖励机制提升检测性能。

Comments Accepted at ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15737 2026-01-23 cs.AI cs.CL 62%

PhysProver: Advancing Automatic Theorem Proving for Physics

PhysProver: 推动物理领域自动定理证明的发展

Hanning Zhang, Ruida Wang, Rui Pan, Wenyuan Wang, Bingxu Meng, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Rutgers University(罗格斯大学)

专题命中 推理与问题求解 :foundation model(abstract);分类 cs.CL、cs.AI

AI总结 PhysProver通过结合可验证语言和强化学习,提升物理领域形式化定理证明的效率和效果。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16163 2026-01-23 cs.AI cs.RO 57%

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Cosmos Policy: 为视觉运动控制和规划微调视频模型

Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shuran Song, Ming-Yu Liu, Chelsea Finn, Jinwei Gu

机构 * NVIDIA Stanford University(斯坦福大学)

专题命中 推理与问题求解 :post-training(abstract);分类 cs.AI

AI总结 Cosmos Policy通过单阶段后训练将预训练视频模型转化为高效机器人策略,实现视觉运动控制与规划的先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15656 2026-01-23 cs.HC 50%

Reflective Motion and a Physical Canvas: Exploring Embodied Journaling in Virtual Reality

反思性运动与物理画布:探索虚拟现实中的具身日记

Michael Yin, Robert Xiao, Nadine Wagener

专题命中 推理与问题求解 :prompting(abstract)

AI总结 本文提出了一种基于虚拟现实的具身日记方法,通过身体运动和语音表达替代传统写作,探索其在情感反思中的独特表现与潜力。

Comments 19 pages, 6 figures, accepted at CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 评测与基准 31 篇

2601.15628 2026-01-23 cs.AI 89%

CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models

CogToM:一种受人类认知启发的大型语言模型全面理论思维基准

Haibo Tong, Zeyang Yue, Feifei Zhao, Erliang Lin, Lu Jia, Ruolin Chen, Yinqian Sun, Qian Zhang, Yi Zeng

机构 * BrainCog Lab, Institute of Automation, Chinese Academy of Sciences(脑认知实验室,自动化研究所,中国科学院) Beijing Institute of AI Safety and Governance (Beijing-AISI)(北京人工智能安全与治理研究所) Beijing Key Laboratory of Safe AI and Superalignment(北京安全人工智能与超对齐重点实验室) School of Artificial Intelligence, UCAS(人工智能学院,中国科学院大学) Long-term AI(长期人工智能)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 CogToM通过8000个双语实例评估LLM的理论思维能力,揭示性能异质性和认知瓶颈,为研究LLM认知边界提供新视角。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13137 2026-01-23 cs.CL 89%

Adversarial Alignment: Ensuring Value Consistency in Large Language Models for Sensitive Domains

对抗对齐:确保大型语言模型在敏感领域中的价值一致性

Yuan Gao, Zhigang Liu, Xinyu Yao, Bo Chen, Xiaobing Zhao

机构 * School of Information Engineering, Minzu University of China(中国民族大学信息工程学院) National Language Resource Monitoring and Research Center of Minority Languages(少数民族语言资源监测与研究中心)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

AI总结 本文提出对抗对齐框架,通过持续预训练、指令微调和对抗训练提升大型语言模型在敏感领域中的价值一致性,实验表明其优于现有主流模型。

Comments 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10309 2026-01-23 cs.AI cs.HC cs.SI 88%

A large-scale evaluation of commonsense knowledge in humans and large language models

对人类和大语言模型常识知识的大规模评估

Tuan Dung Nguyen, Duncan J. Watts, Mark E. Whiting

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.AI;LLM(comments)

AI总结 本文提出了一种评估人类和大语言模型常识知识的方法,发现较小的开放权重模型在常识能力上表现更优,强调常识知识的文化基础与人类群体差异。

Comments Code and data: https://github.com/Watts-Lab/commonsense-llm-eval

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15406 2026-01-23 cs.CV 88%

Evaluating Multimodal Large Language Models for Heterogeneous Face Recognition

评估多模态大语言模型用于异构人脸识别

Hatef Otroshi Shahreza, Anjith George, Sébastien Marcel

机构 * Idiap Research Institute(日内瓦研究所)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract)

AI总结 本文评估了多模态大语言模型在异构人脸识别中的性能,发现其在跨模态条件下表现欠佳,凸显了当前模型的局限性及生物识别评估的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16134 2026-01-23 cs.AI cs.CL 86%

LLM Prompt Evaluation for Educational Applications

用于教育应用的LLM提示评估

Langdon Holmes, Adam Coscia, Scott Crossley, Joon Suh Choi, Wesley Morris

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出了一种系统评估方法,通过测试六个提示模板,发现一个结合角色和上下文管理器模式的提示在教育应用中表现最佳,支持元认知学习策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15556 2026-01-23 cs.CY cs.CL 85%

LLM or Human? Perceptions of Trust and Information Quality in Research Summaries

LLM还是人类?研究摘要中的信任与信息质量感知

Nil-Jana Akpinar, Sandeep Avula, CJ Lee, Brandon Dang, Kaza Razat, Vanessa Murdock

机构 * Microsoft(微软) Amazon AWS AI(亚马逊AWS人工智能)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究探讨了读者对LLM生成摘要的信任与质量感知,发现尽管难以识别LLM内容,但读者的信念显著影响评估,且LLM编辑的摘要更受好评。

Comments Accepted to ACM CHI conference on Human Factors in Computing Systems(CHI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15385 2026-01-23 cs.HC cs.CL 85%

VegaChat: A Robust Framework for LLM-Based Chart Generation and Assessment

VegaChat: 一种基于大语言模型的图表生成与评估稳健框架

Marko Hostnik, Rauf Kurbanov, Yaroslav Sokolov, Artem Trofimov

机构 * JetBrains

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 VegaChat通过引入Spec Score和Vision Score两个指标,解决了LLM基于自然语言生成可视化中的评估难题,实现了高准确率的图表生成与评估。

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏