Commentsv2: reproducibility study (kappa~0.8), agent-security case (PPMF), anchor semantics, 15+ fixes. Code and data: https://github.com/1549080929-debug/math_agent Keywords: LLM verification; verification autonomy; completeness; ground truth; trustworthy AI Writing and implementation assisted by an AI language model; all experiments, data, and research decisions are the author's own
Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation
通过自我场景增强在多模态大语言模型中强化自我中心空间感知
Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng, Zidong Cao, Lutao Jiang, Zixin Zhang, Huiyu Zhou, Xuming Hu
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Guangxi Zhuang Autonomous Region Information Center(广西壮族自治区信息中心)
;
The Hong Kong University of Science and Technology(香港科技大学)
专题命中
推理与问题求解
:large language model(title,abstract);language model(title,abstract)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
探究优化下上下文与参数化思维链忠实性之间的相互作用
Jingyi Sun, Qianli Wang, Pepa Atanasova, Nils Feldhus, Isabelle Augenstein
机构
*
University of Copenhagen(哥本哈根大学)
;
Technische Universität Berlin(柏林技术大学)
;
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))
;
BIFOLD – Berlin Institute for the Foundations of Learning and Data(BIFOLD – 柏林学习与数据基础研究院)
专题命中
推理与问题求解
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
机构
*
Department of Computer Science and Engineering, Texas A&M University(计算机科学与工程系,德克萨斯A&M大学)
;
Department of Materials Science and Engineering, Texas A&M University(材料科学与工程系,德克萨斯A&M大学)
;
Department of Electrical and Computer Engineering, Texas A&M University(电气与计算机工程系,德克萨斯A&M大学)
;
Computing and Data Sciences, Brookhaven National Laboratory(布鲁赫斯国家实验室计算与数据科学部)
;
Department of Physics and Astronomy, Texas A&M University(物理与天文学系,德克萨斯A&M大学)
专题命中
推理与问题求解
:large language model(abstract);language model(abstract);foundation model(abstract);language agent(abstract)
Agentic retrieval-augmented reasoning reshapes collective reliability under model variability in radiology question answering
代理检索增强推理重塑放射学问答中的集体可靠性在模型变异性下
Mina Farajiamiri, Jeta Sopa, Saba Afza, Lisa Adams, Felix Barajas Ordonez, Tri-Thien Nguyen, Mahshad Lotfinia, Sebastian Wind, Keno Bressem, Sven Nebelung, Daniel Truhn, Soroosh Tayebi Arasteh
专题命中
推理与问题求解
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
CommentsPublished as a conference paper at the Conference on Language Modeling (COLM) 2026. 25 pages, 3 figures, 16 tables. Code: https://github.com/DEFENSE-SEU/FlowEvo
Veith Weilnhammer, Jefferson Ortega, David Whitney
机构
*
Helen Wills Neuroscience Institute, University of California Berkeley, USA(加州大学伯克利分校海伦·威尔斯神经科学研究所)
;
Max Planck UCL Centre for Computational Psychiatry and Ageing Research, London, UK(马克斯·普朗克伦敦大学学院计算精神病学与衰老研究中心)
;
Department of Psychology, University of California Berkeley, USA(加州大学伯克利分校心理学系)
;
Vision Science Group, University of California Berkeley, USA(加州大学伯克利分校视觉科学组)
专题命中
推理与问题求解
:large language model(abstract);language model(abstract);分类 cs.AI
AI总结
研究通过 MAILA 框架利用日常人机交互数据预测心理健康状态,实现高精度的数字表型分析。
Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
瞬间动作预见:多模态线索能替代视频到何种程度?
Manuel Benavent-Lledo, Konstantinos Bacharidis, Victoria Manousaki, Konstantinos Papoutsakis, Antonis Argyros, Jose Garcia-Rodriguez
机构
*
Universidad de Alicante(阿利坎特大学)
;
Foundation for Research and Technology-Hellas(希腊基础研究与技术基金会)
;
University of Crete(克里特大学)
;
Hellenic Mediterranean University(希腊地中海大学)