LLM-Assisted Proactive Threat Intelligence for Automated Reasoning
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments 10 Pages, 1 Figure
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments 10 Pages, 1 Figure
专题命中 其他推理 :reasoning(title);分类 cs.CL
专题命中 其他推理 :reasoning(title);分类 cs.CL
Comments 9 pages, 7 figures, 3 tables, submitted to CogSci 2025
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments NeurIPS Pluralistic Alignment Workshop 2024
专题命中 其他推理 :reasoning(title);分类 cs.CL
Comments EMNLP 24 Findings
专题命中 其他推理 :reasoning(title);分类 cs.CL
专题命中 其他推理 :reasoning(title);分类 cs.CL
专题命中 其他推理 :reasoning(title);分类 cs.CL
Comments Published as a conference paper at ACL Findings 2024
专题命中 其他推理 :chain-of-thought(title);分类 cs.CL
专题命中 其他推理 :reasoning(title);分类 cs.AI
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments To appear in Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (March 2024)
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments 14 pages
专题命中 其他推理 :reasoning(title);分类 cs.CL
专题命中 其他推理 :reasoning(title);分类 cs.CL
Comments Accepted to IJCAI 2023 Main Conference (AI for Social Good Track)
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments 5 pages, one figure one table. arXiv admin note: text overlap with arXiv:2303.01067 (longer, prior version of this project)
专题命中 其他推理 :reasoning(title);分类 cs.CL
Comments EACL 2023
专题命中 其他推理 :reasoning(title);分类 cs.CL
Comments EMNLP 2022
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments In Proceedings HCVS 2021, arXiv:2109.03988
Journal ref EPTCS 344, 2021, pp. 79-90
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments In Proceedings AREA 2020, arXiv:2007.11260
Journal ref EPTCS 319, 2020, pp. 117-125
专题命中 其他推理 :reasoning(title);分类 cs.AI
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments 24 pages
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments Forthcoming in the proceedings of AI^3
专题命中 其他推理 :reasoning(title);分类 cs.AI
Comments A preliminary version of the paper appeared in IJCAI '95
专题命中 其他推理 :reasoning(title,comments)
Comments Accepted on 24 September 2025 at NeurIPS 2025 Efficient Reasoning Workshop
大语言模型(LLMs)从针对性的合成多语言数据中提升性能
机构 * UIUC(伊利诺伊大学厄巴纳-香槟分校) ; Uniphore(优尼佛(Uniphore)公司)
专题命中 其他推理 :reasoning(abstract,abstract_cn);分类 cs.CL、cs.AI
AI总结 本研究提出HOTFIXR数据生成框架,通过探测学生模型多语言弱点生成合成多语言数据,可提升大语言模型的多语言性能,降低微调引发的灾难性遗忘。
REHEARSE:大语言模型中用于语言置信度校准的经验式排练
机构 * University of Pennsylvania(宾夕法尼亚大学) ; University of Southern California(南加州大学) ; University of Illinois Chicago(伊利诺伊大学香槟分校)
专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI
AI总结 针对大语言模型置信度与实际正确性不匹配的问题,提出无训练的Rehearse方法,通过置信度校准博弈的反馈生成校准信号,在多模型多基准实验中显著降低了预期校准误差并提升准确率。
AI气象模型是否遗漏极端天气?
专题命中 其他推理 :reasoning(abstract,abstract_cn);分类 cs.AI、cs.LG
AI总结 本研究验证11种物理与AI气象模型,发现极端天气下相对技能缺失是特定模型的特性而非AI气象模型整体的问题,部分AI模型在极端场景表现优于数值模型。
LLM-意识形态可塑性:将LLM的政治行为测量为上下文条件分布
专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI
AI总结 通过系统实证,证明LLM的政治意识形态是上下文条件分布而非固定点,使用VAA-CHES投影模型在六个上下文轴上评估九个LLM,发现其对上下文高度敏感,但整体占据狭窄的Overton窗口。
Comments Under review, 40 pages, 18 figures, 11 tables
SPORK:自我推测性分叉以加速智能体大语言模型推理
机构 * Tsinghua University(清华大学) ; Meituan(美团)
专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG
AI总结 研究LLM智能体推理中串行循环耗时问题,提出SPORK方法,利用模型自身预测提前调度推测工具调用,通过成本模型等组件优化,大幅提升推理效率且不影响任务准确性。
Comments 16 pages, 15 figures. Code: https://github.com/baihuajun24/spork
LLM 谄媚中权威层级的机制性视角
机构 * Independent Research(独立研究)
专题命中 其他推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG
AI总结 通过受控医疗QA实验,发现LLM按感知权威程度分级响应,机制是特定后期层中正确答案表征被权威信号主动擦除,且该擦除与权威水平成比例、抵抗均值向量干预、仅部分可逆。