arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 461 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 461 篇

2512.11839 2026-08-13 cs.LG 版本更新 91%

Grounding Large Language Models as Generalizable Policies in Network Control

大语言模型作为网络优化的通用策略

Duo Wu, Linjia Kang, Zhimin Wang, Fangxin Wang, Wei Zhang, Chongbo Sun, Xuefeng Tao, Wei Yang, Le Zhang, Wenwu Zhu, Peng Cui, Zhi Wang

机构 * Bytedance(字节跳动) Shenzhen International Graduate School(深圳国际研究生院) Tsinghua University(清华大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Department of Computer Science and Technology(计算机科学与技术系)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);foundation model(abstract)

AI总结 本文提出Trailblazer框架,利用大语言模型实现跨任务和环境的通用网络策略,通过网络对齐和策略协作机制提升效率与泛化能力。

Comments Arxiv version. Official version has been submitted to IEEE Transactions on Mobile Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22966 2026-08-11 cs.CL 版本更新 91%

Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks

提示工程并不能普遍提升大型语言模型在临床决策任务中的性能

Mengdi Chai, Ali R. Zomorrodi

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);prompting(abstract)

AI总结 本研究发现提示工程对LLMs在临床决策任务中的性能提升具有高度依赖模型和任务的特性,强调了定制化策略的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26577 2026-07-27 cs.AI cs.CY cs.RO 版本更新 91%

Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

对大型语言模型在机器人健康护理员控制中的安全性的基准测试

Mahiro Nakao, Kazuhiro Takemoto

机构 * Kyushu Institute of Technology(九州工业技术大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn,abstract_cn);分类 cs.AI

AI总结 本文通过270条有害指令评估72个LLM,在模拟环境中测试其安全性,发现模型大小和发布日期影响安全性能,专有模型更安全,但医疗领域微调和提示防御策略效果有限,需将安全性作为首要考量。

Comments 20 pages, 9 figures, 3 tables, 8 pages supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16197 2026-07-21 cs.LG 版本更新 91%

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation

为大规模语言模型的可扩展数据归因与估值绘制读出

Yide Ran, Jianwen Xie, Minghui Wang, Wenjin Zheng, Denghui Zhang, Chuan Li, Zhaozhuo Xu

机构 * Stevens Institute of Technology(史蒂文斯理工学院) Lambda Inc.(Lambda公司) Columbia University(哥伦比亚大学) Mailman School of Public Health(马利曼公共卫生学院) University of Texas Health Science Center at Houston(德克萨斯大学健康科学中心休斯顿分校)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);pretraining(abstract)

AI总结 本文提出RISE方法,通过聚焦输出层影响热点,利用分解的外积形式实现高效归因与估值,减少存储并扩展至32B参数模型,验证了其在数据检测、领域分离和高质量数据选择中的有效性。

Comments 54 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12317 2026-07-03 cs.CL 版本更新 91%

Large language models reshape the language of science

大型语言模型重塑科学语言

Dingkang Lin, Naixuan Zhao, Dan Tian, Jiang Li

机构 * School of Information Management, Nanjing University(南京大学信息管理学院) School of Economics and Management, Shanxi University(山西大学经济管理学院)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(summary_cn,abstract_cn);分类 cs.CL

AI总结 通过分析2020-2024年间2136万篇摘要,发现LLM导致科学写作词汇复杂度上升、句法复杂度下降,且对非英语母语学者影响更大,可能扩大科学话语与公众语言的差距。

Comments 72 pages, 24 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07595 2026-08-12 cs.IR 版本更新 91%

Towards Comprehensible Recommendation with Large Language Model Fine-tuning

基于大语言模型微调的可理解推荐研究

Yunze Luo, Yinjie Jiang, Gaode Chen, Xinghua Zhang, Kaigui Bian

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(title);language model(title);post-training(abstract)

AI总结 该研究针对传统推荐的语义协同差距问题,提出CURec框架,通过微调LLM并结合强化学习优化,提升推荐可理解性与性能,在公开基准上表现更优。

Comments 11 pages, 6 figures, to be published on CIKM '26

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16152 2026-06-11 cs.DL cs.AI cs.CL cs.LG 版本更新 91%

Mapping Scientific Literature with Large Language Models and Topic Modeling

利用大语言模型和主题建模绘制科学文献图谱

Mason Smetana, Lev Khazanovich

机构 * Department of Civil and Environmental Engineering(土木与环境工程系) University of Pittsburgh(匹兹堡大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 提出基于大语言模型的两阶段分类框架,通过主题建模分析PNAS工程类文献,生成语义可解释主题并揭示跨主题关联,性能优于传统方法。

Comments 35 pages, 10 figures. Accepted for publication in Scientometrics. Final version available via DOI

Journal ref Scientometrics (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02910 2026-08-21 cs.HC cs.AI cs.CY 版本更新 91%

The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices

基本B效应:基于大语言模型(LLM)的智能体的使用会降低人们选择的独特性与多样性

Sandra C. Matz, Kimberly Klugescheid, C. Blaine Horton, Sofie Goethals

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究通过实地研究和对照实验发现,使用基于LLM的智能体会降低人们选择的独特性与多样性,且顺序选择、智能体个性化会放大该效应,相关发现对设计维护人类多样性的AI系统具有重要意义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02594 2026-07-28 q-bio.QM cs.AI cs.ET cs.IR 版本更新 91%

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

OpenAIs HealthBench in Action: 评估基于LLM的医疗助手在真实临床查询中的表现

Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam

机构 * OpenAI

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 DR.INFO在HealthBench基准测试中表现优异,优于多个前沿LLM,在复杂临床查询中展现出高准确性和情境感知能力。

Comments 13 pages, two graphs

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16818 2026-08-14 cs.CL cs.AI cs.LG 版本更新 90%

Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries

利用结合大语言模型生成的专家摘要的多表征学习提升住院死亡率预测

Harshavardhan Battula, Jiacheng Liu, Jaideep Srivastava

专题命中 领域大模型 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究提出结合LLM生成的ICU记录专家摘要与生理数据的多表征框架,在MIMIC-III数据集上验证其可提升住院死亡率预测性能,且摘要主要通过重组记录已有信息实现提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20450 2026-06-24 cs.CL cs.AI cs.CY cs.LG 版本更新 90%

Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

允许使用LLM润色同行评审的政策目前无法执行

Rounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan, Nihar B Shah, Danish Pruthi

机构 * Indian Institute of Science(印度科学研究院) University of California, Santa Barbara(加州大学圣芭芭拉分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 领域大模型 :LLM(title,title_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究通过模拟多级人机协作的同行评审数据集,评估五种AI文本检测器,发现它们无法可靠区分LLM润色后的评审与纯人工评审,导致误判风险,表明当前政策不可执行。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01955 2026-06-19 cs.CY 版本更新 90%

Teaching Students to Question the Machine: An AI Literacy Intervention Improves Students' Regulation of LLM Use in a Science Task

教导学生质疑机器:一项AI素养干预措施提升学生在科学任务中调节LLM使用的能力

O. Clerc, R. Abdelghani, C. Desvaux, E. Poisson, P. Y. Oudeyer, H. Sauzéon

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本研究通过两小时的AI素养工作坊,训练中学生(8-9年级)在科学问题解决中更有效地使用大语言模型,减少盲目依赖并提高答案质量。

Comments Workshop paper accepted at ALIT4ALL 2026: 2nd International Workshop on AI Literacy Education For All, co-located with AIED 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16800 2026-08-04 cs.CL 版本更新 90%

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis

大语言模型作为细粒度意见分析的自动标注者和标注裁决者

Gaurav Negi, MA Waskow, John McCrae, Omnia Zayed, Paul Buitelaar

机构 * Data Science Institute(数据科学研究所) University of Galway(Galway大学)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(title);language model(title);分类 cs.CL

AI总结 本文探索使用大语言模型作为自动标注者进行细粒度意见分析,提出声明式标注流水线和LLM裁决方法,实验表明LLM在跨度级别可靠但难以再现关系结构,更适合作为标注助手而非完全替代人类。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23356 2026-07-28 cs.CL cs.HC 版本更新 90%

VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs

VeriLLMed: 基于知识图谱的医疗大语言模型交互式可视化调试

Yurui Xiang, Xingyi Mao, Rui Sheng, Zixin Chen, Zelin Zang, Yuyang Wu, Haipeng Zeng, Huamin Qu, Yushi Sun, Yanna Lin

机构 * IEEE

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL

AI总结 本文提出VeriLLMed系统,通过整合外部生物医学知识,帮助开发者审计和调试医疗大语言模型的诊断推理过程,识别三种常见诊断错误类型,提升模型可靠性。

Comments Accepted by VIS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08532 2026-07-24 cs.AI 版本更新 90%

DN-Hypo-Pipeline: An AI-Driven Workflow for Generating Hypotheses using Large Language Models and Scientific Explanations

DN-Hypo-Pipeline:一种基于大语言模型和科学解释的AI驱动假设生成工作流

Lei Lin, Xinlong Pan, Ronghao Wang, Chunbao Zhou, Jue Wang, Yangang Wang, Ivana Rasovska

机构 * Computer Network Information Center, Chinese Academy of Sciences, China(中国科学院计算机网络信息中心)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);prompting(abstract)

AI总结 提出DN-Hypo-Pipeline,利用大语言模型和科学解释作为先验知识,从现有文献中推导新假设,在数据科学建模中通过统计推断和专家评估证明优于直接生成方法,并验证了生成假设对应的算法性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18957 2026-06-24 q-fin.TR cs.AI cs.MA 版本更新 90%

When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments

当AI遇见金融(StockAgent):基于大语言模型的模拟真实环境股票交易

Chong Zhang, Xinyi Liu, Zhongmou Zhang, Mingyu Jin, Lingyao Li, Zhenting Wang, Wenyue Hua, Dong Shu, Suiyuan Zhu, Xiaobo Jin, Sujian Li, Mengnan Du, Yongfeng Zhang

机构 * University of Liverpool(利物浦大学) Peking University(北京大学) Shanghai University of Finance(上海金融学院) Rutgers University(罗格斯大学) University of Michigan(密歇根大学) Northwestern University(西北大学) New York University(纽约大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) New Jersey Institute of Technology(新泽西理工学院)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI

AI总结 提出基于大语言模型的多智能体系统StockAgent,模拟真实股票交易环境,评估外部因素对交易行为的影响,并避免测试集泄露问题。

Comments 33 pages, 10 figures. Published in ACM Transactions on Intelligent Systems and Technology (TIST)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07609 2026-06-15 cs.IR cs.CL cs.CY 版本更新 90%

Is ChatGPT Fair for Recommendation? Evaluating Fairness in Large Language Model Recommendation

ChatGPT 在推荐中是否公平?评估大语言模型推荐的公平性

Jizhi Zhang, Keqin Bao, Yang Zhang, Wenjie Wang, Fuli Feng, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL

AI总结 针对大语言模型推荐(RecLLM)可能存在的偏见,提出公平性基准 FaiRLLM,包含精心设计的指标和涵盖8个敏感属性的数据集,评估发现 ChatGPT 在推荐中仍存在不公平现象。

Comments Accepted by Recsys 2023 (Short). Typo corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29668 2026-08-17 cs.AI cs.CL 版本更新 90%

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

GRASP: 门控回归感知技能提议器用于自我改进的LLM智能体

Johannes Moll, Jean-Philippe Corbeil, Jiazhen Pan, Martin Hadamitzky, Daniel Rueckert, Lisa Adams, Keno Bressem

机构 * Technical University of Munich and TUM University Hospital(慕尼黑技术大学及慕尼黑大学医院) Microsoft Healthcare & Life Sciences(微软医疗与生命科学)

专题命中 领域大模型 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 提出GRASP方法,通过门控回归感知技能库编辑,在硬回归预算下确保每次技能更新带来净改进,显著提升LLM智能体在结构化环境中的操作可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25718 2026-07-30 cs.LG cs.AI cs.IR 版本更新 90%

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

工具并非孤岛:通过查询条件超边预测为语言模型智能体进行集合级工具检索

Xinyi Hong, Pinjun Dong, Xinyang Yu, Binyan Jiang

专题命中 领域大模型 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 研究LLM智能体的工具检索问题,提出HYSET方法,将其表述为查询条件超边预测,通过特定基数交互捕捉工具兼容性,设计为预选择模块,实验证明该方法在工具检索性能及任务成功率上优于基线,还支持零样本/少样本迁移。

Comments 9 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05132 2026-07-28 cs.CL cs.AI 版本更新 90%

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment

PrinciplismQA: 一种基于哲学的评估LLM与人类临床医学伦理对齐的方法

Chang Hong, Minghao Wu, Qingying Xiao, Yuchi Wang, Xiang Wan, Guangjun Yu, Benyou Wang, Yan Hu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) National Health Data Institute, Shenzhen(深圳国家健康数据研究院) Shenzhen Research Institute of Big Data(深圳大数据研究院)

专题命中 领域大模型 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 本文提出PrinciplismQA,一种基于哲学框架的评估方法,用于评估LLM在临床医学伦理上的对齐情况,通过专家验证的问题集揭示模型在伦理推理上的不足。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10576 2026-07-28 cs.LG cs.AI 版本更新 90%

LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization

基于大语言模型的科学方程发现:通过物理引导的令牌正则化策略优化

Boxiao Wang, Kai Li, Tianyi Liu, Chen Li, Junzhe Wang, Yifan Zhang, Jian Cheng

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) State Key Laboratory of Aerodynamics(航空动力国家重点实验室) School of Mathematical Sciences, University of Chinese Academy of Sciences(中国科学院大学数学科学学院)

专题命中 领域大模型 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出PiT-PO框架,通过强化学习使LLM成为适应性生成器,生成科学一致且结构简洁的方程,提升科学发现性能。

Comments Accepted at KDD 2026. Code is available at https://github.com/CAS-CLab/PiT-PO

Journal ref Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Vol. 2, pp. 12195-12206, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21134 2026-07-28 cs.CL cs.CY cs.LG 版本更新 90%

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

TRIDENT:评估金融、医学和法律领域大语言模型的安全性

Zheng Hui, Yijiang River Dong, Ehsan Shareghi, Nigel Collier

专题命中 领域大模型 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 研究针对大语言模型在金融、医学和法律领域的安全评估问题,基于相关伦理准则定义安全原则并引入Trident-Bench基准进行评估,揭示了不同模型的安全差距,为LLM安全研究提供了系统资源和研究基础。

Comments COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12184 2026-08-18 cs.IR 版本更新 90%

Making Collaborative Signals Count: Graph-Aware Large Language Models for Sequential Recommendation

让协同信号发挥作用:面向序列推荐的图感知大语言模型

Fenglin Yan, Bohao Wang, Jian Zhang, Yu Cui, Tongya Zheng, Ye Feng, Can Wang, Jiawei Chen

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract);pretraining(abstract)

AI总结 针对现有序列推荐方法难以捕捉全局协同模式的问题,提出图感知大语言模型框架GALLM,通过构建协同图建模三类关系并整合至注意力机制,在四个基准上取得最优性能,HR@5较最强基线平均提升9.76%。

Comments 10 pages, 5 figures, 4 tables, includes appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20763 2026-07-10 cs.SE 版本更新 90%

Unveiling Large Language Model Supply Chain: Structure, Domain, and Vulnerabilities

揭示大语言模型供应链:结构、领域与漏洞

Yanzhe Hu, Shenao Wang, Tianyuan Nie, Yanjie Zhao, Haoyu Wang

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn)

AI总结 研究大语言模型供应链的结构、领域和漏洞,通过分析PyPI和NPM的开源包数据集构建依赖图,发现其拓扑结构特点及安全风险传播规律,为增强生态系统弹性提供定量见解和策略依据。

Comments Accepted by Internetware 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02964 2026-07-03 cs.CL cs.AI cs.CR cs.LG 版本更新 90%

Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens

更少数据,更多安全:通过最小令牌的资源高效领域自适应持续预训练推进网络安全LLM专业化

Salahuddin Salahuddin, Ahmed Hussain, Jussi Löppönen, Toni Jutila

机构 * SSH Communications Security(SSH通讯安全公司) KTH Royal Institute of Technology(皇家理工学院) Aalto University(阿尔托大学)

专题命中 领域大模型 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);pretraining(abstract)

AI总结 提出资源高效的领域自适应持续预训练方法,利用分布式FSDP流水线在126M词网络安全语料上微调LLM,以最少118.8M令牌实现超越大规模预训练模型的最新性能。

Comments 19 Pages; Updated content and authors list

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07036 2026-06-18 physics.ed-ph 版本更新 90%

Using Large Language Models to Analyze Engagement in Computational Thinking via Computational Physics Essays

使用大型语言模型通过计算物理论文分析计算思维中的参与度

Sean Savage, Amir Bralin, Paul Hur, N. Sanjay Rebello

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn)

AI总结 本研究利用多模态大型语言模型自动评估100篇学生计算物理论文中的计算思维参与度,在明确子任务上达到84%的准确率,但主观整体质量评估准确率仅71%。

Comments 13 pages, 3 figures, 3 tables. Submitted to Physical Review Physics Education Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19008 2026-08-21 cs.AI 版本更新 90%

Computational Phenomenology of Borderline Personality Disorder: A Comparative Evaluation of LLM-Simulated Expert Personas and Human Clinical Experts

边缘型人格障碍的计算现象学:对LLM模拟专家人设与人类临床专家的比较评估

Marcin Moskalewicz, Anna Sterna, Karolina Drożdż, Kacper Dudzic, Marek Pokropski, Paula Flores

机构 * IDEAS Research Institute(IDEAS研究机构) Adam Mickiewicz University(亚当·密茨凯维奇大学) AMU Center for Artificial Intelligence(AMU人工智能中心) Poznań University of Medical Sciences(波兹南医科大学) Maria Curie-Skłodowska University(玛丽·居里-斯洛多夫斯卡大学) University of Warsaw(华沙大学)

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究比较了LLM模拟专家与人类专家在边缘型人格障碍定性分析中的表现,发现模型在某些方面与人类难以区分,并能识别人类遗漏的主题,展示了AI增强分析的潜力。

Comments 28 pages, 8 tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22664 2026-08-20 cs.AI 版本更新 90%

MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance

WorkstreamBench: 评估LLM代理在金融领域的端到端电子表格任务

Thomson Yen, Julian Poeltl, Harshith Srinivas Gear, Yilin Meng, Joshua Fan, Adam Shen, Yili Liu, Ali Bauyrzhan, Patrick Shea, Siri Du, Haoyang Liu, Daniel Guetta, Hongseok Namkoong

机构 * Decision, Risk, and Operations Division, Columbia Business School(哥伦比亚商学院决策、风险与运营部门) ESB Business School, Reutlingen University(图宾根大学ESB商学院)

专题命中 领域大模型 :LLM(title,title_cn);分类 cs.AI

AI总结 本文提出WorkstreamBench,用于评估LLM代理在金融领域复杂端到端电子表格任务中的能力,重点在于财务建模和情景分析等关键流程,通过三个维度(准确性、公式、格式)的细粒度标准来衡量解决方案质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10016 2026-08-04 cs.IR cs.LG 版本更新 90%

Tokenizing Numerical and Embedding Features for LLM RecSys

用于语言模型推荐系统的数值和嵌入特征分词

Zhe Xu, Ankit Peshin, Chiyu Zhang, Feng Qi, Johnson Lui, Anil Ramakrishna, Justin Johnson, Carl Hu, Kaushik Rangadurai, Luke Simon

机构 * Meta

专题命中 领域大模型 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 研究针对多数基于大语言模型的推荐器无法利用非文本信号的问题,提出软令牌融合框架,将数值和嵌入特征映射到LLM嵌入空间,在基于共享参数LLM的双塔检索模型中实例化该框架,实验证明该方法有效提升了检索性能。

Journal ref Second Tokenization Workshop @ COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07053 2026-07-22 cs.AI cs.SY eess.SY 版本更新 90%

Animating Petascale Time-varying Data on Commodity Hardware with LLM-assisted Scripting

在消费级硬件上利用LLM辅助脚本动画化exascale时变数据

Ishrat Jahan Eliza, Xuan Huang, Aashish Panta, Alper Sahistan, Zhimin Li, Amy A. Gooch, Valerio Pascucci

机构 * University of Utah(犹他大学) Vanderbilt University(范德比大学) ViSOAR LLC

专题命中 领域大模型 :LLM(title,title_cn);分类 cs.AI

AI总结 本文提出一个用户友好的框架,用于在消费级工作站上生成exascale时变数据的3D动画,通过通用动画描述符、高效数据访问、定制渲染系统和LLM辅助交互接口,使非可视化专家科学家能快速生成高质量动画。

Comments ©2026 IEEE. Personal use of this material is permitted. 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses. N.B. Due to the limitation "The abstract field cannot be longer than 1,920 characters", the abstract here is shorter than that in the original PDF file

详情

展开后加载摘要…

URL PDF HTML 收藏