arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-09-01 至 2026-09-01 共收录 248 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 248 篇

2608.30086 2026-09-01 cs.CL cs.LG 新提交 93%

When Does a Classifier Help an LLM? Classifier-Guided Prompting and Hybrid Classifier-LLM Models for Credit-Default Prediction

分类器何时能助力大语言模型?用于信用卡违约预测的分类器引导提示及混合分类器-大语言模型模型

Rishi Datta, Lavanya Prahallad

机构 * Amador Valley High School(阿马多尔谷高中) Research Spark Hub Inc.(Research Spark Hub 公司)

专题命中 评测与基准 :LLM(title,summary_cn);prompting(title,abstract);large language model(abstract);language model(abstract)

AI总结 本研究针对信用卡违约预测任务,对比了单模型性能,提出将分类器的预测概率加入LLM提示的方法,可提升其AUC-ROC至与随机森林相当,同时保持更高召回率,推荐使用该分类器引导提示。

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28626 2026-09-01 cs.CL cs.AI 新提交 93%

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

大型语言模型是否会仔细审查其评审内容?一项关于评分校准、错误检测及作者身份影响的多模态审计研究

Emad Alharbi

机构 * University of Tabuk(塔布克大学)

专题命中 评测与基准 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);prompting(abstract)

AI总结 本研究以两款多模态LLM为评审者评估2026 ICLR投稿,发现其评分高于人类、错误检测率低,提供图表会降低错误检测率,作者身份对评审无影响,编辑决策与简单评分平均一致。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30110 2026-09-01 cs.CL cs.AI 新提交 92%

Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators

大语言模型(LLM)能否监测经济状况?对宏观经济指标的LLM实时预报的评估

Xinyue Zhao, Ruiyi Zhang, Liqin Ye, Rui Cao, Pengtao Xie, Sudheer Chava

机构 * Georgia Institute of Technology(佐治亚理工学院) University of California San Diego(加利福尼亚大学圣迭戈分校)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);pretraining(abstract)

AI总结 本研究推出抗污染基准LiveMacroEval,对比多类基准发现,具备网络搜索能力的LLM智能体对美国16项宏观经济指标的实时预报准确性与专业基准相当,展现出宏观经济实时估算潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30065 2026-09-01 cs.CL cs.AI 新提交 92%

Pak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextualized Urdu Benchmark

Pak3H:使用人类语境化乌尔都语基准评估LLM对齐中的文化不匹配成本

Abdullah Hashmat, Usman Naseem, Agha Ali Raza

机构 * Macquarie University(麦考瑞大学) Lahore University of Management Sciences(拉合尔管理科学大学)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 该研究针对LLM对齐中低资源语言的文化不匹配问题,构建了首个人类验证的乌尔都语3H对齐基准Pak3H1,经评估发现多语言LLM存在对齐差距,凸显人工引导本地化评估的必要性。

Comments We introduce Pak3H, a human-validated Urdu benchmark for helpfulness, harmlessness, and honesty. Zero-shot evaluations show LLM performance degrades across all three dimensions in low-resourced contextualized settings

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23651 2026-09-01 cs.CL 版本更新 92%

How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

大型语言模型有多像人类?一个语域感知的语言评估框架

Björn Nieth, Marianna Gracheva, Michaela Mahlberg, Bjoern Eskofier, Emmanuelle Salin

机构 * Department Artificial Intelligence in Biomedical Engineering (AIBE)(人工智能生物医学工程系) Department of Digital Humanities and Social Studies (DHSS)(数字人文与社会科学系) University of Birmingham(伯明翰大学) Chair of AI-supported Therapy Decisions(人工智能支持治疗决策教授职位) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Institute of AI for Health, Helmholtz Zentrum München(健康人工智能研究所,海德堡中心慕尼黑)

专题命中 评测与基准 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 提出一个基于语域感知的评估框架,通过比较人类参考语料库与LLM生成文本的词汇语法特征分布(使用最大均值差异和Biber的67个特征),发现LLM偏离人类基线,且最接近人类的模型取决于语域而非模型大小。

Comments 9 pages (main) + 31 pages appendix, 29 figures, 10 tables. Code and data: this https URL (https://github.com/BjoernNieth/Register_Aware_LLMs)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16206 2026-09-01 cs.AI cs.CL cs.CY cs.HC 版本更新 92%

Beyond Helpfulness: A Teaching-over-Solving Diagnostic for Measuring Educational Impact in LLM Tutors

衡量LLM导师是教学还是解题:教育影响的诊断方法

Junyi Yao, Zihao Zheng, Baichuan Li

机构 * Washington University in St. Louis(圣路易斯华盛顿大学) Department of Operations Research and Engineering Management, Southern Methodist University(南卫理公会大学运筹学与工程管理系)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 针对LLM作为教育导师时解题能力不等于教学支持的问题,提出基于解题导向与教学导向基准性能差距的诊断方法,通过MathTutorBench分析表明两者仅部分对齐,建议分开报告评分并明确保护学生能动性的标准。

Comments accepted to 2026 EMNLP NLP4PI workshop, in proceedings to ACL anthology

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29224 2026-09-01 cs.CL cs.AI cs.CR 版本更新 92%

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

相关性即漏洞:网络检索如何削弱LLM智能体的安全对齐

Aditya Nawal, Manit Baser, Mohan Gurusamy

机构 * Department of Electrical and Computer Engineering(电子与计算机工程系) National University of Singapore(新加坡国立大学)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出AgentREVEAL框架,分析检索集成方式和内容属性如何导致LLM智能体安全退化,发现相关性是共同激活条件,并引入HarmURLBench基准。

Comments Accepted to EMNLP 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30856 2026-09-01 cs.CL cs.HC 新提交 92%

You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals

你本不该问:一种受语用学启发的评估大语言模型(LLM)拒绝行为的分类法

Ruoxuan Li, Pinqiao Wang, Sheng Li, Cameron Robert Jones

机构 * Stony Brook University(石溪大学) University of Virginia(弗吉尼亚大学)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 该研究提出首个基于语用学理论的LLM拒绝分类法,通过分析16个LLM在14类有害请求下的拒绝回复,发现其拒绝的特点及存在的问题,呼吁开展兼顾情境适应性与社会问责的对齐评估。

Comments To appear in the Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.25920 2026-09-01 cs.AI cs.SE 版本更新 92%

Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems

修复还是重采样?重新思考LLM多智能体系统中的故障调试

Zhongwen Luan, Xiaoyu Zhang, Ming Hu, Yue Yang, Jiongchi Yu, Xiaohong Chen

机构 * East China Normal University(华东师范大学) Nanyang Technological University(南洋理工大学) Singapore Management University(新加坡管理大学) Xi’an University of Architecture and Technology(西安建筑科技大学)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究针对LLM多智能体系统的故障调试,提出SymTrace评估框架与SymFail数据集,发现现有无指导重运行方法不可靠,提出的症状驱动干预方法可显著提升故障修复率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08531 2026-09-01 cs.AI 版本更新 92%

ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

VESTA: 一种全自动的LLM智能体场景生成与安全评估框架

Lu Jia, Haibo Tong, Feifei Zhao, Jindong Li, Dongqi Liang, Ping Wu, Qian Zhang, Yi Zeng

机构 * BrainCog AI Lab, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所类脑人工智能实验室) Beijing Institute of AI Safety and Governance (Beijing-AISI)(北京人工智能安全与治理研究院) Beijing Key Laboratory of Safe AI and Superalignment(北京市安全人工智能与超级对齐重点实验室) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Long-term AI(长期人工智能)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出VESTA框架,基于五个风险维度自动生成1072个可执行场景,评估12个LLM智能体在任务执行中的行为安全风险,平均攻击成功率达47.1%。

Comments Accepted to EMNLP 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06443 2026-09-01 cs.CL cs.MM cs.SI 版本更新 92%

Revising Context, Shifting Simulated Stance: Auditing LLM-Based Stance Simulation in Online Discussions

修正语境,转变模拟立场:审计基于LLM的在线讨论立场模拟

Xinnong Zhang, Wanting Shan, Hanjia Lyu, Zhongyu Wei, Jiebo Luo

机构 * Fudan University(复旦大学) University of Rochester(罗切斯特大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本研究通过反事实语境修正框架审计LLM立场模拟,对比纯文本与多模态策略,评估平均方向性立场转变和立场转换率,揭示语境敏感性的有效性与鲁棒性。

Comments Accepted for publication in the Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27887 2026-09-01 cs.AI q-fin.PM 版本更新 92%

PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management

PortBench: 一种相关性感知的、全流水线的LLM驱动投资组合管理基准

Yuxuan Zhao, Sijia Chen, Ningxin Su

机构 * Yantai Research Institute of Harbin Engineering University(哈尔滨工程大学烟台研究院) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);pretraining(abstract)

AI总结 提出PortBench基准,通过静态QA和动态五阶段分配流水线评估LLM在投资组合管理中的表现,发现多数模型无法超越等权重分配,且存在推理错误累积和压力下大幅回撤的问题。

Comments Project page: this https URL (https://portbench.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18486 2026-09-01 cs.CL cs.CY 版本更新 92%

Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias

不同的人口统计线索导致对LLM个性化和偏见的结论不一致

Manuel Tonneau, Neil K. R. Sehgal, Niyati Malhotra, Sharif Kazemi, Victor Orozco-Olvera, Ana María Muñoz Boudet, Lakshmi Subramanian, Samuel P. Fraiberger, Sharath Chandra Guntuku, Valentin Hofmann

机构 * World Bank(世界银行) University of Oxford(牛津大学) New York University(纽约大学) University of Pennsylvania(宾夕法尼亚大学) LMU Munich(慕尼黑大学) Allen Institute for AI(人工智能研究院)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究通过1480万个提示测试了人口统计线索对LLM响应的影响,发现同一群体的不同线索导致响应变化不一致,偏见结论不稳定,揭示了身份线索与语言信号的关系。

Comments Accepted to EMNLP 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04191 2026-09-01 cs.AI 版本更新 92%

Towards Natural Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions

迈向真实个性化:评估个性化用户-LLM交互中的长周期偏好跟随

Qianyun Guo, Yibo Li, Yue Liu, Bryan Hooi

机构 * National University of Singapore, Singapore(新加坡国立大学)

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 RealPref提出一个评估个性化用户-LLM交互中长周期偏好跟随的基准,揭示LLM在长上下文和隐性偏好下的表现下降问题。

Comments Accepted to EMNLP 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29613 2026-09-01 cs.CL cs.LG 新提交 91%

Cross-lingual Functional Vectors for Emotion Detection in Large Language Models

用于大语言模型情感检测的跨语言功能向量

Jieying Xue, Phuong Minh Nguyen, Minh Le Nguyen, Shogo Okada

机构 * Japan Advanced Institute of Science and Technology(日本先进科学技术学院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.LG

AI总结 本研究以多语言多标签情感识别为基准,探究跨语言功能向量(FVs)在大语言模型中的应用,发现FVs可跨语言提升情感检测性能,是轻量可迁移的多语言任务适配机制。

Comments Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30659 2026-09-01 cs.AR cs.MA cs.SE 新提交 91%

LLM-based Hardware Development with Hierarchical IRs and End-to-End Multi-Agent Workflow

基于大型语言模型(LLM)的硬件开发:分层中间表示(IR)与端到端多智能体工作流

Chenyang Yin, Agasthi Haputhanthri, Aditya Anirudh Jonnalagadda, Zhenyu Bai, Yuanming Song, Saranyu Chattopadhyay, Mohammad Fadiheh, Tom Zelazny, Subhasish Mitra, Tulika Mitra

专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 针对LLM在硬件设计中应用受限的问题,本文提出基于分层IR与多智能体工作流的LLM硬件开发框架,在Verilog-Eval基准获95.5% pass@5,可生成符合标准的功能性复杂硬件设计。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01456 2026-09-01 cs.LG cs.CL cs.GT 版本更新 91%

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

诚实的人工智能顾问:偏好错位下大语言模型诚实性的预设基准

Hamidreza Hasani Balyani, Seyed Pouyan Mousavi Davoudi, Alireza Amiri-Margavi, Amin Gholami Davodi, Arshia Gharagozlou

机构 * Amazon Lab126, HW Tech Org.(亚马逊实验室126,硬件技术组织) Computational Modeling and Simulation University of Pittsburgh(计算建模与仿真大学匹兹堡分校) Mathematics & Statistics Department University of Minnesota Duluth(数学与统计学系明尼苏达大学 Duluth 分校)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.LG

AI总结 通过Crawford-Sobel廉价谈话模型构建基准,评估大语言模型在偏好冲突时是否诚实,发现模型过度揭示信息,偏离策略最优。

Comments 24 pages, 6 figures. Code and data: this https URL (https://github.com/iHamidHasani/cheap-talk-llm-benchmark)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24564 2026-09-01 cs.AI cs.CE cs.LG 版本更新 91%

Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models

召唤神谕以屠之:利用大语言模型缓解金融回测中的前瞻偏差

Weixian Waylon Li, Mengyu Wang, Tiejun Ma

机构 * University of Edinburgh(爱丁堡大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 提出FinCAD方法,通过对抗性偏差发现和实体日期自适应规则,在不重新训练的情况下抑制大语言模型对历史结果的记忆,从而缓解金融回测中的参数化前瞻偏差。

Comments EMNLP 2026, Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23213 2026-09-01 cs.CL cs.AI 版本更新 91%

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process

评分、推理与选择最佳!通过同行评审过程进行大型语言模型的集成

Zhijun Chen, Zeyu Ji, Qianren Mao, Hao Wu, Jinhuan Song, Junhang Cheng, Bangjie Qin, Zhuoran Li, Jingzheng Li, Kai Sun, Zizhe Wang, Yikun Ban, Zhu Sun, Xiangyang Ji, Hailong Sun, Xiao Huang

机构 * Beihang University, Beijing, China(北京航空航天大学) Zhongguancun Laboratory, Beijing, China(中关村实验室) Beijing University of Posts and Telecommunications(北京邮电大学) Hong Kong University of Science and Technology(香港科学与技术大学) Xi'an Jiaotong University, Xi'an, China(西安交通大学) Tsinghua University, Beijing, China(清华大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 评测与基准 :LLM(summary_cn,abstract);large language model(title);language model(title);分类 cs.CL、cs.AI

AI总结 本文提出LLM-PeerReview方法,通过同行评审机制集成多个大型语言模型,提升响应质量。实验显示其在多个数据集上优于Smoothie-Global,适用于事实性问答、数学推理等任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16753 2026-09-01 cs.CL cs.AI 版本更新 91%

Standardizing Longitudinal Radiology Report Evaluation via Large Language Model Annotation

通过大型语言模型标注标准化纵向放射学报告评估

Xinyi Wang, Grazziela Figueredo, Ruizhe Li, Xin Chen

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文提出基于大型语言模型的标注方法,用于标准化放射学报告中的纵向信息评估,通过对比实验验证了其在疾病进展检测和跟踪方面的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29249 2026-09-01 cs.AI cs.CL cs.IR cs.LG 新提交 90%

Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge

验证FKG.in:大语言模型增强的印度食品知识的合理性评估

Saransh Kumar Gupta, Armaan Shah, Lipika Dey, Partha Pratim Das, Ramesh Jain

机构 * Ashoka University(阿肖克大学) Mphasis AI and Applied Tech Lab(美斐西斯人工智能与应用技术实验室) Koita Centre for Digital Health, Ashoka University(阿肖克大学科伊塔数字健康中心) Institute for Future Health, UC Irvine(加州大学欧文分校未来健康研究所)

专题命中 评测与基准 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究针对LLM生成食谱存在的问题,提出半自动化评估框架,验证LLM增强的印度食品知识图谱FKG.in的合理性,该框架可推广至多语言多元文化烹饪领域,为食品知识基础设施提供支撑。

Comments 15 pages, 2 figures, 5 tables, 27 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30903 2026-09-01 cs.CL 新提交 90%

MMDS-Bench: Benchmarking Multimodal Large Language Models on Dynamic Stance in Social Media Interactions

MMDS-Bench:面向社交媒体互动中动态立场的多模态大语言模型基准测试

Yuzhe Ding, Kang He, Li Zheng, Shengwu Zheng, Teng Shi, Fei Li, Chong Teng, Donghong Ji

机构 * School of Cyber Science and Engineering, Wuhan University(武汉大学网络安全学院) Shanghai Innovation Institute(上海创新研究院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL

AI总结 本研究构建多模态动态立场分类基准MMDS-Bench,评估12个多模态大语言模型,发现其在关系推理类多模态动态立场理解上仍存在不足。

Comments Accepted by EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.28833 2026-09-01 cs.AI 新提交 90%

Evaluating the Hidden Costs of Personalization in Large Language Models

评估大语言模型个性化的隐性成本

Yumeng Wang, Yuchen Wu, Cheng Qian, Zhiyuan Fan, Hyeonjeong Ha, Shujin Wu, Jiayu Liu, Heng Ji, Ge Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) HKUST(香港科技大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI

AI总结 本研究针对大语言模型个性化的三种隐性风险,提出PRISK评估框架,经13个模型实证分析,发现个性化信息会加剧偏见并导致特定指标下降。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29535 2026-09-01 physics.soc-ph cs.AI 新提交 90%

Integrating adaptive human behavior into epidemic models with large language models

结合大语言模型将自适应人类行为整合进流行病模型

Yicheng Mao, Haoyang Li, Rob Deardon, Hongru Du

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI

AI总结 研究人员提出结合大语言模型(LLMs)的流行病生成式自适应行为层(GABLE),将自适应人类行为整合进流行病模型,其生成的接触矩阵预测效果优于流动性驱动矩阵,还可用于前瞻性政策评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30730 2026-09-01 cs.LG cs.CL 新提交 90%

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

电商基准测试集:评估大语言模型智能体在长周期自主商业运营中的表现

Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu

机构 * Qwen Team, Alibaba Group(通义千问团队,阿里巴巴集团) Department of Computer Science and Engineering, HKUST(香港科技大学计算机科学与工程系) Taobao & Tmall Group, Alibaba Group(阿里巴巴集团淘宝天猫集团)

专题命中 评测与基准 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 该研究推出首个开源电商长周期自主商业运营基准E-Commerce Bench,评估18个前沿LLM智能体,发现无单一模型占优,GPT-5.6 Sol盈利最高,Qwen3.8-Max-Preview为开源模型最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07469 2026-09-01 cs.CL cs.AI 版本更新 90%

SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation

SynthAVE:通过大语言模型领域验证实现电子商务的可扩展合成标注

Andrea Scarinci, Virginia Negri, Brayan Impata, Suleiman Khan, Victor Martinez, Marcello Federico

机构 * Amazon(亚马逊)

专题命中 评测与基准 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究针对电子商务属性提取微调大语言模型需大量标注数据、人工标注成本高的问题,提出SynthAVE基准及多LLM领域框架,通过多数投票验证合成标注,实现经济高效且质量与人工审核相当的大规模属性值提取验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28183 2026-09-01 cs.CL cs.AI 版本更新 90%

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

BenGER:德国法律中基于归入的法律推理的LLM系统基准测试

Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach, Aleyna Koçak, Anne Zettelmeier, Elly Breu, Angelina Greiner, Sofija Milijas, Matthias Grabmair

机构 * Technical University of Munich (TUM)(慕尼黑技术大学) Ludwig Maximilian University of Munich (LMU)(慕尼黑路德维希-马克西米利安大学) University of Konstanz(康斯坦茨大学) University of Saarbrücken(萨尔布吕肯大学)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 提出BenGER数据集,用于评估LLM系统在德国法律归入推理中的表现,通过自动和基于法官的指标比较12个LLM系统与人类基线。

Comments Pre-Print - Accepted at EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13899 2026-09-01 cs.CL cs.AI 版本更新 90%

Do We Still Need Humans in the Loop? Human vs. LLM Annotation in Active Learning for TikTok Hate Speech Detection

我们是否仍然需要人在回路中?比较主动学习中用于敌意检测的人类与LLM标注

Ahmad Dawar Hakimi, Lea Hirlimann, Isabelle Augenstein, Hinrich Schütze

机构 * Center for Information and Language Processing, LMU Munich(慕尼黑大学信息与语言处理中心) Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 评测与基准 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 研究比较了LLM与人类在主动学习中的标注效果,发现LLM标注成本更低且性能更优,但主动学习在LLM标注下无优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29802 2026-09-01 cs.CV 新提交 90%

Foundation and Multimodal Large Language Models for Face Presentation and Morph Attack Detection

用于人脸呈现与变形攻击检测的基础模型及多模态大语言模型

Hatef Otroshi Shahreza, Asif Hussain Khan, Peter Lorenz, Alain Komaty, Sébastien Marcel

机构 * Idiap Research Institute(Idiap研究所) Université de Lausanne(洛桑大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);foundation model(abstract);prompting(abstract)

AI总结 本文研究通用基础模型和多模态大语言模型是否包含人脸呈现与变形攻击相关信息,通过五种方法在8个公开数据集上测试,发现微调后模型跨数据集检测性能达SOTA,将开源代码。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29773 2026-09-01 cs.CR 新提交 90%

HSMLog: Small Language Model-Assisted Hardware Security Module Log Anomaly Detection with Behavioral Analysis

HSMLog:基于小型语言模型辅助的硬件安全模块日志异常检测与行为分析

Chia-Hsuan Wu, Dar-Hsin Dustin Wu, Rui Fang, Yi-Ting Lee, Chia-Chih Lin, Ming-Syan Chen

专题命中 评测与基准 :language model(title,abstract);small language model(title,abstract);SLM(abstract,abstract_cn)

AI总结 本文提出HSMLog两阶段框架,结合小型语言模型与检索式行为分析,在真实工业HSM日志上实现98.97%精确率等优异指标,可有效完成HSM日志异常检测与事件分诊。

Comments Accepted to the ISSRE 2026 Industry Track; 6 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏