arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-01-23 至 2026-01-23 共收录 63 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 多智能体 6 篇

2502.17366 2026-01-23 eess.SY cs.SY 67%

Distributed Coordination for Heterogeneous Non-Terrestrial Networks

异构非地网络的分布式协调

Jikang Deng, Hui Zhou, Mohamed-Slim Alouini

专题命中 多智能体 :agent(abstract);multi-agent(abstract)

AI总结 本文研究了异构非地网络中分布式协调问题,分析了各层通信挑战,并提出基于多智能体深度强化学习的解决方案,通过案例研究验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 工作流自动化 10 篇

2601.15521 2026-01-23 quant-ph 82%

NWQWorkflow: The Northwest Quantum Workflow

NWQWorkflow:西北量子工作流

Ang Li

专题命中 工作流自动化 :workflow(title,abstract);planning(abstract)

AI总结 NWQWorkflow是一个集成量子应用开发、编译、纠错和模拟的端到端工作流系统,旨在推动量子计算领域的开源协作与规模化发展。

Comments This whitepaper reflects the author's own perspectives on the NWQ software ecosystem and does not represent an official position of PNNL, UW, or the U.S. Department of Energy

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21184 2026-01-23 cs.LG cs.AI cs.CL 75%

Can Language Models Discover Scaling Laws?

语言模型能否发现扩展定律?

Haowei Lin, Haotian Ye, Wenzheng Feng, Quzhe Huang, Yujun Li, Hubert Lim, Zhengrui Li, Xiangyu Wang, Jianzhu Ma, Yitao Liang, James Zou

专题命中 工作流自动化 :agent(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出SLDAgent,一种基于进化的代理,能够自动发现比人类衍生定律更准确的扩展定律,展示了AI在科学发现中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15687 2026-01-23 cs.SE cs.AI 73%

FARM: Field-Aware Resolution Model for Intelligent Trigger-Action Automation

FARM:面向智能触发-动作自动化的场感知分辨率模型

Khusrav Badalov, Young Yoon

机构 * Neouly Co., Ltd.(Neouly公司) Hongik University(弘国大学)

专题命中 工作流自动化 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.SE

AI总结 FARM通过双阶段架构实现智能触发-动作自动化,生成具有正确字段绑定的可执行小应用,提升自动化配置效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24592 2026-01-23 cs.AI cs.SE 62%

BPMN Assistant: An LLM-Based Approach to Business Process Modeling

BPMN助手:基于大语言模型的业务流程建模方法

Josip Tomo Licardo, Nikola Tankovic, Darko Etinger

机构 * Faculty of Informatics(信息学院) Juraj Dobrila University of Pula(朱拉·多布里拉大学)

专题命中 工作流自动化 :function calling(abstract);分类 cs.AI、cs.SE

AI总结 BPMN助手通过基于JSON的中间表示和大语言模型,提升BPMN图编辑效率与可靠性

Comments 22 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16045 2026-01-23 cs.AI 57%

AgriPINN: A Process-Informed Neural Network for Interpretable and Scalable Crop Biomass Prediction Under Water Stress

AgriPINN:一种过程指导的神经网络,用于在水分胁迫下可解释且可扩展的作物生物量预测

Yue Shi, Liangxiu Han, Xin Zhang, Tam Sobeih, Thomas Gaiser, Nguyen Huu Thuy, Dominik Behrend, Amit Kumar Srivastava, Krishnagopal Halder, Frank Ewert

机构 * Department of Computing, and Mathematics, Faculty of Science and Engineering(计算与数学系,科学与工程学院) Manchester Metropolitan University(曼彻斯特 Metropolitan 大学) Leibniz Centre for Agricultural Landscape Research (ZALF)(莱比锡农业景观研究中心(ZALF)) Institute of Crop Science and Resource Conservation (INRES)(作物科学与资源保护研究所)

专题命中 工作流自动化 :planning(abstract);分类 cs.AI

AI总结 AgriPINN通过整合生物物理作物生长方程,实现可解释且可扩展的水分胁迫下作物生物量预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15626 2026-01-23 eess.SY cs.AI cs.SY 57%

Bridging Qualitative Rubrics and AI: A Binary Question Framework for Criterion-Referenced Grading in Engineering

连接定性评价标准与AI:一种二元问题框架用于工程中的标准参照评分

Lili Chen, Winn Wing-Yiu Chow, Stella Peng, Bencheng Fan, Sachitha Bandara

专题命中 工作流自动化 :workflow(abstract);分类 cs.AI

AI总结 本研究提出一种二元问题框架,利用生成式AI提升工程数学评估的评分准确性与反馈质量,实现与人类专家相当的评分效果。

Comments Proceedings of the 36th Annual Conference of the Australasian Association for Engineering Education (AAEE 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17841 2026-01-23 cs.SD cs.LG eess.AS 57%

acoupi: An Open-Source Python Framework for Deploying Bioacoustic AI Models on Edge Devices

acoupi:一种用于在边缘设备上部署生物声学AI模型的开源Python框架

Aude Vuilliomenet, Santiago Martínez Balvanera, Oisin Mac Aodha, Kate E. Jones, Duncan Wilson

机构 * The Bartlett Centre for Advanced Spatial Analysis, Faculty of the Build Environment, University College London(大学学院伦敦大学学院)

专题命中 工作流自动化 :workflow(abstract);分类 cs.LG

AI总结 acoupi是一个开源Python框架,旨在简化在边缘设备上部署生物声学AI模型,通过模块化设计提升生物多样性监测的灵活性和可定制性。

Comments 21 pages, 3 figures, 1 table, to be submitted to BES Methods in Ecology and Evolution

Journal ref Methods in Ecology and Evolution, Volume 17, Issue 1 pp. 67-76, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08185 2026-01-23 cs.HC cs.CY cs.IR 50%

Exploring Multidimensional Checkworthiness: Designing AI-assisted Claim Prioritization for Human Fact-checkers

探索多维可信度:为人类事实核查员设计AI辅助的声明优先级排序

Houjiang Liu, Jacek Gwizdka, Matthew Lease

专题命中 工作流自动化 :workflow(abstract)

AI总结 本研究通过设计和混合方法评估,开发AI辅助声明优先级排序原型,揭示事实核查员的多维可信度分层策略,并提出改进事实核查流程的设计建议。

Comments Accepted at CSCW 2025

Journal ref Proceedings of the ACM on Human-Computer Interaction 9, no. 7 (2025): 1-49

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15918 2026-01-23 cs.CV 50%

A Multi-View Pipeline and Benchmark Dataset for 3D Hand Pose Estimation in Surgery

用于外科手术中3D手姿态估计的多视角流程及基准数据集

Valery Fischer, Alan Magdaleno, Anna-Katharina Calek, Nicola Cavalcanti, Nathan Hoffman, Christoph Germann, Joschua Wüthrich, Max Krähenmann, Mazda Farshad, Philipp Fürnstahl, Lilian Calvet

机构 * University Hospital Balgrist, University of Zurich(苏黎世大学附属医院巴尔格斯特医院,苏黎世大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 工作流自动化 :workflow(abstract)

AI总结 本文提出了一种无需微调的多视角流程和外科手术专用基准数据集,用于提升3D手姿态估计的准确性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15528 2026-01-23 cs.DC cs.CR 50%

Securing LLM-as-a-Service for Small Businesses: An Industry Case Study of a Distributed Chatbot Deployment Platform

为中小企业保障LLM即服务的安全性:一个分布式聊天机器人部署平台的行业案例研究

Jiazhu Xie, Bowen Li, Heyu Fu, Chong Gao, Ziqi Xu, Fengling Han

专题命中 工作流自动化 :workflow(abstract)

AI总结 本文提出一个开源多租户平台,帮助中小企业低成本安全部署定制LLM聊天机器人,通过分布式集群和加密网络实现资源池化与隔离,同时集成防提示注入攻击机制。

Comments Accepted by AISC 2026

Journal ref Australasian Information Security Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 软件智能体 3 篇

2510.23761 2026-01-23 cs.SE cs.AI cs.MA 86%

TDFlow: Agentic Workflows for Test Driven Development

TDFlow: 为测试驱动开发设计的代理工作流

Kevin Han, Siddharth Maddikayala, Tim Knappe, Om Patel, Austen Liao, Amir Barati Farimani

机构 * Carnegie Mellon University(卡内基梅隆大学) UC San Diego(南加州大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 软件智能体 :agentic(title,abstract);agent(abstract);workflow(abstract);分类 cs.AI、cs.SE

AI总结 TDFlow通过测试驱动的工作流实现人类水平的测试解析,提升软件修复性能。

Comments Published in the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026 Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06247 2026-01-23 cs.SE cs.AI cs.AR 84%

DUET: Agentic Design Understanding via Experimentation and Testing

DUET:通过实验和测试进行代理设计理解

Gus Henry Smith, Sandesh Adhikary, Vineet Thumuluri, Karthik Suresh, Vivek Pandit, Kartik Hegde, Hamid Shojaei, Chandra Bhagavatula

机构 * Southmountain Research(Southmountain研究机构) ChipStack

专题命中 软件智能体 :agentic(title);agent(abstract);AI agent(abstract);分类 cs.AI、cs.SE

AI总结 DUET通过实验和测试提升AI代理在硬件设计验证中的性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15494 2026-01-23 econ.GN q-fin.EC 67%

Vibe Coding Kills Open Source

vibe编码摧毁开源生态系统

Miklós Koren, Gábor Békés, Julian Hinz, Aaron Lohmann

专题命中 软件智能体 :agent(abstract);AI agent(abstract)

AI总结 本文研究了vibe编码对开源生态系统的影响,发现其虽提高生产力,但削弱用户参与,导致OSS质量和可用性下降,需改变维护者支付方式以维持现状。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. GUI与网页智能体 1 篇

2508.04037 2026-01-23 cs.AI 83%

Evolving in Tasks: Empowering the Multi-modality Large Language Model as the Computer Use Agent

任务演化:使多模态大语言模型成为计算机使用代理

Yuhao Cheng, Liang Tang, Shuxian Li, Yukang Huo, Tiaonan Duan, Kaer Huang, Yanzhe Jing, Yiqiang Yan

机构 * Lenovo Research(联想研究) China Agricultural University(中国农业大学)

专题命中 GUI与网页智能体 :agent(title,abstract);planning(abstract);分类 cs.AI

AI总结 本文提出自演化代理(SEA),通过数据生成、强化学习和模型增强三个创新,使7B参数模型在计算机使用任务中超越同类模型并接近更大规模模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 记忆与上下文管理 7 篇

2601.15290 2026-01-23 cs.HC cs.AI 87%

Agentic Persona Control and Task State Tracking for Realistic User Simulation in Interactive Scenarios

代理人格控制与任务状态跟踪用于交互场景中逼真用户模拟

Hareeshwar Karthikeyan

机构 * Toast Inc.(Toast公司)

专题命中 记忆与上下文管理 :agentic(title,abstract);agent(abstract);AI agent(abstract);multi-agent(abstract)

AI总结 本文提出了一种多代理框架,通过人格控制和任务状态跟踪模拟逼真用户交互,实验表明其在任务完成和真实性方面优于单LLM基线。

Comments - Accepted to 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Scaling Environments for Agents (SEA) - Paper contains 12 pages with 3 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15709 2026-01-23 cs.AI cs.DB cs.LG 84%

AgentSM: Semantic Memory for Agentic Text-to-SQL

AgentSM: 语义记忆用于代理文本到SQL

Asim Biswal, Chuan Lei, Xiao Qin, Aodong Li, Balakrishnan Narayanaswamy, Tim Kraska

机构 * Amazon Web Services(亚马逊网络服务) University of California, Berkeley(加州大学伯克利分校) Oracle Corporation(甲骨文公司) Snowflake Inc.(Snowflake公司)

专题命中 记忆与上下文管理 :agentic(title,abstract);agent(abstract);分类 cs.AI、cs.LG

AI总结 AgentSM通过构建可解释的语义记忆,提升文本到SQL任务的效率和准确性,减少令牌使用和轨迹长度,达到更高执行精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06037 2026-01-23 cs.CL cs.AI cs.CV 76%

TeleMem: Building Long-Term and Multimodal Memory for Agentic AI

TeleMem: 构建面向代理AI的长期和多模态记忆

Chunliang Chen, Ming Guan, Xiao Lin, Jiaxu Li, Luxi Lin, Qiyi Wang, Xiangyu Chen, Jixiang Luo, Changzhi Sun, Dell Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信)

专题命中 记忆与上下文管理 :agentic(title);分类 cs.AI、cs.CL

AI总结 TeleMem通过统一的长期和多模态记忆系统,提升代理AI在长对话和多模态任务中的表现,实现更高的准确率、更低的令牌使用和更快的操作速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11298 2026-01-23 cs.NI cs.AI cs.CL 73%

Integrating Language Models for Enhanced Network State Monitoring in DRL-Based SFC Provisioning

在基于深度强化学习的SFC配置中整合语言模型以增强网络状态监控

Parisa Fard Moshiri, Murat Arda Onsu, Poonam Lohan, Burak Kantarci, Emil Janulewicz

机构 * University of Ottawa(渥太华大学) Ciena(Ciena公司)

专题命中 记忆与上下文管理 :agent(abstract);planning(abstract);分类 cs.AI、cs.CL

AI总结 本文提出将深度强化学习与BERT结合,以提升基于DRL的SFC配置中网络状态监控的效率和准确性。

Comments 6 pages, 5 figures, submitted to 30th IEEE International Symposium on Computers and Communications (ISCC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15677 2026-01-23 quant-ph 50%

Quantum-HPC hybrid computation of biomolecular excited-state energies

量子-高性能计算混合计算用于生物分子激发态能量

Kentaro Yamamoto, Riku Masui, Takahito Nakajima, Miwako Tsuji, Mitsuhisa Sato, Peter Schow, Lukas Heidemann, Matthew Burke, Philipp Seitz, Oliver J. Backhouse, Juan W. Pedersen, John Children, Craig Holliman, Nathan Lysne, Daichi Okuno, Seyon Sivarajah, David Muñoz Ramo, Alex Chernoguzov, Ross Duncan

专题命中 记忆与上下文管理 :workflow(abstract)

AI总结 本研究提出了一种量子-高性能计算混合计算方法,用于准确模拟生物分子激发态能量,结合超级计算机与量子计算机实现复杂反应的高效模拟。

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05810 2026-01-23 cond-mat.stat-mech 50%

Simulation of a generalized asset exchange model with investment and income mechanisms

带有投资和收入机制的广义资产交换模型模拟

Jan Tobochnik, Harvey Gould, William Klein

专题命中 记忆与上下文管理 :agent(abstract)

AI总结 本文提出一个包含投资和收入机制的经济模型,通过模拟发现现实财富分布及非平衡行为特征。

Comments 22 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04938 2026-01-23 math.OC 50%

Initial Error Tolerant Distributed Mean Field Control under Partial and Discrete Information

初始误差容忍的分布式均场控制(部分和离散信息下)

Yuxin Jin, Haotian Wang, Wang Yao, Xiao Zhang

专题命中 记忆与上下文管理 :agent(abstract)

AI总结 本文提出了一种在部分和离散信息下,能够容忍初始误差的分布式均场控制方法,通过最大似然估计实现分布式误差估计,并展示了状态估计的一致性性质。

详情

展开后加载摘要…

URL PDF HTML 收藏

6. Agent评测 11 篇

2601.15778 2026-01-23 cs.AI cs.CL 86%

Agentic Confidence Calibration

代理置信度校准

Jiaxin Zhang, Caiming Xiong, Chien-Sheng Wu

机构 * Salesforce AI Research(Salesforce人工智能研究)

专题命中 Agent评测 :agentic(title,abstract);agent(abstract);AI agent(abstract);分类 cs.AI、cs.CL

AI总结 本文提出Holistic Trajectory Calibration (HTC)方法,通过提取代理轨迹上的过程级特征,提升AI代理的置信度校准能力,实现更可靠的自主系统。

Comments 37 pages, 15 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15679 2026-01-23 cs.AI 85%

Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats

提升跨领域的代理评估方法:敏感信息泄露、欺诈和网络安全隐患

Ee Wei Seah, Yongsen Zheng, Naga Nikshith, Mahran Morsidi, Gabriel Waikin Loh Matienzo, Nigel Gay, Akriti Vij, Benjamin Chua, En Qi Ng, Sharmini Johnson, Vanessa Wilfred, Wan Sie Lee, Anna Davidson, Catherine Devine, Erin Zorer, Gareth Holvey, Harry Coppock, James Walpole, Jerome Wynee, Magda Dubois, Michael Schmatz, Patrick Keane, Sam Deverett, Bill Black, Bo Yan, Bushra Sabir, Frank Sun, Hao Zhang, Harriet Farlow, Helen Zhou, Lingming Dong, Qinghua Lu, Seung Jang, Sharif Abuadbba, Simon O'Callaghan, Suyu Ma, Tom Howroyd, Cyrus Fung, Fatemeh Azadi, Isar Nejadgholi, Krishnapriya Vishnubhotla, Pulei Xiong, Saeedeh Lohrasbi, Scott Buffett, Shahrear Iqbal, Sowmya Vajjala, Anna Safont-Andreu, Luca Massarelli, Oskar van der Wal, Simon Möller, Agnes Delaborde, Joris Duguépéroux, Nicolas Rolin, Romane Gallienne, Sarah Behanzin, Tom Seimandi, Akiko Murakami, Takayuki Semitsu, Teresa Tsukiji, Angela Kinuthia, Michael Michie, Stephanie Kasaon, Jean Wangari, Hankyul Baek, Jaewon Noh, Kihyuk Nam, Sang Seo, Sungpil Shin, Taewhi Lee, Yongsu Kim

专题命中 Agent评测 :agentic(title,abstract);agent(abstract);AI agent(abstract);分类 cs.AI

AI总结 该研究通过跨领域合作,提升代理评估方法,聚焦敏感信息泄露、欺诈和网络安全等共同风险,推动AI系统测试的最佳实践。

Comments The author/contributor list organises contributors by country and alphabetical order within each country. In some places, the order has been altered to match other related publications

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14691 2026-01-23 cs.AI cs.CL 81%

Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation

操纵法官:不忠的推理链可能损害智能体评估

Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Sungryull Sohn, Yunxiang Zhang, Moontae Lee, Hao Peng, Lu Wang, Honglak Lee

机构 * University of Michigan(密歇根大学) LG AI Research(LG人工智能研究) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.CL

AI总结 本文揭示了LLM法官对智能体推理轨迹操纵的脆弱性,表明基于内容的操纵能显著提高假阳性率,凸显了需验证推理与证据的评估机制的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15551 2026-01-23 cs.AI cs.MA 70%

ALIGNAgent: Adaptive Learner Intelligence for Gap Identification and Next-step guidance

ALIGNAgent: 适应性学习智能用于知识缺口识别与下一步指导

Bismack Tokoli, Luis Jaimes, Ayesha S. Dina

机构 * Department Of Data Science and Business Analytics, *Department Of Computer Science, Florida Polytechnic University(数据科学与商业分析系、计算机科学系,佛罗里达理工学院)

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI

AI总结 ALIGNAgent是一种多智能体教育框架,通过整合知识估计、技能缺口识别和定向资源推荐,实现个性化学习,实验证明其在知识熟练度估计上的高精度和高F1分数。

Comments 35 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22777 2026-01-23 cs.CL 70%

MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators

MEDAL:一个用于评估LLM作为多语言开放领域对话评估器的框架

John Mendonça, Alon Lavie, Isabel Trancoso

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.CL

AI总结 MEDAL框架通过多语言多代理方法,评估LLM作为开放领域对话评估者的性能,揭示了现有评判者在检测细微问题上的不足。

Comments EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13710 2026-01-23 cs.LG cs.AI cs.CL 67%

Who Benefits From Sinus Surgery? Comparing Generative AI and Supervised Machine Learning for Predicting Surgical Outcomes in Chronic Rhinosinusitis

哪些人会从鼻窦手术中受益?比较生成式AI与监督学习在预测慢性鼻窦炎手术结果中的应用

Sayeed Shafayet Chowdhury, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan

机构 * Department of Computer Science, Purdue University(计算机科学系,普渡大学) Department of Computer Science, Indiana University Indianapolis(计算机科学系,印第安纳大学印第安纳波利斯分校) Department of Otolaryngology—Head and Neck Surgery, Indiana University School of Medicine(耳鼻喉科—头颈外科,印第安纳大学医学院)

专题命中 Agent评测 :workflow(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文比较生成式AI与监督学习在预测慢性鼻窦炎手术预后中的表现,发现监督学习在准确率和校准上优于生成式AI,后者在解释性上与临床经验一致,支持以ML为主、生成式AI辅助的决策流程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15487 2026-01-23 cs.AI cs.CL cs.MA 62%

MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation

MiRAGE:一种多智能体框架,用于生成多模态多跳问题-答案数据集以评估RAG系统

Chandan Kumar Sahu, Premith Kumar Chilukuri, Matthew Hetrich

机构 * ABB Inc(ABB公司)

专题命中 Agent评测 :agent(abstract);分类 cs.AI、cs.CL

AI总结 MiRAGE通过多智能体框架生成多模态多跳问题-答案数据集,提升RAG系统评估的准确性和复杂性。

Comments 12 pages, 2 figures, Submitted to ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12471 2026-01-23 cs.CL cs.AI 62%

Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

知何时退避:医疗大语言模型在临床不确定性中的表现

Sravanthi Machcha, Sushrita Yerra, Sahil Gupta, Aishwarya Sahoo, Sharmin Sultana, Hong Yu, Zonghai Yao

机构 * Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(马萨诸塞大学阿姆赫斯特曼宁信息与计算机科学学院) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(医疗组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(米纳尔计算机与信息科学学院)

专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.CL

AI总结 本文提出MedAbstain基准,探讨医疗LLM在临床不确定性中的退避能力,发现显式退避选项能显著提升安全性,而模型规模和提示方法效果有限。

Comments Equal contribution for the first two authors; To appear in proceedings of the Main Conference of the European Chapter of the Association for Computational Linguistics (EACL) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏