机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
Large Language Model Department, Tencent(腾讯大语言模型部)
;
Nankai University(南开大学)
Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research
Andy:用于严谨证明与自主研究的数学智能体
Zi'an Wang
机构
*
School of Mathematical Sciences, Tongji University(同济大学数学科学学院)
;
Key Laboratory of Intelligent Computing and Applications (Tongji University)(同济大学智能计算与应用重点实验室)
Comments20 pages, 4 figures, 3 tables, and 3 algorithms. Research logs, reports, and simulation code are available at https://github.com/mowaiwaim/Andy
PULSE: Agentic Investigation with Passive Sensing for Proactive Affective Intervention in Cancer Survivorship
PULSE:基于被动感知的代理探究用于癌症幸存者的主动干预
Zhiyuan Wang, Subigya Nepal, Ariful Islam, Indrajeet Ghosh, Xinyu Chen, Katharine E. Daniel, Laura E. Barnes, Philip Chow
机构
*
Department of Systems and Information Engineering, University of Virginia(系统与信息工程系,弗吉尼亚大学)
;
Center for Behavioral Health and Technology, University of Virginia(行为健康与技术中心,弗吉尼亚大学)
;
Department of Computer Science, University of Virginia(计算机科学系,弗吉尼亚大学)
"LLM Agent Performance" Is Not a Single Evaluation Target
基于LLM的智能体评估统一框架的必要性
Pengyu Zhu, Li Sun, Philip S. Yu, Sen Su
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
;
Chongqing University of Posts and Telecommunications(重庆邮电大学)
CommentsAfter further review of the current submission, we have identified potential legal and intellectual property concerns associated with keeping the manuscript publicly available as a preprint. In particular, there are ongoing considerations regarding institutional affiliation information, intellectual property ownership, and related compliance matters
DeepPAAC: A New Deep Galerkin Method for Principal-Agent Problems
深度主从问题的新深度伽辽金方法
Michael Ludkovski, Changgen Xie, Zimu Zhu
机构
*
Department of Statistics and Applied Probability, University of California, Santa Barbara(加州大学圣巴巴拉分校统计学与应用概率系)
;
Fintech Thrust, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)金融科技方向)
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
这是一个陷阱!面向网络代理的任务重定向说服基准
Karolina Korgul, Yushi Yang, Arkadiusz Drohomirecki, Piotr Błaszczyk, Will Howard, Lukas Aichberger, Chris Russell, Philip H. S. Torr, Adam Mahdi, Adel Bibi
FORTIS: Benchmarking Over-Privilege in Agent Skills
FORTIS:评估代理技能中的过度特权
Shawn Li, Chenxiao Yu, Han Wang, Wei Yang, Ryan Rossi, Franck Dernoncourt, Xiyang Hu, Philip Yu, Chaowei Xiao, Huan Zhang, Yue Zhao
机构
*
University of Southern California(南加州大学)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Adobe Research(Adobe研究)
;
Arizona State University(亚利桑那州立大学)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
;
Johns Hopkins University(约翰霍普金斯大学)