PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
PortBench: 一种相关性感知的、全流水线的LLM驱动投资组合管理基准
Yuxuan Zhao, Sijia Chen, Ningxin Su
机构
*
Yantai Research Institute of Harbin Engineering University(哈尔滨工程大学烟台研究院)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);pretraining(abstract)
ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
VESTA: 一种全自动的LLM智能体场景生成与安全评估框架
Lu Jia, Haibo Tong, Feifei Zhao, Jindong Li, Dongqi Liang, Ping Wu, Qian Zhang, Yi Zeng
机构
*
BrainCog AI Lab, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所类脑人工智能实验室)
;
Beijing Institute of AI Safety and Governance (Beijing-AISI)(北京人工智能安全与治理研究院)
;
Beijing Key Laboratory of Safe AI and Superalignment(北京市安全人工智能与超级对齐重点实验室)
;
School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院)
;
Long-term AI(长期人工智能)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
CommentsThis paper has already been accepted and presented at the IEEE 27th International Conference on Information Reuse and Integration for Data Science (IRI 2026) from July 31 to August 2, 2026, in Seattle, WA, USA
Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?
基于LLM的金融投资策略能否长期跑赢市场?
Weixian Waylon Li, Hyeonjun Kim, Mihai Cucuringu, Tiejun Ma
机构
*
AIAI, School of Informatics The University of Edinburgh Edinburgh United Kingdom
;
Global Finance Research Center Sungkyunkwan University Seoul Republic of Korea
;
Dept. of Statistics \& OMI University of California, Los Angeles
;
University of Oxford United States
;
The University of Edinburgh
;
Sungkyunkwan University
;
University of California, Los Angeles
;
University of Oxford
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation
需要多少次迭代才能突破限制?多轮LLM评估中的动态预算分配
Shai Feldman, Yaniv Romano
机构
*
Department of Computer Science(计算机科学系)
;
Technion, Israel(技术ion, 以色列)
;
Departments of Electrical and Computer Engineering and of Computer Science(电气与计算机工程系和计算机科学系)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG
LLM-PDESR: Robust PDE Discovery via Subdomain Weighted Residuals and LLM-Guided Symbolic Hypothesis Generation
LLM-PDESR:通过子域加权残差和大语言模型引导的符号假设生成进行稳健的偏微分方程发现
Jinyang Du, Hao Ma, Xiaohu Shi, Bo Yang, Yanchun Liang, Heow Pueh Lee, Chunguo Wu
机构
*
College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院)
;
School of Big Data and Artificial Intelligence, Guangdong University of Finance and Economics(广东财经大学大数据与人工智能学院)
;
Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(吉林大学符号计算与知识工程教育部重点实验室)
;
School of Computer Science, Zhuhai College of Science and Technology(珠海科技学院计算机科学学院)
;
Department of Mechanical Engineering, National University of Singapore(新加坡国立大学机械工程系)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG
CommentsAccepted at the Second Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD), ICML 2026. Revised to match the camera-ready version. OpenReview: https://openreview.net/forum?id=aUSUvS4dPL
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
Clotho:用于LLM输入的任务特定预生成测试充分性度量
Juyeon Yoon, Somin Kim, Robert Feldt, Shin Yoo
机构
*
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
Chalmers University of Technology Gothenburg Sweden(楚德大学哥德堡瑞典)
;
Chalmers University of Technology(楚德大学)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG