Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation
Cutscene Agent:一种用于自动化3D场景生成的LLM代理框架
Lanshan He, Haozhou Pang, Qi Gan, Xin Shen, Ziwei Zhang, Yibo Liu, Gang Fang, Bo Liu, Kai Sheng, Shengfeng Zeng, Chaofan Li, Zhen Hui, Keer Zhou, Lan Zhou, Shujun Dai
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
EmbodiedMidtrain: 通过中训练弥合视觉语言模型与视觉语言动作模型之间的差距
Yiyang Du, Zhanqiu Guo, Xin Ye, Liu Ren, Chenyan Xiong
机构
*
Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所)
;
Bosch Research North America & Bosch Center for Artificial Intelligence (BCAI)(博世北美研究部及博世人工智能中心(BCAI))
Measuring Representation Robustness in Large Language Models for Geometry
在大型语言模型中测量几何表示的鲁棒性
Vedant Jawandhia, Yash Sinha, Murari Mandal, Ankan Pal, Dhruv Kumar
机构
*
Department of Computer Science and Information Systems, BITS Pilani(比特学院计算机科学与信息系统系)
;
School of Computer Science, KIIT University(KIIT大学计算机科学学院)
;
Department of Mathematics, BITS Pilani(比特学院数学系)
Enhancing Geo-localization for Crowdsourced Flood Imagery via LLM-Guided Attention
通过LLM引导注意力增强 crowdsourced 洪水影像的地理定位
Fengyi Xu, Jun Ma, Waishan Qiu, Cui Guo, Jack C. P. Cheng
机构
*
Department of Urban Planning and Design, The University of Hong Kong(香港大学城市规划与设计系)
;
Urban Systems Institute, The University of Hong Kong(香港大学都市系统研究所)
;
Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(香港理工大学土木与环境工程系)
Seeing the Intangible: Survey of Image Classification into High-Level and Abstract Categories
看见无形:图像分类到高级和抽象类别的调查
Delfina Sol Martinez Pandiani, Valentina Presutti
机构
*
University of Bologna(博洛尼亚大学)
;
University of Bologna Department of Computer Science(博洛尼亚大学计算机科学系)
;
University of Bologna Department of Modern Languages, Literatures(博洛尼亚大学现代语言文学系)
;
Centrum Wiskunde en Informatica(数学与信息学研究中心)
;
Institute for Clarity in Documentation(文档清晰研究所)
;
Inria Paris-Rocquencourt(巴黎-罗克奎恩特研究所)
;
Rajiv Gandhi University(拉贾·甘地大学)
;
Tsinghua University(清华大学)
;
Palmer Research Laboratories(帕尔默研究实验室)
LLMs for Text-Based Exploration and Navigation Under Partial Observability
基于部分可观测性的文本基于探索与导航中的大型语言模型
Stephan Sandfuchs, Maximilian Melchert, Jörg Frochte
机构
*
AKIS -- Interdisciplinary Institute for Applied AI and Data Science Ruhr(AKIS——鲁尔跨学科应用人工智能与数据科学研究所)
;
Bochum University of Applied Sciences(波鸿应用科学大学)
Comments15 pages, (to be published Springer Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering [LNICST] )
机构
*
Yale University(耶鲁大学)
;
Broad Institute of MIT and Harvard(麻省理工学院-哈佛大学博德研究所)
;
Google DeepMind(谷歌DeepMind)
;
Stanford University(斯坦福大学)
;
Genentech(基因泰克)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Cornell University(康奈尔大学)
;
Harvard University(哈佛大学)
Concept frustration: Aligning human concepts and machine representations
概念冲突:对齐人类概念与机器表示
Enrico Parisini, Christopher J. Soelistyo, Ahab Isaac, Alessandro Barp, Christopher R. S. Banerji
机构
*
The Francis Crick Institute(弗朗西斯·克里克研究所)
;
Department of Statistical Science, University College London(伦敦大学学院统计科学系)
;
The Alan Turing Institute(艾伦·图灵研究所)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
LVRPO:基于GRPO的语言-视觉对齐用于多模态理解和生成
Shentong Mo, Sukmin Yun
机构
*
Department of Machine Learning, CMU, USA(卡内基梅隆大学机器学习系,美国)
;
Department of Machine Learning, MBZUAI, UAE(穆罕默德·本·扎耶德人工智能大学机器学习系,阿联酋)
;
Department of Artificial Intelligence, Hanyang University ERICA, South Korea(汉阳大学ERICA校区人工智能系,韩国)