arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-01-16 至 2026-01-16 共收录 45 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 45 篇

2601.10122 2026-01-16 cs.CL cs.AI cs.HC 91%

Role-Playing Agents Driven by Large Language Models: Current Status, Challenges, and Future Trends

由大型语言模型驱动的角色扮演代理:现状、挑战与未来趋势

Ye Wang, Jiaxing Chen, Hongjiang Xiao

机构 * Communication University of China(中国传媒大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

AI总结 本文系统回顾了由大型语言模型驱动的角色扮演代理的现状、挑战及未来趋势,探讨了关键技术路径与评估方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10660 2026-01-16 cs.CL 89%

Detecting Winning Arguments with Large Language Models and Persuasion Strategies

利用大型语言模型和说服策略检测获胜论点

Tiziano Labruna, Arkadiusz Modzelewski, Giorgio Satta, Giovanni Da San Martino

机构 * University of Padua(帕多瓦大学) Polish-Japanese Academy of Information Technology(波兰-日本信息科技学院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL

AI总结 本文利用大型语言模型和说服策略分析论证文本,通过多策略评分方法提升说服力预测,并公开发布主题注释数据集以促进研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10194 2026-01-16 quant-ph physics.chem-ph 89%

Autonomous Quantum Simulation through Large Language Model Agents

通过大语言模型代理实现自主量子模拟

Weitang Li, Jiajun Ren, Lixue Cheng, Cunxi Gong

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本文提出利用大语言模型代理实现自主量子模拟,通过上下文学习和多代理架构提升模拟效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09972 2026-01-16 cs.AI 88%

Chinese Labor Law Large Language Model Benchmark

中国劳动法大语言模型基准

Zixun Lan, Maochun Xu, Yifan Ren, Rui Wu, Jianghui Zhou, Xueyang Cheng, Jianan Ding Ding, Xinheng Wang, Mingmin Chi, Fei Ma

机构 * Department of Mechatronics and Robotics, Xi’an Jiaotong–Liverpool University(机械与机器人系,西安交通大学利物浦大学) Department of Applied Mathematics, Xi’an Jiaotong–Liverpool University(应用数学系,西安交通大学利物浦大学) Hansheng Lawyers Building(翰盛律师事务所) Suzhou Gewu Digital Technology Co., Ltd.(苏州葛武数字技术有限公司) Kenneth Wang School of Law, Soochow University(肯尼斯·王法学院,苏州大学) College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出LabourLawLLM和LabourLawBench,专门针对中国劳动法任务,通过精确的法律知识和复杂推理提升法律AI的应用效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09833 2026-01-16 cs.CL 88%

Stable and Explainable Personality Trait Evaluation in Large Language Models with Internal Activations

在大型语言模型中利用内部激活实现稳定且可解释的人格特质评估

Xiaoxu Ma, Xiangbo Zhang, Zhenyu Weng

机构 * Georgia Institute of Technology(佐治亚理工学院) South China University of Technology(华南理工大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文提出PVNI方法,通过内部激活提取人格向量并插值得分,实现LLMs中稳定且可解释的人格特质评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10460 2026-01-16 cs.CL cs.AI cs.CY cs.LG 88%

Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models

上下文立体集:在大型语言模型中压力测试偏见对齐的鲁棒性

Abhinaba Basu, Pavan Chakraborty

机构 * Indian Institute of Information Technology, Allahabad (IIITA)(印度信息与技术研究所(Allahabad)) National Institute of Electronics and Information Technology (NIELIT)(国家电子与信息技术研究所)

专题命中 评测与基准 :large language model(title);language model(title);prompting(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 上下文立体集通过压力测试大型语言模型在不同上下文中的偏见对齐鲁棒性,揭示固定条件测试的偏见分数可能无法推广。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09905 2026-01-16 cs.SE cs.CL 87%

Self-reflection in Automated Qualitative Coding: Improving Text Annotation through Secondary LLM Critique

在自动化定性编码中的自我反思:通过二次LLM批评提升文本标注

Zackary Okun Dunivin, Mobina Noori, Seth Frey, Curtis Atkinson

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 通过二次LLM批评提升文本标注精度,减少错误率并提高分类器性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20721 2026-01-16 cs.CL cs.AI cs.HC 86%

User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios

用户感知与代理LLM评判:隐私与帮助性在隐私敏感场景中的LLM响应

Xiaoyuan Wu, Roshni Kaushik, Wenkai Li, Lujo Bauer, Koichi Onoue

机构 * Fujitsu Research of America Inc.(富士通美国研究所) Carnegie Mellon University(卡内基梅隆大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究发现用户对LLM响应的隐私和帮助性感知与代理LLM评判存在显著差异,需加强用户为中心的评估以提升隐私保护与实用性平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05465 2026-01-16 cs.AI cs.CL 86%

VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models

VAL-Bench:信念一致性作为语言模型价值观对齐的度量标准

Aman Gupta, Denny O'Shea, Fazl Barez

机构 * MasterClass University of Oxford(牛津大学) WhiteBox Martian

专题命中 评测与基准 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 VAL-Bench通过评估语言模型在现实价值相关提示中的信念一致性,提出了一种衡量价值观对齐的新基准,揭示了不同模型在一致性上的显著差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02097 2026-01-16 cs.CL cs.AI 86%

JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation

JudgeAgent: 超越静态基准的面向知识驱动和动态LLM评估

Zhichao Shi, Xuhui Jiang, Chengjin Xu, Cangli Yao, Shengjia Ma, Yinghan Shen, Zixuan Li, Jian Guo, Yuanzhuo Wang

机构 * DataArc Tech Ltd.(DataArc科技有限公司) IDEA Research, International Digital Economy Academy(国际数字经济学院IDEA研究所) School of Advanced Interdisciplinary Sciences, UCAS(北京大学交叉学科学院) State Key Lab of AI Safety, Institute of Computing Technology, CAS(中国科学院计算技术研究所人工智能安全国家重点实验室)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 JudgeAgent通过知识驱动和动态评估框架,解决LLM评估中知识覆盖不足和难度不匹配的问题,实现更全面的评估和模型优化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10025 2026-01-16 cs.AI 85%

Structured Personality Control and Adaptation for LLM Agents

结构化人格控制与适应用于大语言模型代理

Jinpeng Wang, Xinyu Jia, Wei Wei Heng, Yuquan Li, Binbin Shi, Qianlei Chen, Guannan Chen, Junxia Zhang, Yuyu Yin

机构 * Key Laboratory of Complex Systems Modeling and Simulation of Ministry of Education(教育部复杂系统建模与仿真重点实验室) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出基于荣格心理学类型模型的结构化人格控制框架,通过协调、强化补偿和反思机制实现LLM人格的细腻表达与动态适应,提升人机交互中的自然代理设计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09717 2026-01-16 cs.CL cs.AI 84%

SALP-CG: Standard-Aligned LLM Pipeline for Classifying and Grading Large Volumes of Online Conversational Health Data

SALP-CG:面向在线医疗对话数据分类与分级的标准对齐大语言模型流水线

Yiwei Yan, Hao Li, Hua He, Gong Kai, Zhengyi Yang, Guanfeng Liu

机构 * School of Computing, Macquarie University, Australia(麦考瑞大学计算机学院) Department of Biomedical Engineering, National University of Singapore(新加坡国立大学生物医学工程系) School of Mathematics and Statistics, Shandong University of Technology, China(山东理工大学数学与统计学院) Digital Intelligence Center, Fuzhou University Affiliated Provincial Hospital, China(福州市附属省医院数字智能中心) School of Computer Science and Engineering, UNSW, Australia(新南威尔士大学计算机科学与工程学院)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 SALP-CG通过统一标准和自动化方法,实现在线医疗对话数据的分类与敏感性分级,提升健康数据治理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10458 2026-01-16 cs.HC cs.LG stat.CO 83%

LangLasso: Interactive Cluster Descriptions through LLM Explanation

LangLasso:通过LLM解释实现交互式聚类描述

Raphael Buchmüller, Dennis Collaris, Linhao Meng, Angelos Chatzimparmpas

机构 * University of Konstanz(康斯坦茨大学) Utrecht University(乌得勒支大学) Eindhoven University of Technology(埃因霍温理工大学)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 LangLasso通过LLM生成交互式自然语言描述,使非专家也能理解聚类结构并整合外部知识。

Comments This manuscript is accepted for publication in VIS 2025 VISxGenAI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18454 2026-01-16 cs.CL cs.AI cs.LG 82%

Fairness Definitions in Language Models Explained

语言模型中的公平性定义解析

Zhipeng Yin, Zichong Wang, Avash Palikhe, Wenbin Zhang

机构 * Florida International University(佛罗里达国际大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文系统解析了语言模型中公平性定义的分类与应用,介绍了现有公平性概念及新分类方法,通过实验探讨其实际影响,旨在推动公平性研究的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10646 2026-01-16 cs.SE cs.AI 81%

CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance

CodeAssistBench (CAB): 用于多轮聊天式代码协助的数据库与基准测试

Myeongsoo Kim, Shweta Garg, Baishakhi Ray, Varun Kumar, Anoop Deoras

机构 * AWS AI Labs(AWS人工智能实验室)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 CodeAssistBench首次提出多轮、项目导向的编程协助基准测试,揭示当前LLM在真实项目情境中的性能差距。

Comments Accepted to NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10421 2026-01-16 cs.CL cs.AI 81%

Are Language Models Models?

语言模型是模型吗?

Philip Resnik

专题命中 评测与基准 :language model(title);LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文质疑语言模型是否应被视为认知模型,指出其在不同层次上的不足,并强调其作为工具的适用性而非认知模型的合理性。

Comments 5 pages. This is an invited commentary under review at Behavioral and Brain Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09713 2026-01-16 cs.CL 79%

LLM-Driven Preference Data Synthesis for Proactive Prediction of the Next User Utterance in Human-Machine Dialogue

基于大语言模型的偏好数据合成用于人机对话中下一用户话语的前瞻性预测

Jinqiang Wang, Huansheng Ning, Jianguo Ding, Tao Zhu, Liming Chen, Chris Nugent

机构 * School of Computer & Communication Engineering, University of Science and Technology Beijing(计算机与通信工程学院,北京科技大学) Department of Computer Science, Blekinge institute of Technology(计算机科学系,布莱金厄理工大学) School of Computer Science, University of South China(计算机科学学院,南方大学) School of Computer Science and Technology, Dalian University of Technology(计算机科学与技术学院,大连理工大学) School of Computing, Ulster University(计算机科学学院,乌斯特大学)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL

AI总结 ProUtt通过意图树建模和偏好数据合成,提升人机对话中下一次用户发言的预测精度。

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21357 2026-01-16 cs.CV cs.LG 79%

AgriFM: A Multi-source Temporal Remote Sensing Foundation Model for Agriculture Mapping

AgriFM:一种多源时序遥感基础模型用于农业制图

Wenyuan Li, Shunlin Liang, Keyan Chen, Yongzhe Chen, Han Ma, Jianglei Xu, Yichuan Ma, Shikang Guan, Husheng Fang, Zhenwei Shi

机构 * Jockey Club STEM Lab of Quantitative Remote Sensing, Department of Geography, The University of Hong Kong, Hong Kong, China(香港大学地理系金钟STEM实验室) Department of Aerospace Intelligent Science and Technology, School of Astronautics, Beihang University, Beijing, China(北航航天智能科学与技术学院) School of Remote Sensing and Information Engineering, Wuhan University, China(武汉大学遥感与信息工程学院)

专题命中 评测与基准 :foundation model(title,abstract);分类 cs.LG

AI总结 AgriFM是一种多源时序遥感基础模型,通过改进的Video Swin Transformer架构实现多尺度时空特征提取,提升农业制图的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16821 2026-01-16 cs.NI cs.LG eess.SP 79%

LLM-Based Emulation of the Radio Resource Control Layer: Towards AI-Native RAN Protocols

基于大语言模型的无线资源控制层仿真:迈向AI原生的无线接入网络协议

Ziming Liu, Bryan Liu, Alvaro Valcarce, Xiaoli Chu

机构 * School of Electrical and Electronic Engineering, The University of Sheffield(电子与电气工程学院,谢菲尔德大学) Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 评测与基准 :LLM(title);prompting(abstract);分类 cs.LG

AI总结 本文提出基于大语言模型的RRC层仿真方法,通过参数高效适应和模式受限提示,实现高精度协议仿真,展示了AI原生RAN协议的可行性。

Comments This work has been submitted to the IEEE for possible publication. Focuses on applying LLMs to 5G RRC protocol generation; primary: cs.NI; cross-list: eess.SP, cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08355 2026-01-16 cs.CV 78%

Semantic Misalignment in Vision-Language Models under Perceptual Degradation

视觉-语言模型在感知退化下的语义错位

Guo Cheng

专题命中 评测与基准 :language model(title,abstract)

AI总结 本研究探讨了视觉-语言模型在感知退化下的语义错位问题,发现传统分割指标的下降并未影响下游行为,揭示了像素鲁棒性与多模态语义可靠性之间的脱节。

Comments 10 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10589 2026-01-16 cs.CR cs.CL 77%

Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay

做自己的红队:通过自play和反思经验回放实现安全对齐

Hao Wang, Yanting Wang, Hao Li, Rui Li, Lei Sha

机构 * Beihang University(北京航空航天大学) Peking University(北京大学) Zhongguancun Laboratory(中关村实验室)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 通过自play和反思经验回放机制,使模型自主进化防御能力,提升安全对齐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10496 2026-01-16 cs.SE cs.AI 77%

Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs

模型看见,模型做?基于暴露的bug-修复偏好评估

Ali Al-Kaswan, Claudio Spiess, Prem Devanbu, Arie van Deursen, Maliheh Izadi

机构 * Delft University of Technology(代尔夫特理工大学) University of California at Davis(加州大学戴维斯分校)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究发现LLM在生成代码时更倾向于重复bug而非修复,暴露于bug的示例加剧了这一倾向,而修复暴露则仅有微小改进,揭示了LLM可能传播记忆错误的风险。

Comments MSR 2026 Technical Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10033 2026-01-16 cs.CL 77%

EmplifAI: a Fine-grained Dataset for Japanese Empathetic Medical Dialogues in 28 Emotion Labels

EmplifAI:一个用于日语共情医疗对话的细粒度数据集,包含28种情绪标签

Wan Jou She, Lis Kanashiro Pereira, Fei Cheng, Sakiko Yahata, Panote Siriaraya, Eiji Aramaki

机构 * Kyoto Institute of Technology, Japan(京都技术大学) National Institute of Information and Communications Technology (NICT), Japan(信息通信技术国家研究所) Kyoto University, Japan(京都大学) Nara Institute of Science and Technology, Japan(奈良科学技术大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 EmplifAI是一个用于日语共情医疗对话的细粒度数据集,包含28种情绪标签,通过微调提升了日本LLM的流畅性和共情能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09883 2026-01-16 cs.AI 77%

Beyond Rule-Based Workflows: An Information-Flow-Orchestrated Multi-Agents Paradigm via Agent-to-Agent Communication from CORAL

超越基于规则的工作流:通过CORAL的agent-to-agent通信实现的信息流协调多智能体范式

Xinxing Ren, Quagmire Zang, Caelum Forder, Suman Deb, Ahsen Tahir, Roman J. Georgio, Peter Carroll, Zekun Guo

机构 * Coral Protocol(Coral协议) Brunel University of London(伦敦布鲁内尔大学) Universitéit Lëtzebuerg(列日大学) University of Hull(霍尔姆大学) National University of Computer and Emerging Sciences(国家计算机与新兴科学大学)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出了一种基于信息流协调的多智能体范式,通过agent-to-agent通信替代传统工作流,提升任务处理的灵活性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02737 2026-01-16 cs.CV 75%

Unveiling and Bridging the Functional Perception Gap in MLLMs: Atomic Visual Alignment and Hierarchical Evaluation via PET-Bench

揭示和弥合MLLMs中的功能感知差距:通过PET-Bench实现原子视觉对齐和分层评估

Zanting Ye, Xiaolong Niu, Xuanbin Wu, Xu Han, Shengyuan Liu, Jing Hao, Zhihao Peng, Hao Sun, Jieqin Lv, Fanghu Wang, Yanchao Huang, Hubing Wu, Yixuan Yuan, Habib Zaidi, Arman Rahmim, Yefeng Zheng, Lijun Lu

机构 * School of Biomedical Engineering, Southern Medical University(生物医学工程学院,南方医科大学) School of Biomedical Engineering, Shanghai Jiaotong University(生物医学工程学院,上海交通大学) Department of Electronic Engineering, Chinese University of Hong Kong(电子工程系,中国香港大学) Faculty of Dentistry, The University of Hong Kong(牙科学院,香港大学) Department of Nuclear Medicine, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine(核医学科,广州中医药大学第二附属医院) PET Center, Department of Nuclear Medicine, Guangdong Provincial People’s Hospital, Southern Medical University(PET中心,核医学科,广东省人民医院,南方医科大学) Department of Nuclear Medicine, Nanfang Hospital, Southern Medical University(核医学科,南芳医院,南方医科大学) Division of Nuclear Medicine and Molecular Imaging, Geneva University Hospitals(核医学与分子影像学部,日内瓦大学医院) Departments of Radiology, Physics, and Biomedical Engineering, The University of British Columbia(放射学、物理和生物医学工程系,不列颠哥伦比亚大学) Medical Artificial Intelligence Laboratory, Westlake University(医学人工智能实验室,西湖大学)

专题命中 评测与基准 :large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出AVA方法,通过原子视觉对齐解决MLLMs在功能成像中的感知差距,提升诊断准确性14.83%。

Comments 9 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04675 2026-01-16 cs.AI cs.CL cs.LG 75%

Scalable Oversight for Superhuman AI via Recursive Self-Critiquing

通过递归自我批评实现超人类AI的可扩展监督

Xueru Wen, Jie Lou, Xinyu Lu, Junjie Yang, Yanjiang Liu, Yaojie Lu, Debing Zhang, Xing Yu

机构 * Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所信息处理实验室) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 评测与基准 :SFT(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出通过递归自我批评实现超人类AI的可扩展监督,探索批评的批评比直接批评更容易,并验证递归批评在AI监督中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10524 2026-01-16 cs.AI 70%

Diagnosing Generalization Failures in Fine-Tuned LLMs: A Cross-Architectural Study on Phishing Detection

在微调大语言模型中诊断泛化失败: phishing检测的跨架构研究

Frank Bobe, Gregory D. Vetaw, Chase Pavlick, Darshan Bryner, Matthew Cook, Jose Salas-Vernis

机构 * Naval Surface Warfare Center Panama City Division(海军水面 warfare 中心巴拿马城分部)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过跨架构研究,揭示了微调大语言模型在钓鱼检测任务中泛化失败的原因,发现架构与数据多样性协同作用、架构依赖性及某些架构固有泛化能力。

Comments 16 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10323 2026-01-16 cs.CV cs.CL 70%

ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding

ROMA: 实时多模态助手与交互式流式理解

Xueyun Tian, Wei Li, Bingbing Xu, Heng Dong, Yuanzhuo Wang, Huawei Shen

机构 * CAS Key Laboratory of AI Safety, Institute of Computing Technology, CAS, Beijing, China(中国科学院人工智能安全重点实验室,计算技术研究所,中国科学院,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) Tsinghua University, Beijing, China(清华大学,北京,中国)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 ROMA通过实时多模态处理和交互式流式理解,实现统一的反应性和主动性交互,展现强大的多模态处理能力。

Comments Our project page is available at https://eureka-maggie.github.io/ROMA_show

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13142 2026-01-16 cs.AI cs.HC 70%

Can LLMs Understand What We Cannot Say? Measuring Multilevel Alignment Through Abortion Stigma Across Cognitive, Interpersonal, and Structural Levels

LLMs能否理解我们无法表达的内容?通过堕胎污名的多层级对齐进行测量

Anika Sharma, Malavika Mampally, Chidaksh Ravuru, Kandyce Brennan, Neil Gaikwad

机构 * Society-Centered AI Lab(以社会为中心的人工智能实验室) School of Nursing(护理学院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究发现LLMs在认知、人际和结构性层面对堕胎污名的理解不一致,缺乏对多维度心理构造的 coherent 理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00010 2026-01-16 cs.CL 70%

PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization

PlotCraft: 推动大语言模型在复杂和交互式数据可视化中的极限

Jiajun Zhang, Jianke Zhang, Zeyu Cui, Jiaxi Yang, Lei Zhang, Binyuan Hui, Qiang Liu, Zilei Wang, Liang Wang, Junyang Lin

机构 * USTC(University of Science and Technology of China) THU(Tsinghua University) Alibaba Group(阿里巴巴集团) CASIA(Chinese Academy of Sciences Institute of Automation) SIAT(State Key Laboratory of Information Security)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 PlotCraft 提出了一种新的基准和数据集,用于评估大语言模型在复杂和交互式数据可视化任务中的性能,展示了 PlotCraftor 在复杂任务中的显著改进。

详情

展开后加载摘要…

URL PDF HTML 收藏