CommentsAccepted at WISE 2026 (27th International Conference on Web Information Systems Engineering), Special Track on Innovative Web Technologies. 15 pages. This version has been accepted for publication after peer review but is not the Version of Record; the Version of Record will appear in the Springer LNCS proceedings
TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts
TailSieve:面向大语言模型(LLM)rollout的部分rollout引导式长尾路由
Tianqi Xu, Lu Lv, Haoyang Huang, Wenjie Huang, Zhanming Shen, Yuhao Shen, Baolin Zhang, Xinyi Hu, Shuang Ge, Jun Dai, Tianyu Liu, Suorong Yang, Zhikai Li, Ye Bai, Jun Zhang, Lei Chen, Yue Li, Mingchen Wan
机构
*
Qwen Business Unit of Alibaba(阿里巴巴通义千问业务部)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Zhejiang University(浙江大学)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization Levels
Hydra:面向边缘SoC代际、后端及量化级别的LLM推理的阶段感知工作负载表征
Amir Taherin, Sana Taghipour Anvari, Charles Amante, Yixiao Chen, Ruben Noroian, Zlatan Feric, Nicolas Bohm Agostini, Pu Zhao, José Cano, Bin Ren, Yanzhi Wang, David Kaeli
机构
*
Nanjing University of Science and Technology(南京理工大学)
;
C 2 \text{C}^{2} DL, Institute of Automation, Chinese Academy of Sciences(C2 DL,自动化研究所,中国科学院)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
School of Computer Science, University of Sydney(悉尼大学计算机学院)
专题命中
效率与部署
:LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);post-training(abstract)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
大语言模型还能自我解释吗?探讨量化对自我解释的影响
Qianli Wang, Nils Feldhus, Pepa Atanasova, Fedor Splitt, Simon Ostermann, Sebastian Möller, Vera Schmitt
机构
*
Quality and Usability Lab, Technische Universität Berlin(柏林技术大学质量与可用性实验室)
;
University of Copenhagen(哥本哈根大学)
;
Saarland Informatics Campus(萨尔州信息学校园)
;
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
;
Centre for European Research in Trusted AI (CERTAIN)(可信AI欧洲研究中心)
;
BIFOLD – Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究院)
专题命中
效率与部署
:large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG
The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models
CommentsSubstantially revised version with a formal non-identifiability result for the Inference Attribution Problem, expanded related work, and extended analysis of Probability Placement, auditing, and runtime transparency
Comments16 pages, 9 figures, 12 tables. Published in Proceedings of ACL 2026 (Volume 1: Long Papers), pages 11480-11497. Demo and API: this https URL (https://witness-ai-jpt-ner.hf.space/)
Multi-Turn Reasoning LLMs for Task Offloading in Mobile Edge Computing
面向移动边缘计算的任务卸载的多轮推理LLM
Ning Yang, Chuangxin Cheng, Haijun Zhang
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Machano-Electronic Engineering, Xidian University(西安电子科技大学机电工程学院)
;
School of Computer and Communication Engineering, University of Science and Technology Beijing(北京科技大学计算机与通信工程学院)
专题命中
效率与部署
:LLM(title_cn);SFT(abstract,abstract_cn);large language model(abstract);language model(abstract)