arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

ByteDance(字节跳动)

2026-04-21 至 2026-04-21 共收录 6
2604.18375 2026-04-21 cs.CL cs.AI

IceBreaker for Conversational Agents: Breaking the First-Message Barrier with Personalized Starters

对话代理的IceBreaker:通过个性化开场白打破第一消息障碍

Hongwei Zheng, Weiqi Wu, Zhengjia Wang, Guanyu Jiang, Haoming Li, Tianyu Wu, Yongchun Zhu, Jingwu Chen, Feng Zhang

机构 * ByteDance(字节跳动)

AI总结 本文提出IceBreaker,通过生成个性化开场白帮助对话代理克服用户初始对话时的障碍,提升用户活跃度和点击率。

Comments ACL 2026 Accepted Paper (Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18292 2026-04-21 cs.AI cs.CL

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

Agent-World:通过可扩展环境提升进化通用智能

Guanting Dong, Junting Lu, Junjie Huang, Wanjun Zhong, Longxiang Liu, Shijue Huang, Zhenyu Li, Yang Zhao, Xiaoshuai Song, Xiaoxi Li, Jiajie Jin, Yutao Zhu, Hanbin Wang, Fangyu Lei, Qinyu Luo, Mingyang Chen, Zehui Chen, Jiazhan Feng, Ji-Rong Wen, Zhicheng Dou

机构 * Renmin University of China(中国人民大学) ByteDance Seed(字节跳动种子)

AI总结 本文提出Agent-World,通过可扩展环境实现通用智能进化,展示其在23个挑战性基准测试中超越现有模型和基线。

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21228 2026-04-21 cs.CL cs.AI

ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction Following

ImpRIF:更强的隐式推理导致更好的复杂指令遵循

Yuancheng Yang, Lin Yang, Xu Wang, Chao Tong, Haihua Yang

机构 * ByteDance China(字节跳动中国) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室)

AI总结 本文提出ImpRIF方法,通过构建可验证推理图提升大语言模型对隐式推理指令的理解,从而增强复杂指令遵循能力,在五个基准测试中表现优异。

Comments Accepted at ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16756 2026-04-21 cs.AI cs.CL cs.CV cs.RO eess.AS

End-to-end Listen, Look, Speak and Act

端到端听、看、说和行动

Siyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Lu Lu, Chao Zhang

机构 * Tsinghua University(清华大学) ByteDance(字节跳动)

AI总结 本文提出ELLSA模型,首个端到端全双工系统,实现视觉、文本、语音和行动的同步感知与生成,支持对话轮替、指令拒绝等复杂交互行为,推动更自然的人工智能发展。

Comments 22 pages, 8 figures

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19857 2026-04-21 cs.SD cs.CV

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts

FoleyDirector: 通过结构化脚本实现视频到音频生成的细粒度时间控制

You Li, Dewei Zhou, Fan Ma, Fu Li, Dongliang He, Yi Yang

机构 * The State Key Lab of Brain-Machine Intelligence, Zhejiang University(浙江大学脑机智能State Key Lab) ReLER, CCAl, Zhejiang University(浙江大学ReLER, CCAl) Intelligent Creation, ByteDance, China(字节跳动智能创作)

AI总结 本文提出FoleyDirector框架,通过结构化时间脚本提升视频到音频生成的细粒度时间控制能力,同时保持音频质量并支持无缝切换生成与受控合成。

Comments Accepted at IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026, 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09885 2026-04-21 cs.CV

The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge

你依赖得越少,学到的越多:从稀疏、未定位图像中合成新视角的方法,仅需最小的3D知识

Haoru Wang, Kai Ye, Minghan Qin, Yangyan Li, Wenzheng Chen, Baoquan Chen

机构 * Peking University(北京大学) ByteDance Seed(字节跳动种子) Ant Group(蚂蚁集团) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

AI总结 本文探讨了在数据量增加时,依赖较少3D知识的方法性能提升更快的现象,提出了一种无需显式场景结构和姿态标注的端到端NVS框架,实现了更高效的视角合成。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏