arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

2026-04-10 至 2026-04-10 共收录 4
2604.07655 2026-04-10 cs.LG cs.CL

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs

顾问式守护者:推进下一代守护模型以实现可信的大语言模型

Yue Huang, Haomin Zhuang, Jiayi Ye, Han Bao, Yanbo Wang, Hang Hua, Siyuan Wu, Pin-Yu Chen, Xiangliang Zhang

机构 * University of Notre Dame(圣母大学) University of California, Los Angeles(加利福尼亚大学洛杉矶分校) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室) IBM Research(IBM研究院)

AI总结 本文提出Guardian-as-an-Advisor框架,通过预测风险标签和解释并前置建议,提升模型安全性和实用性,同时降低误拒率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07092 2026-04-10 cs.CV

Location Is All You Need: Continuous Spatiotemporal Neural Representations of Earth Observation Data

位置即一切:连续时空神经表示法用于地球观测数据

Mojgan Madadikhaljan, Jonathan Prexl, Isabelle Wittmann, Conrad M Albrecht, Michael Schmitt

机构 * University of the Bundeswehr Munich(慕尼黑联邦国防军大学) IBM Research – Europe(IBM欧洲研究院) Columbia University(哥伦比亚大学) German Aerospace Center (DLR)(德国航空航天中心)

AI总结 本文提出LIANet,通过连续时空神经场建模多时相遥感数据,无需原始数据即可实现卫星影像重建,并在各种下游任务中取得与从头训练或使用现有GFMs相当的性能。

Comments Updated the affiliation of one of the authors, no changes to the technical content

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15118 2026-04-10 cs.CV

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents

VAREX:多模态结构提取文档的基准

Udi Barzelay, Ophir Azulai, Inbar Shapira, Idan Friedman, Foad Abo Dahood, Madison Lee, Abraham Daniels

机构 * IBM Research(IBM研究院)

AI总结 VAREX基准通过四类输入模态评估多模态基础模型在政府表格结构数据提取中的性能,揭示参数规模、输入格式对提取准确率的影响,强调结构合规性是关键瓶颈。

Comments 9 pages, 4 figures, 4 tables, plus 12-page supplementary. Dataset: https://huggingface.co/datasets/ibm-research/VAREX Code: https://github.com/udibarzi/varex-bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04500 2026-04-10 cs.AI cs.RO

"Don't Do That!": Guiding Embodied Systems through Large Language Model-based Constraint Generation

别这样做!

Amin Seffo, Aladin Djuhera, Masataro Asai, Holger Boche

机构 * Technical University Munich(慕尼黑工业大学) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)

AI总结 本文提出STPR框架,利用LLM将自然语言中的约束转化为可执行Python代码,解决复杂约束的翻译问题,并在仿真环境中验证其高效性和兼容性。

Comments ICLR 2026 Workshop -- Agentic AI in the Wild: From Hallucinations to Reliable Autonomy

详情

展开后加载摘要…

URL PDF HTML 收藏