TACTIC:面向路侧激光雷达攻击的时序与上下文感知大语言模型战术规划
TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks
- University of Michigan(密歇根大学)
- Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TACTIC利用多模态大语言模型结合路侧感知,自适应选择攻击基元,在CARLA试验中实现100%碰撞率,显著优于固定策略,并大幅缩短场景突变响应时间。
AI中文摘要:
物理激光雷达攻击通常使用固定基元和手动选择的参数进行评估,尽管这些攻击强烈依赖于周围交通状况。我们提出TACTIC,一个场景感知框架,利用多模态大语言模型(MLLM)协调状态自适应的路侧激光雷达攻击。在灰盒威胁模型下,TACTIC仅依赖攻击者操作的路侧感知栈,无需访问受害者激光雷达的原生点云或内部处理。局部感知提供度量车辆状态,而MLLM将这些测量与路侧图像结合,推断关系型交通上下文并构建语义场景图。基于该表示,TACTIC选择并配置两种互补基元:\u201c推离\u201d(push-away),用于偏移前车的感知距离;以及\u201c幻影障碍物制动\u201d(phantom-obstacle braking),通过注入障碍物触发紧急制动。测量的交通状态和经验校准的约束将生成的战术限制在物理可行的操作区域内。为适应MLLM延迟,TACTIC异步重叠推理与执行,同时高速局部感知检测场景变化并触发重新规划。在280次随机CARLA试验中,完整策略实现100%碰撞率,而固定规则为35%,随机选择为60%,使用默认参数进行模式选择的受限LLM为75%。联合物理与图像输入实现100%成功率,而仅物理测量为65%,仅图像为75%;异步Δ刷新将场景突变响应时间从7.4秒降至2.0秒。这些结果表明,场景依赖的战术规划能够暴露固定攻击策略可能遗漏的上下文敏感激光雷达故障模式。
英文摘要:
Physical LiDAR attacks are often evaluated using fixed primitives and manually selected parameters, despite their strong dependence on surrounding traffic. We present TACTIC, a scene-aware framework that uses a multimodal large language model (MLLM) to coordinate state-adaptive roadside LiDAR attacks. Under a gray-box threat model, TACTIC relies only on an attacker-operated roadside perception stack, without accessing the victim LiDAR's native point clouds or internal processing. Local perception provides metric vehicle states, while the MLLM combines these measurements with roadside imagery to infer relational traffic context and construct a semantic scene graph. Based on this representation, TACTIC selects and configures two complementary primitives: \emph{push-away}, which shifts the perceived range of a lead vehicle, and \emph{phantom-obstacle braking}, which triggers emergency braking through obstacle injection. Measured traffic states and empirically calibrated constraints ground the generated tactics in physically feasible operating regions. To accommodate MLLM latency, TACTIC overlaps reasoning and execution asynchronously while high-rate local perception detects scene changes and triggers replanning. Across 280 randomized CARLA trials, the full policy achieves a 100% collision rate, compared with 35% for a fixed rule, 60% for random selection, and 75% for a restricted LLM using mode selection with default parameters. Joint physical-and-image input achieves 100% success, versus 65% with physical measurements alone and 75% with imagery alone, while asynchronous $Δ$ refresh reduces scene-mutation response from 7.4 s to 2.0 s. These results show that scene-dependent tactical planning can expose context-sensitive LiDAR failure modes that fixed attack policies may miss.