[论文] PRIME: Perception Feedback with Situational Memory Embeddings in VLA M…

## 论文概要 **研究领域**: CV **作者**: Erik Deinzer, Naya Baslan,...

论文概要

研究领域: CV 作者: Erik Deinzer, Naya Baslan, Luca Paparusso, Narunas Vaskevicius, Peter Knott, Luigi Palmieri 发布时间: 2026-09-18 arXiv: 2609.22040

中文摘要

当前用于自动驾驶的视觉-语言-动作(VLA)模型主要通过感知-推理-规划层级的前馈推理运行。现代架构虽在感知模块内保持时间递归,但早期感知对下游推理和导航目标视而不见,以无差别方式处理视觉输入,不优先处理由先前决策告知的线索。为弥合这一差距,本文提出 PRIME——一种学习到的反馈机制,用一种新颖的情境记忆(Situational Memory)来条件化 VLA 的感知查询。通过对过去 L 步窗口内的感知、推理、导航目标和预测行为的潜在表征进行跨注意力聚合,PRIME 以极小的计算成本实现意图驱动的感知注意:最多仅增加 2,970 万参数(73 亿参数基座模型的 0.41%)。在 Bench2Drive 闭环基准上评估,PRIME 取得最优的驾驶分数 82.47(比 ORION 高 4.73)和 60.00% 的成功率(高 5.38 个百分点),是在 Think2Drive 演示数据上训练的所有已发表 VLA 中报告的最高驾驶分数。

原文摘要

Current Vision-Language-Action (VLA) models for autonomous driving operate primarily through feedforward inference across the perception–reasoning–planning hierarchy. While modern architectures maintain temporal recurrence within the perceptual module, early perception remains blind to downstream reasoning and navigation goals, processing visual inputs agnostically without prioritizing cues informed by prior decisions. To bridge this gap, this paper introduces PRIME, a learned feedback mechanism that conditions the VLA perceptual queries on a novel Situational Memory. By aggregating latent representations of past perception, reasoning, navigation goals, and predicted behaviors across an L-step window via cross-attention, PRIME enables intent-driven perceptual attention at minimal computa…

— 自动采集于 2026-09-22

#论文 #arXiv #CV #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1