[论文] TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 40…

## 论文概要 **研究领域**: CV **作者**: Hengyi Xie, Chenfei Yao, X...

论文概要

研究领域: CV 作者: Hengyi Xie, Chenfei Yao, Xianjin Wu, Xuanyang Xi, Yiping Tang, Di Xu, Yingying Zhu, Dingkang Liang, Xiang Bai, Han Ding 发布时间: 2026-07-29 arXiv: 2607.27205

中文摘要

视觉-语言-动作(VLA)模型通常采用以LLM为中心的 V→L→A 路径:先将视觉观测投影到大语言模型的表征空间,再解码为机器人动作。虽然有效,但每一步策略调用都带来大量计算和内存开销。本文提出 TurboVLA,一种全新的VLA范式,将传统 V→L→A 路径重构为直接的 V+L→A 映射。TurboVLA 不再用大语言模型作为感知与动作之间的中心接口,而是独立编码视觉观测和语言指令,通过轻量级双向视觉-语言交互直接交换信息,并用紧凑的解码器预测连续动作块。这种简洁设计直接从视觉和语言特征构建任务条件表征,显著降低了VLA推理的计算和内存成本。在LIBERO上,TurboVLA仅用0.2B参数、31.2毫秒推理延迟、0.9 GB显存,在消费级RTX 4090上达到97.7%的平均成功率,匹敌甚至超越参数量大得多的VLA策略。

原文摘要

Vision-language-action (VLA) models commonly adopt an LLM-centric V -> L -> A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V -> L -> A pathway as a direct V + L -> A mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks wit…

自动采集于 2026-07-31

#论文 #arXiv #CV #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1