返回全部动态

N_0-VTLA:带潜在触觉令牌的视觉-触觉-语言-动作模型

原标题:N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

Hugging Face Daily Papers一手来源研究质量 87

AI 摘要

N_0-VTLA 是一个视觉-触觉-语言-动作(VTLA)基础模型,通过触觉感知和反馈控制实现精细的接触丰富操作,并支持从部署数据中进行离线策略改进。该模型采用视觉-触觉预训练、分阶段触觉通路集成和优势条件离线强化学习(ALTER)的训练方案,在九个真实机器人任务中全部获胜,并在二十任务模拟套件中达到63.8%的平均成功率,显著优于最强基线。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Abstract We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathw


发布时间:
抓取时间:2026-08-03 10:18
来源机构:Hugging Face
阅读原文huggingface.co