返回全部动态
WCM:用于视觉-语言-动作强化学习的世界评论家模型
原标题:WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
AI 摘要
WCM是一种用于视觉-语言-动作(VLA)模型强化学习的世界评论家模型,通过轻量级LeJEPA架构联合预测未来潜在状态和估计价值,解决部分可观测性下的价值估计问题。实验在149个任务和四个基准上取得最先进性能,并在七个真实世界操作任务中验证了部署稳定性。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Abstract Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM backbone latents, which is a fundamental mismatch with the partially observable nature of robot control. A naive approach to incorporate o
发布时间:—
抓取时间:2026-08-04 11:04
来源机构:Hugging Face