返回全部动态

WCM:用于视觉-语言-动作强化学习的世界评论家模型

原标题:WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

Hugging Face Daily Papers一手来源研究质量 83

AI 摘要

WCM是一种用于视觉-语言-动作(VLA)模型强化学习的世界评论家模型,通过轻量级LeJEPA架构联合预测未来潜在状态和估计价值,解决部分可观测性下的价值估计问题。实验在149个任务和四个基准上取得最先进性能,并在七个真实世界操作任务中验证了部署稳定性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Abstract Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM backbone latents, which is a fundamental mismatch with the partially observable nature of robot control. A naive approach to incorporate o


发布时间:
抓取时间:2026-08-04 11:04
来源机构:Hugging Face
阅读原文huggingface.co