返回全部动态

Capek 0.5:面向具身智能的执行中心视觉语言模型

原标题:Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

Hugging Face Daily Papers一手来源研究质量 84

AI 摘要

Capek 0.5 是一个面向具身智能的执行中心视觉语言模型,围绕执行能力分类法组织训练,涵盖空间推理、时间理解、动作引导和状态验证四类能力。每个能力先通过强化学习由共享骨干网络训练出专家,再通过权重空间合并和路由策略空间蒸馏整合为单一推理模型。Capek 0.5 在 2B 和 35B-A3B 规模上进行了评估,在多数基准上优于初始化模型,并能在模拟环境中完成闭环任务执行。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence Abstract Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands requires complementary capabilities that differ in supervision signals, prediction formats, and verification criteria. Existing appr


发布时间:—
抓取时间:2026-08-10 17:49
来源机构:Hugging Face
阅读原文huggingface.co