Token导航 LogoToken导航TokenDH.com
研究检索只读clawhub未标认证来源可访问clear审计通过

robotics-vlarobotics VLA 搜索

Agent Skill

robotics-vla 用于查找、检索和筛选相关信息,适合在 OpenClaw 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

6,336

周安装

264

GitHub Stars

1

下载量

2,112
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:robotics-vla(robotics VLA 搜索)
来源仓库:https://github.com/arden2010/robotics-vla
安装命令:
openclaw skills install robotics-vla
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install robotics-vla

简介

robotics-vla 提供视觉-语言-动作(VLA)机器人基础模型的专业指导,涵盖架构设计、训练与部署全流程。

  • 适用于需要构建或优化 VLA 系统的开发者和研究者,支持多模态任务场景下的模型应用。
  • 通过关键词检索和来源仓库文档获取技术细节,结合 OpenClaw 环境进行技能调用。
  • 安装前需确认权限范围和维护状态,注意可能涉及联网访问和模型推理资源消耗。
  • 建议参考原始 README 和 SKILL.md 了解具体接口和使用限制。

SKILL.md

name
robotics-vla
description
>

Robotics VLA Skill

Expert guidance for building generalist robot policies using Vision-Language-Action (VLA) flow models, based on the π0 architecture.

Core Architecture

π0 model = VLM backbone + action expert + flow matching

ComponentDetail
VLM backbonePaliGemma (3B) — provides visual + language understanding
Action expertSeparate transformer weights (~300M) for robot state + actions
Total params~3.3B
Action outputChunks of H=50 actions; 50Hz or 20Hz robots
Inference speed~73ms on RTX 4090

See references/architecture.md for full technical details (attention masks, flow matching math, MoE design).

Training Pipeline

Two-phase approach (mirrors LLM training):

  1. Pre-training → broad physical capabilities + recovery behaviors across many tasks/robots
  2. Fine-tuning → fluent, task-specific execution on target task

Key rule: combining both phases outperforms either alone. Pre-training gives robustness; fine-tuning gives precision.

See references/training.md for data mixture ratios, loss functions, and fine-tuning dataset sizing.

Action Representation

Use flow matching, not autoregressive discretization.

  • Flow matching models continuous action distributions → essential for high-frequency dexterous control
  • Autoregressive token prediction (e.g. RT-2 style) cannot produce action chunks efficiently
  • Action chunks allow open-loop execution at 50Hz without temporal ensembling

Multi-Embodiment Support

Single model handles 7+ robot configurations via:

  • Zero-padding smaller action spaces to match the largest (17-dim)
  • Shared VLM backbone; embodiment-specific behavior learned via data
  • Weighted task sampling: n^0.43 to handle imbalanced data across robot types

See references/embodiments.md for robot platform specs and action space details.

High-Level Policy Integration

For long-horizon tasks, use a two-tier approach:

  • High-level VLM: decomposes task ("bus the table") → subtasks ("pick up napkin")
  • Low-level π0: executes each subtask as a language-conditioned action sequence

Analogous to SayCan. Intermediate language commands significantly boost performance vs. flat task descriptions.

Related & Complementary Research (2025)

π0 has been extended and complemented by several key works. See references/related-work.md for the full landscape, including:

  • π0-FAST / π0.5 / π0.6 — direct successors with faster training, open-world generalization, and RL fine-tuning
  • RTC — async action chunking to eliminate inference pauses (plug-in, no retraining)
  • UniVLA — unsupervised action extraction from raw video (no action labels needed)
  • ManiFlow / Streaming Flow — smoother action generation
  • GR00T N1, Helix, OpenVLA-OFT, DiVLA, RDT-1B — parallel approaches from NVIDIA, Figure AI, and academia

Evaluation Checklist

When evaluating a robot manipulation policy:

  • [ ] Out-of-box generalization (no fine-tuning) vs. baselines
  • [ ] Language following accuracy with flat / human-guided / HL commands
  • [ ] Fine-tuning efficiency (success rate vs. hours of data)
  • [ ] Complex multi-stage tasks (5–20 min, recovery from failure)
  • [ ] Compare: OpenVLA, Octo, ACT, Diffusion Policy as baselines

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

76.21%
按下载量换算1,610

安全审计

VirusTotal

通过

ClawScan

通过

Static analysis

通过

权限和风险

只读

该 Skill 主要提供规则、说明或参考内容,本身偏只读;真正读写文件、联网或执行命令仍取决于宿主 Agent 的任务。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills