Token导航 LogoToken导航TokenDH.com
效率只读clawhub未标认证来源可访问clear审计通过

see-video看视频

Agent Skill

用于辅助视频生成、动画合成、脚本化剪辑或 Remotion 等视频项目开发。它适合让 Agent 组织镜头、生成素材说明、维护合成代码或排查渲染问题。使用时需要确认分辨率、时长、素材路径和导出格式;涉及外部素材、人物肖像或商业发布时,应先核对版权授权和内容审核要求。

总安装

3,461

周安装

140

GitHub Stars

公开资料未说明

下载量

1,086
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:see-video(看视频)
来源仓库:https://github.com/john-ver/see-video
安装命令:
openclaw skills install see-video
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install see-video

简介

用于辅助视频内容理解与视觉素材处理。see-video 属于效率类 Skill,可作为该场景下的辅助能力补充。

  • 适合分析视频帧、提取关键画面或辅助剪辑决策。
  • 使用时需确认分辨率、时长及导出格式要求。适用宿主包括 OpenClaw,接入前应确认版本、权限和运行环境要求。
  • 涉及人物肖像或商业素材时应核对版权授权情况。
  • 将帧图像网格注入 LLM 上下文以增强内容理解。

SKILL.md

name
see-video
description
Use when the user sends a video file or asks about video content. Extracts frames and injects them as an image grid directly into the LLM context — no proxy model, no description handoff. uniform mode (default): evenly spaced sampling. highlight mode: scene-change biased sampling. ⚠️ Requires a vision-capable (multimodal) model.
metadata

see-video

Extract frames from a video and inject them as a grid image + XML timestamps into LLM context.

Setup (first time only)

cd <skill directory>
npm install

Usage

node {baseDir}/scripts/inject.mjs <video_path> [--mode uniform|highlight] [--start N] [--end N]

On success, outputs JSON to stdout:

{
  "gridPath": "/tmp/video_llm-frames.jpg",
  "description": "<video_frames>...</video_frames>",
  "duration": 1326,
  "frameCount": 28,
  "layout": { "cols": 4, "rows": 7, "cellW": 384, "cellH": 216 },
  "videoWidth": 854,
  "videoHeight": 480,
  "inputSizeMb": 42.3
}

If the video exceeds 10 minutes and uniform mode was used without --start/--end, a hint field is included:

{
  "hint": "Video is 30 minutes long. This is a uniform overview. For better scene coverage re-run with --mode highlight, or use --start/--end to zoom into a specific section."
}

Recommended workflow for long videos:

  1. First run with --mode highlight — shows key scene changes across the whole video
  2. If the user wants detail on a specific section, re-run with --start N --end N

On error, writes ERROR: <message> + Hint: <diagnosis> to stderr and exits 1.

Injection procedure

Step 1 — Run the script (bash tool):

node {baseDir}/scripts/inject.mjs "/path/to/video.mp4"

Step 2 — Parse JSON: Extract gridPath and description.

Step 3 — Inject image (read tool):

read <gridPath>

The read tool injects the jpg as a native multimodal image block into context. After viewing the grid, use the description XML timestamps to reference frames:

"Look at the grid image above. Use the timestamps in the description XML to analyze the video. The number in the top-left of each cell is the frame index."

On error:

  • Translate the Hint: message into natural language for the user. Do not paste raw error output.
  • If read <gridPath> fails — /tmp/ files are ephemeral. Re-run the script and read immediately.

Options

OptionDefaultDescription
--mode uniformEvenly spaced frames
--mode highlightScene-change biased sampling
--start N0Segment start (seconds)
--end Nend of videoSegment end (seconds)

Diagnostics

ErrorCauseAction
Input file not foundFile missing or dropped by channel media size limitAsk the user to share the file path directly as text
corrupt, incomplete, or unsupported formatDamaged file, interrupted transfer, or unsupported codecTry a different file, or use --start/--end to skip problematic sections
moov atom not foundIncomplete mp4 (streaming not finished)Retry with a complete file
ffmpeg not foundffmpeg not installedCheck ffmpeg installation

Notes

  • Frame count and cell size are determined automatically from video duration and aspect ratio
  • Grid is ~1500×1500px, cell long side 384–512px
  • Timestamps are in the description XML only, not overlaid on the image
  • Portrait and landscape videos both supported
  • Telegram users: if a video file is not attached to the message, check channels.telegram.mediaMaxMb in the OpenClaw config — the file may have been dropped at the channel level before reaching the agent

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

71.12%
按下载量换算772

安全审计

VirusTotal

通过

ClawScan

通过

Static analysis

通过

权限和风险

只读

该 Skill 主要提供规则、说明或参考内容,本身偏只读;真正读写文件、联网或执行命令仍取决于宿主 Agent 的任务。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills