Token导航 LogoToken导航TokenDH.com
图像处理操作浏览器clawhub未标认证来源可访问clear审计通过

image-to-video-runcomfy图像到视频运行舒适

Agent Skill

用于辅助图像生成、图片编辑、视觉素材处理或图像模型工作流。它适合让 Agent 根据文本生成图片、处理背景、整理视觉提示词或调用相关图像工具。使用时需要确认输入图片、版权来源、输出格式和模型限制;涉及人物、品牌、商品或公开展示素材时,应额外核对授权、真实性和内容合规边界。

总安装

294

周安装

12

GitHub Stars

公开资料未说明

下载量

95
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:image-to-video-runcomfy(图像到视频运行舒适)
来源仓库:https://github.com/kalvinrv/image-to-video-runcomfy
安装命令:
openclaw skills install image-to-video-runcomfy
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install image-to-video-runcomfy

简介

智能路由RunComfy目录中的i2v模型,适配不同图像类型。

  • 精选快乐路径模型,提升动画生成成功率与效果。
  • 上传任意静态图并描述意图,系统自动选择最优模型。
  • 依赖RunComfy服务可用性,可能存在延迟或排队情况。
  • image-to-video-runcomfy 属于图像处理类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

name
image-to-video-runcomfy
displayName
🫧 Image-to-Video — Pro Pack on RunComfy
description
>
emoji
🫧
homepage
https://www.runcomfy.com
license
MIT
clawdis
requires
bins
env
config

🫧 Image-to-Video — Pro Pack on RunComfy

runcomfy.com · docs · Image-to-video models

Image-to-video generation on RunComfy. This skill is the canonical image-to-video entry point for the RunComfy Model API: give it a still image and a motion description, and it returns a short video clip. Image-to-video on RunComfy means turning any image — portrait, product photo, environment, illustration — into a video, with the motion driven by your prompt.

What "image-to-video" means here

Image-to-video (often abbreviated i2v or image2video) is the task of generating a short video starting from a single still image. The image fixes the look — face, wardrobe, product, scene geometry — and the prompt drives the motion. Image-to-video is distinct from text-to-video (no input image) and from video-to-video (which transforms an existing clip).

Image-to-video on RunComfy supports three patterns:

  • General image-to-video: animate any still — portrait drift, product reveal, environment motion, illustration coming alive. The default image-to-video pipeline.
  • Lip-sync image-to-video: a custom voiceover drives mouth movement on a generated talking-head image-to-video clip. Input: image + audio. Output: lip-synced image-to-video.
  • Multi-modal image-to-video: combine subject image + reference scene video + reference voice audio into one image-to-video output.

This skill picks the right image-to-video endpoint for the user's intent and calls runcomfy run <model>/image-to-video with the matching schema.

When to use image-to-video on RunComfy

Pick image-to-video on RunComfy whenever:

  • You have a still image and want it to move — image-to-video is the right task.
  • You want identity-stable image-to-video — the face / product / brand from your input image must survive into the output video.
  • You want fast iteration on image-to-video — RunComfy hosts the GPU; you don't deploy or rent.
  • You're building image-to-video at scale — multi-language image-to-video dubs, multi-shot image-to-video sequences, batch image-to-video jobs.

If the user said "image to video", "i2v", "animate this image", "image2video", "make a video from this", or showed an image and asked for video — route here.

Image-to-video routes

User intentImage-to-video modelWhy
Default image-to-video — portraits, products, environmentshappyhorse-1-0/image-to-video#1 on Arena (Elo 1392 i2v); strong identity preservation; native synchronized audio in image-to-video output
Image-to-video with custom voiceover lip-syncwan-ai/wan-2-7/text-to-video + audio_urlDrives lip-sync on the image-to-video frame from your audio file
Multi-modal image-to-video (image + ref video + ref audio)bytedance/seedance-v2/proMulti-input image-to-video with up to 9 image refs and 3 audio refs

The agent reads this table, classifies the user's image-to-video intent, and picks the matching endpoint.

Prerequisites

  1. RunComfy CLInpm i -g @runcomfy/cli
  2. RunComfy accountruncomfy login opens a browser device-code flow.
  3. CI / containers — set RUNCOMFY_TOKEN=<token>.
  4. A source image URL — JPEG/PNG/WebP, min 300px, ≤10MB; aspect 1:2.5 to 2.5:1 for the default image-to-video model.

Default image-to-video — HappyHorse 1.0 i2v

The default image-to-video endpoint. Use for any general image-to-video task: portrait drift, product reveal, environment motion, character animation. Image-to-video output includes synchronized audio in the same generation pass.

Schema

FieldTypeRequiredDefaultNotes
image_urlstringyesThe source still for image-to-video. JPEG/PNG/WebP, min 300px, aspect 1:2.5–2.5:1, ≤10MB.
promptstringyesMotion / camera / lighting description for the image-to-video output. ≤5000 chars.
resolutionenumno1080P720P or 1080P.
durationintno53–15 seconds per image-to-video clip.
seedintno0Reuse for image-to-video variant comparisons.
watermarkboolnotrueProvider watermark on image-to-video output.

Output aspect of the image-to-video clip equals input image aspect.

Invoke

runcomfy run happyhorse/happyhorse-1-0/image-to-video \
  --input '{
    "image_url": "https://.../portrait.jpg",
    "prompt": "Gentle camera drift around the subject'\''s face, subtle breathing motion, identity-stable features, soft natural light."
  }' \
  --output-dir <absolute/path>

Lip-sync image-to-video — custom voiceover

When the image-to-video output needs to lip-sync to a custom audio track, use Wan 2.7 with audio_url. The image-to-video clip is generated around your voiceover so mouth movement matches.

FieldTypeRequiredNotes
promptstringyesDescribe the talking-head shot for the image-to-video output.
audio_urlstringyesWAV/MP3, 3–30s, ≤15MB. Drives lip-sync on the image-to-video frame.
aspect_ratioenumno16:9, 9:16, 1:1, 4:3, 3:4.
resolutionenumno720p or 1080p.
durationenumno2–15 seconds. Match audio length for clean image-to-video lip-sync.
runcomfy run wan-ai/wan-2-7/text-to-video \
  --input '{
    "prompt": "Medium close-up, soft key light, locked tripod, shallow DOF.",
    "audio_url": "https://.../voiceover-en.mp3",
    "duration": 12,
    "aspect_ratio": "9:16"
  }' \
  --output-dir <absolute/path>

For multi-language image-to-video dubs: same prompt, swap audio_url per call, lock seed for visual consistency across all image-to-video outputs.

Multi-modal image-to-video — image + ref video + ref audio

When the image-to-video output should fuse a subject image with a scene reference and voice reference, use Seedance 2.0 Pro. Multi-modal image-to-video accepts up to 9 image refs.

FieldTypeRequiredNotes
promptstringyesDescription for the image-to-video output. EN ≤1000 words.
image_urlarrayyes0–9 source images for image-to-video. First is the primary subject.
video_urlarrayno0–3 reference clips (2–15s each) for image-to-video scene cues.
audio_urlarrayno0–3 reference audio (2–15s, <15MB each) for image-to-video voice cues.
durationintno4–15 seconds.
resolutionenumno480p or 720p.
runcomfy run bytedance/seedance-v2/pro \
  --input '{
    "prompt": "Subject from image 1 walks through the scene from video 1, voice from audio 1.",
    "image_url": ["https://.../subject.jpg"],
    "video_url": ["https://.../scene.mp4"],
    "audio_url": ["https://.../voice.mp3"],
    "duration": 8
  }' \
  --output-dir <absolute/path>

Prompting image-to-video — what works

Image-to-video prompts behave differently from text-to-video prompts. The image already fixes the look — your prompt should drive motion, not redescribe the image.

  • Lead with motion verbs. "drift", "dolly in", "orbit", "tilt up", "blink", "breathe" — front-load what's MOVING in the image-to-video output.
  • Don't restate the image. The image-to-video model sees the input. Spend tokens on what changes, not what already exists.
  • Preservation goals explicit. "identity-stable features", "packaging unchanged", "background geometry stable" — tell the image-to-video model what NOT to change.
  • One beat per image-to-video clip. Single primary motion (orbit OR dolly OR tilt OR character action). Compound motion drifts.
  • Lighting evolution. "rim light intensifying", "shadows shortening as camera rises" — image-to-video output reads lighting cues well.

Image-to-video FAQ

What's the max duration of an image-to-video clip? 15 seconds across all image-to-video routes here. For longer image-to-video sequences, generate multiple clips and stitch.

What image formats does image-to-video accept? JPEG, PNG, WebP. Min 300px, ≤10MB, aspect 1:2.5 to 2.5:1.

Does image-to-video preserve face identity? Yes — the default image-to-video model has strong identity preservation. For best identity hold, the face should fill at least 5% of the frame in the input image.

Can image-to-video include audio? Yes. The default image-to-video model generates synchronized audio in the same pass. The lip-sync image-to-video route accepts your custom audio. The multi-modal image-to-video route accepts reference audio.

Image-to-video vs text-to-video on RunComfy? Image-to-video starts from your image (look fixed). Text-to-video starts from your prompt only (look generated). Use image-to-video when you have an exact reference; use text-to-video for novel content.

Image-to-video output resolution? 720p or 1080p depending on the route.

Limitations

  • Image-to-video clip length is 15s per call. Longer image-to-video output requires stitching multiple calls.
  • Image-to-video output aspect = input image aspect on the default route. For independent reframing, crop the input first.
  • Image-to-video doesn't blend across routes in one call. If you need multi-modal image-to-video + custom voiceover lip-sync in one clip, that's two image-to-video calls plus a stitch.

Exit codes

codemeaning
0image-to-video succeeded
64bad CLI args
65bad input JSON for image-to-video / schema mismatch
69upstream 5xx
75retryable: timeout / 429
77not signed in or token rejected

Full reference: docs.runcomfy.com/cli/troubleshooting.

How it works

The skill picks one of three image-to-video endpoints based on user intent (general image-to-video, lip-sync image-to-video, or multi-modal image-to-video) and invokes runcomfy run <endpoint> with the matching JSON body. The CLI POSTs to the RunComfy Model API, polls the image-to-video request status every 2 seconds, and downloads the resulting image-to-video file from the *.runcomfy.net / *.runcomfy.com URL into --output-dir. Ctrl-C cancels the in-flight image-to-video request.

Security & Privacy

  • Token storage: runcomfy login writes the API token to ~/.config/runcomfy/token.json with mode 0600. Set RUNCOMFY_TOKEN env var to bypass the file in CI.
  • Input boundary: the image-to-video prompt is passed as JSON via --input. The CLI does NOT shell-expand. No shell-injection surface.
  • Third-party content: image / video / audio URLs are fetched by the RunComfy server. Treat external URLs as untrusted; image-based prompt injection is a known risk for any image-to-video model.
  • Outbound endpoints: only model-api.runcomfy.net and *.runcomfy.net / *.runcomfy.com. No telemetry.
  • Generated-file size cap: the CLI aborts any image-to-video download > 2 GiB.

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

77.01%
按下载量换算73

安全审计

VirusTotal

通过

ClawScan

通过

Static analysis

通过

权限和风险

操作浏览器

该 Skill 可能涉及浏览器控制能力,使用时可能读取或操作网页内容,需要在受控环境中确认权限边界。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills