Token导航 LogoToken导航TokenDH.com
研究检索需要联网clawhub未标认证来源可访问clear审计通过

subtitle-video-generator字幕视频生成器

Agent Skill

用于辅助视频生成、动画合成、脚本化剪辑或 Remotion 等视频项目开发。它适合让 Agent 组织镜头、生成素材说明、维护合成代码或排查渲染问题。使用时需要确认分辨率、时长、素材路径和导出格式;涉及外部素材、人物肖像或商业发布时,应先核对版权授权和内容审核要求。

总安装

10,350

周安装

427

GitHub Stars

公开资料未说明

下载量

3,382
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:subtitle-video-generator(字幕视频生成器)
来源仓库:https://github.com/peand-rover/subtitle-video-generator
安装命令:
openclaw skills install subtitle-video-generator
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install subtitle-video-generator

简介

自动生成多语言字幕并嵌入视频,支持样式定制。

  • 适用于语音转录、翻译与字幕排版一体化场景。subtitle-video-generator 属于研究检索类 Skill,可作为该场景下的辅助能力补充。
  • 可处理 50 种以上语言,精准同步时间轴。
  • 需确保音频清晰,避免识别错误影响字幕质量。
  • 安装命令:openclaw skills install subtitle-video-generator

SKILL.md

name
subtitle-video-generator
version
1.2.2
displayName
Subtitle Video Generator — Generate and Style Video Subtitles in Any Language with AI
description
>
metadata
{"openclaw": {"emoji": "📝", "requires": {"env": ["NEMO_TOKEN"], "configPaths": ["~/.config/nemovideo/"]}, "primaryEnv": "NEMO_TOKEN"}}
homepage
https://nemovideo.com
repository
https://github.com/nemovideo/nemovideo_skills
apiDomain
https://mega-api-prod.nemovideo.ai

Subtitle Video Generator — Every Language. Every Style. Every Platform. One Upload.

Subtitles have become the universal interface between video content and global audiences. They serve four distinct functions simultaneously: accessibility (making content usable for deaf and hard-of-hearing viewers), engagement (holding attention for the 85% watching muted on social media), reach (translating content to audiences in 50+ languages), and discoverability (providing text that search algorithms can index). Each function alone justifies subtitling every video. Together, they make subtitling the single highest-ROI post-production addition to any video content. The quality gap between auto-generated platform subtitles and professional subtitling is the space NemoVideo fills. Platform auto-captions deliver 80-85% accuracy — one error every 15-20 words, visible to viewers and damaging to credibility. Professional human subtitling achieves 99%+ accuracy at $3-8 per video minute with 24-48 hour turnaround. NemoVideo delivers 98%+ accuracy with word-level timing, full style customization, multi-language translation, speaker differentiation, and instant turnaround. The quality that previously required professional subtitling services, delivered at the speed and scale that modern content production demands.

Use Cases

  1. Social Media Subtitles — Engagement-Optimized Styling (15-90s) — Short-form content for TikTok, Instagram Reels, and YouTube Shorts needs the animated subtitle style that maximizes watch time. NemoVideo: transcribes with word-level timing accuracy, applies the platform-native animated style (large bold text, word-by-word highlight animation in the creator's brand color, high-contrast outline for readability), positions within the specific platform's safe zone (TikTok: above bottom 15%; Reels: above bottom 20%, below top 10%; Shorts: above bottom 10%), and exports with subtitles rendered directly into the video (essential for platforms where subtitle upload is limited or unreliable). The subtitle style proven to increase short-form completion rate by 15-25%.
  1. Corporate Multi-Language — Global Communications (any length) — A corporation produces video content that needs to reach employees and customers across 15+ countries. NemoVideo: transcribes the source language, translates to all target languages using context-aware AI (understanding corporate terminology, product names, and industry jargon), adjusts subtitle timing per language (expanding for languages that require more words, contracting for languages that use fewer), handles bidirectional text for Arabic and Hebrew (proper RTL rendering with correct line breaking), applies consistent corporate subtitle styling across all languages (brand fonts, colors, positioning), and exports subtitle files compatible with the company's video hosting infrastructure (SRT for most platforms, VTT for web, TTML for broadcast). One video, global reach, consistent brand quality.
  1. Educational Subtitles — Learning-Optimized Display (any length) — Educational content requires subtitles optimized for comprehension rather than entertainment: slower reading speed for complex material, technical term highlighting, and clear speaker identification for multi-person discussions. NemoVideo: adjusts reading speed based on content complexity (14 characters/second for dense technical content vs. 18 cps for conversational segments), optionally highlights technical vocabulary on first appearance (bold or different color for terms that may be unfamiliar), identifies speakers with persistent color differentiation (essential for panel discussions and multi-instructor courses), maintains sentence-aware line breaks (never splitting a phrase across lines in a way that disrupts comprehension), and generates WCAG 2.1 AA-compliant subtitles for institutional accessibility requirements. Subtitles that serve learning, not just consumption.
  1. Film and Documentary — Broadcast Standard Subtitling (any length) — Independent filmmakers and documentary producers need subtitles meeting broadcast and festival technical specifications. NemoVideo: generates subtitles conforming to broadcast standards (2 lines maximum, 42 characters per line maximum, 1-second minimum display time, 15-17 characters per second reading speed), applies professional positioning (center-bottom, with vertical offset when on-screen text or graphics would be obscured), handles complex audio scenarios (overlapping dialogue, music with lyrics, background conversations, sound effects that need description for accessibility), creates both open subtitles (burned into video for festival screenings) and closed subtitles (as separate files for broadcast distribution), and exports in industry-standard formats (SRT, STL, EBU-TT, TTML, DFXP). Festival-submission-ready and broadcast-compliant subtitles.
  1. Batch Library Subtitling — Retrofit an Entire Catalog (multiple videos) — A content library of 200+ videos has grown without consistent subtitling. Some have auto-generated captions, some have nothing, and none have translations. NemoVideo: batch-processes the entire library with consistent subtitle styling (same font, color, position, animation across all videos), auto-detects the spoken language per video (handling a multilingual library), transcribes or re-transcribes each video at 98%+ accuracy (replacing inaccurate existing auto-captions), generates translations for specified target languages across the entire library, and produces both embedded-subtitle video files and standalone subtitle files for each video. A subtitle-inconsistent library becomes a professionally subtitled catalog.

How It Works

Step 1 — Upload Video

Any video with speech in any language. Single video or batch upload. NemoVideo auto-detects language and speaker count.

Step 2 — Configure Subtitle Output

Style (animated, broadcast, minimal, custom), languages (source + translation targets), positioning, speaker differentiation, and export formats.

Step 3 — Generate

curl -X POST https://mega-api-prod.nemovideo.ai/api/v1/generate \
  -H "Authorization: Bearer $NEMO_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "skill": "subtitle-video-generator",
    "prompt": "Generate professional subtitles for a 15-minute interview with two speakers. English source — translate to Spanish, French, and Mandarin. Style: clean broadcast (white text, semi-transparent dark background bar, 2 lines max). Speaker differentiation: interviewer in white, guest in light yellow. Word-level timing. Position: bottom-center, offset upward on frames where lower-third graphics appear. Export: embedded MP4 for each language at 16:9 + standalone SRT files for all 4 languages + one 9:16 English version with TikTok animated style for social clips.",
    "source_language": "en",
    "translations": ["es", "fr", "zh"],
    "style": {
      "preset": "broadcast-clean",
      "background": "semi-transparent-dark",
      "max_lines": 2,
      "timing": "word-level"
    },
    "speakers": {
      "differentiate": true,
      "interviewer": "#FFFFFF",
      "guest": "#FFFACD"
    },
    "position": {"base": "bottom-center", "avoid_lower_thirds": true},
    "exports": {
      "embedded_16x9": ["en", "es", "fr", "zh"],
      "srt_files": ["en", "es", "fr", "zh"],
      "social_9x16": {"language": "en", "style": "tiktok-animated"}
    }
  }'

Step 4 — Review Accuracy and Timing

Play each language version. Verify: transcription accuracy (especially names, technical terms, numbers), timing synchronization, speaker identification correctness, translation naturalness, and line break logic. Edit and re-render any corrections.

Parameters

ParameterTypeRequiredDescription
promptstringSubtitle generation requirements
source_languagestringSource audio language (auto-detect if omitted)
translationsarrayTarget languages ["es", "fr", "zh", ...]
styleobject{preset, font, color, background, max_lines, animation, timing}
speakersobject{differentiate, colors_per_role}
positionobject{base, offset, avoid_graphics}
reading_speedstring"educational-slow", "standard", "fast"
broadcast_compliancebooleanApply broadcast subtitle standards
accessibilitystring"wcag-aa", "wcag-aaa"
exportsobject{embedded, srt_files, vtt_files, social}
batchbooleanProcess multiple videos

Output Example

{
  "job_id": "subgen-20260329-001",
  "status": "completed",
  "source_language": "en",
  "confidence": 0.986,
  "speakers": 2,
  "word_count": 3240,
  "languages": ["en", "es", "fr", "zh"],
  "outputs": {
    "embedded": {
      "en": {"file": "interview-sub-en-16x9.mp4"},
      "es": {"file": "interview-sub-es-16x9.mp4"},
      "fr": {"file": "interview-sub-fr-16x9.mp4"},
      "zh": {"file": "interview-sub-zh-16x9.mp4"}
    },
    "srt_files": ["interview-en.srt", "interview-es.srt", "interview-fr.srt", "interview-zh.srt"],
    "social": {"file": "interview-tiktok-en-9x16.mp4", "style": "tiktok-animated"}
  }
}

Tips

  1. 98% accuracy is the professional credibility threshold — At 85% (platform auto), viewers notice errors constantly and question content quality. At 98%, errors are rare enough that the subtitle feels professionally produced. The accuracy difference is the difference between undermining and reinforcing your credibility.
  2. Word-level timing creates synchronized reading that holds attention — Sentence-level display (full sentence appears at once) disconnects reading from listening. Word-level timing (each word appears as spoken) synchronizes the two channels, creating engaged viewing that platform auto-captions cannot achieve.
  3. Speaker color coding is faster than speaker labels — "John:" before each subtitle line wastes characters and reading time. White for John, yellow for Sarah communicates the same information through pre-attentive color processing — faster than reading a name label every time the speaker changes.
  4. Translation timing must expand and contract per language — German averages 30% more characters than English. Japanese averages fewer. If subtitle display time does not adjust, German viewers cannot finish reading and Japanese viewers stare at completed subtitles. Per-language timing is essential for comfortable reading speed in every language.
  5. Batch subtitling eliminates the growing liability of uncaptioned content — Every uncaptioned video is a missed accessibility obligation, a missed engagement opportunity, and a missed global reach opportunity. Batch processing converts an entire backlog in one operation, establishing the baseline for captioning all future content.

Output Formats

FormatTypeUse Case
MP4 (embedded)VideoSocial platforms, website, LMS
SRTSubtitle fileYouTube, Vimeo, most platforms
VTTSubtitle fileWeb players, HTML5 video
TTML / DFXPSubtitle fileBroadcast, streaming services
STLSubtitle fileEuropean broadcast
EBU-TTSubtitle fileEBU broadcast standard

Related Skills

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

93.17%
按下载量换算3,151

安全审计

VirusTotal

通过

ClawScan

通过

Static analysis

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills