Token导航 LogoToken导航TokenDH.com
开发敏感数据clawhub未标认证来源可访问clear审计提醒

runbook-generator运行手册生成器

Agent Skill

runbook-generator 用于辅助部署、云资源、容器和基础设施运维,适合在 OpenClaw 中需要检查配置、整理部署步骤或排查环境问题时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

10,827

周安装

438

GitHub Stars

公开资料未说明

下载量

3,399
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:runbook-generator(运行手册生成器)
来源仓库:https://github.com/alirezarezvani/runbook-generator
安装命令:
openclaw skills install runbook-generator
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install runbook-generator

简介

runbook-generator 生成运维操作手册,用于部署与故障排查。

  • 适用于 OpenClaw 中云资源、容器与基础设施管理任务。
  • 通过 clawhub 安装,命令为 openclaw skills install runbook-generator。
  • 使用前应确认权限范围,避免生成包含敏感操作的指令。
  • 建议结合 README 了解模板结构与变量替换规则。

SKILL.md

name
runbook-generator
description
Runbook Generator

Runbook Generator

Tier: POWERFUL Category: Engineering Domain: DevOps / Site Reliability Engineering


Overview

Analyze a codebase and generate production-grade operational runbooks. Detects your stack (CI/CD, database, hosting, containers), then produces step-by-step runbooks with copy-paste commands, verification checks, rollback procedures, escalation paths, and time estimates. Keeps runbooks fresh with staleness detection linked to config file modification dates.


Core Capabilities

  • Stack detection — auto-identify CI/CD, database, hosting, orchestration from repo files
  • Runbook types — deployment, incident response, database maintenance, scaling, monitoring setup
  • Format discipline — numbered steps, copy-paste commands, ✅ verification checks, time estimates
  • Escalation paths — L1 → L2 → L3 with contact info and decision criteria
  • Rollback procedures — every deployment step has a corresponding undo
  • Staleness detection — runbook sections reference config files; flag when source changes
  • Testing methodology — dry-run framework for staging validation, quarterly review cadence

When to Use

Use when:

  • A codebase has no runbooks and you need to bootstrap them fast
  • Existing runbooks are outdated or incomplete (point at the repo, regenerate)
  • Onboarding a new engineer who needs clear operational procedures
  • Preparing for an incident response drill or audit
  • Setting up monitoring and on-call rotation from scratch

Skip when:

  • The system is too early-stage to have stable operational patterns
  • Runbooks already exist and only need minor updates (edit directly)

Stack Detection

When given a repo, scan for these signals before writing a single runbook line:

# CI/CD
ls .github/workflows/     → GitHub Actions
ls .gitlab-ci.yml         → GitLab CI
ls Jenkinsfile            → Jenkins
ls .circleci/             → CircleCI
ls bitbucket-pipelines.yml → Bitbucket Pipelines

# Database
grep -r "postgresql\|postgres\|pg" package.json pyproject.toml → PostgreSQL
grep -r "mysql\|mariadb"           package.json               → MySQL
grep -r "mongodb\|mongoose"        package.json               → MongoDB
grep -r "redis"                    package.json               → Redis
ls prisma/schema.prisma            → Prisma ORM (check provider field)
ls drizzle.config.*                → Drizzle ORM

# Hosting
ls vercel.json                     → Vercel
ls railway.toml                    → Railway
ls fly.toml                        → Fly.io
ls .ebextensions/                  → AWS Elastic Beanstalk
ls terraform/  ls *.tf             → Custom AWS/GCP/Azure (check provider)
ls kubernetes/ ls k8s/             → Kubernetes
ls docker-compose.yml              → Docker Compose

# Framework
ls next.config.*                   → Next.js
ls nuxt.config.*                   → Nuxt
ls svelte.config.*                 → SvelteKit
cat package.json | jq '.scripts'   → Check build/start commands

Map detected stack → runbook templates. A Next.js + PostgreSQL + Vercel + GitHub Actions repo needs:

  • Deployment runbook (Vercel + GitHub Actions)
  • Database runbook (PostgreSQL backup, migration, vacuum)
  • Incident response (with Vercel logs + pg query debugging)
  • Monitoring setup (Vercel Analytics, pg_stat, alerting)

Runbook Types

1. Deployment Runbook

# Deployment Runbook — [App Name]
**Stack:** Next.js 14 + PostgreSQL 15 + Vercel  
**Last verified:** 2025-03-01  
**Source configs:** vercel.json (modified: git log -1 --format=%ci -- vercel.json)  
**Owner:** Platform Team  
**Est. total time:** 15–25 min  

---

## Pre-deployment Checklist
- [ ] All PRs merged to main
- [ ] CI passing on main (GitHub Actions green)
- [ ] Database migrations tested in staging
- [ ] Rollback plan confirmed

## Steps

### Step 1 — Run CI checks locally (3 min)

pnpm test pnpm lint pnpm build

✅ Expected: All pass with 0 errors. Build output in `.next/`

### Step 2 — Apply database migrations (5 min)

Staging first

DATABASE_URL=$STAGING_DATABASE_URL npx prisma migrate deploy

✅ Expected: `All migrations have been successfully applied.`

Verify migration applied

psql $STAGING_DATABASE_URL -c "\d" | grep -i migration

✅ Expected: Migration table shows new entry with today's date

### Step 3 — Deploy to production (5 min)

git push origin main

OR trigger manually:

vercel --prod

✅ Expected: Vercel dashboard shows deployment in progress. URL format:
`https://app-name-<hash>-team.vercel.app`

### Step 4 — Smoke test production (5 min)

Health check

curl -sf https://your-app.vercel.app/api/health | jq .

Critical path

curl -sf https://your-app.vercel.app/api/users/me \ -H "Authorization: Bearer $TEST_TOKEN" | jq '.id'

✅ Expected: health returns `{"status":"ok","db":"connected"}`. Users API returns valid ID.

### Step 5 — Monitor for 10 min
- Check Vercel Functions log for errors: `vercel logs --since=10m`
- Check error rate in Vercel Analytics: < 1% 5xx
- Check DB connection pool: `SELECT count(*) FROM pg_stat_activity;` (< 80% of max_connections)

---

## Rollback

If smoke tests fail or error rate spikes:

Instant rollback via Vercel (preferred — < 30 sec)

vercel rollback [previous-deployment-url]

Database rollback (only if migration was applied)

DATABASE_URL=$PROD_DATABASE_URL npx prisma migrate reset --skip-seed

WARNING: This resets to previous migration. Confirm data impact first.


✅ Expected after rollback: Previous deployment URL becomes active. Verify with smoke test.

---

## Escalation
- **L1 (on-call engineer):** Check Vercel logs, run smoke tests, attempt rollback
- **L2 (platform lead):** DB issues, data loss risk, rollback failed — Slack: @platform-lead
- **L3 (CTO):** Production down > 30 min, data breach — PagerDuty: #critical-incidents

2. Incident Response Runbook

# Incident Response Runbook
**Severity levels:** P1 (down), P2 (degraded), P3 (minor)  
**Est. total time:** P1: 30–60 min, P2: 1–4 hours  

## Phase 1 — Triage (5 min)

### Confirm the incident

Is the app responding?

curl -sw "%{http_code}" https://your-app.vercel.app/api/health -o /dev/null

Check Vercel function errors (last 15 min)

vercel logs --since=15m | grep -i "error\|exception\|5[0-9][0-9]"

✅ 200 = app up. 5xx or timeout = incident confirmed.

Declare severity:
- Site completely down → P1 — page L2/L3 immediately
- Partial degradation / slow responses → P2 — notify team channel
- Single feature broken → P3 — create ticket, fix in business hours

---

## Phase 2 — Diagnose (10–15 min)

Recent deployments — did something just ship?

vercel ls --limit=5

Database health

psql $DATABASE_URL -c "SELECT pid, state, wait_event, query FROM pg_stat_activity WHERE state != 'idle' LIMIT 20;"

Long-running queries (> 30 sec)

psql $DATABASE_URL -c "SELECT pid, now() - pg_stat_activity.query_start AS duration, query FROM pg_stat_activity WHERE state = 'active' AND now() - pg_stat_activity.query_start > interval '30 seconds';"

Connection pool saturation

psql $DATABASE_URL -c "SELECT count(*), max_conn FROM pg_stat_activity, (SELECT setting::int AS max_conn FROM pg_settings WHERE name='max_connections') t GROUP BY max_conn;"


Diagnostic decision tree:
- Recent deploy + new errors → rollback (see Deployment Runbook)
- DB query timeout / pool saturation → kill long queries, scale connections
- External dependency failing → check status pages, add circuit breaker
- Memory/CPU spike → check Vercel function logs for infinite loops

---

## Phase 3 — Mitigate (variable)

Kill a runaway DB query

psql $DATABASE_URL -c "SELECT pg_terminate_backend(<pid>);"

Scale DB connections (Supabase/Neon — adjust pool size)

Vercel → Settings → Environment Variables → update DATABASE_POOL_MAX

Enable maintenance mode (if you have a feature flag)

vercel env add MAINTENANCE_MODE true production vercel --prod # redeploy with flag


---

## Phase 4 — Resolve & Postmortem

After incident is resolved, within 24 hours:

1. Write incident timeline (what happened, when, who noticed, what fixed it)
2. Identify root cause (5-Whys)
3. Define action items with owners and due dates
4. Update this runbook if a step was missing or wrong
5. Add monitoring/alert that would have caught this earlier

**Postmortem template:** `docs/postmortems/YYYY-MM-DD-incident-title.md`

---

## Escalation Path

| Level | Who | When | Contact |
|-------|-----|------|---------|
| L1 | On-call engineer | Always first | PagerDuty rotation |
| L2 | Platform lead | DB issues, rollback needed | Slack @platform-lead |
| L3 | CTO/VP Eng | P1 > 30 min, data loss | Phone + PagerDuty |

3. Database Maintenance Runbook

# Database Maintenance Runbook — PostgreSQL
**Schedule:** Weekly vacuum (automated), monthly manual review  

## Backup

Full backup

pg_dump $DATABASE_URL \ --format=custom \ --compress=9 \ --file="backup-$(date +%Y%m%d-%H%M%S).dump"

✅ Expected: File created, size > 0. `pg_restore --list backup.dump | head -20` shows tables.

Verify backup is restorable (test monthly):

pg_restore --dbname=$STAGING_DATABASE_URL backup.dump psql $STAGING_DATABASE_URL -c "SELECT count(*) FROM users;"

✅ Expected: Row count matches production.

## Migration

Always test in staging first

DATABASE_URL=$STAGING_DATABASE_URL npx prisma migrate deploy

Verify, then:

DATABASE_URL=$PROD_DATABASE_URL npx prisma migrate deploy

✅ Expected: `All migrations have been successfully applied.`

⚠️ For large table migrations (> 1M rows), use `pg_repack` or add column with DEFAULT separately to avoid table locks.

## Vacuum & Reindex

Check bloat before deciding

psql $DATABASE_URL -c " SELECT schemaname, tablename, pg_size_pretty(pg_total_relation_size(schemaname||'.'||tablename)) AS total_size, n_dead_tup, n_live_tup, ROUND(n_dead_tup::numeric / NULLIF(n_live_tup + n_dead_tup, 0) * 100, 1) AS dead_ratio FROM pg_stat_user_tables ORDER BY n_dead_tup DESC LIMIT 10;"

Vacuum high-bloat tables (non-blocking)

psql $DATABASE_URL -c "VACUUM ANALYZE users;" psql $DATABASE_URL -c "VACUUM ANALYZE events;"

Reindex (use CONCURRENTLY to avoid locks)

psql $DATABASE_URL -c "REINDEX INDEX CONCURRENTLY users_email_idx;"

✅ Expected: dead_ratio drops below 5% after vacuum.

Staleness Detection

Add a staleness header to every runbook:

## Staleness Check
This runbook references the following config files. If they've changed since the
"Last verified" date, review the affected steps.

| Config File | Last Modified | Affects Steps |
|-------------|--------------|---------------|
| vercel.json | `git log -1 --format=%ci -- vercel.json` | Step 3, Rollback |
| prisma/schema.prisma | `git log -1 --format=%ci -- prisma/schema.prisma` | Step 2, DB Maintenance |
| .github/workflows/deploy.yml | `git log -1 --format=%ci -- .github/workflows/deploy.yml` | Step 1, Step 3 |
| docker-compose.yml | `git log -1 --format=%ci -- docker-compose.yml` | All scaling steps |

Automation: Add a CI job that runs weekly and comments on the runbook doc if any referenced file was modified more recently than the runbook's "Last verified" date.


Runbook Testing Methodology

Dry-Run in Staging

Before trusting a runbook in production, validate every step in staging:

# 1. Create a staging environment mirror
vercel env pull .env.staging
source .env.staging

# 2. Run each step with staging credentials
# Replace all $DATABASE_URL with $STAGING_DATABASE_URL
# Replace all production URLs with staging URLs

# 3. Verify expected outputs match
# Document any discrepancies and update the runbook

# 4. Time each step — update estimates in the runbook
time npx prisma migrate deploy

Quarterly Review Cadence

Schedule a 1-hour review every quarter:

  1. Run each command in staging — does it still work?
  2. Check config drift — compare "Last Modified" dates vs "Last verified"
  3. Test rollback procedures — actually roll back in staging
  4. Update contact info — L1/L2/L3 may have changed
  5. Add new failure modes discovered in the past quarter
  6. Update "Last verified" date at top of runbook

Common Pitfalls

PitfallFix
Commands that require manual copy of dynamic valuesUse env vars — $DATABASE_URL not postgres://user:pass@host/db
No expected output specifiedAdd ✅ with exact expected string after every verification step
Rollback steps missingEvery destructive step needs a corresponding undo
Runbooks that never get testedSchedule quarterly staging dry-runs in team calendar
L3 escalation contact is the former CTOReview contact info every quarter
Migration runbook doesn't mention table locksCall out lock risk for large table operations explicitly

Best Practices

  1. Every command must be copy-pasteable — no placeholder text, use env vars
  2. ✅ after every step — explicit expected output, not "it should work"
  3. Time estimates are mandatory — engineers need to know if they have time to fix before SLA breach
  4. Rollback before you deploy — plan the undo before executing
  5. Runbooks live in the repodocs/runbooks/, versioned with the code they describe
  6. Postmortem → runbook update — every incident should improve a runbook
  7. Link, don't duplicate — reference the canonical config file, don't copy its contents into the runbook
  8. Test runbooks like you test code — untested runbooks are worse than no runbooks (false confidence)

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

95.55%
按下载量换算3,248

安全审计

VirusTotal

通过

ClawScan

可疑

Static analysis

通过

权限和风险

敏感数据

该 Skill 可能接触密钥、Token、环境变量或敏感配置,应进入高风险复核队列,默认不自动发布。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills