Token导航 LogoToken导航TokenDH.com
研究检索external-servicegithub未标认证来源可访问许可证需确认审计提醒

running-chaos-tests运行混沌测试

Agent Skill

用于辅助测试设计、自动化测试、用例整理和回归验证。它适合让 Agent 编写单元测试、端到端测试、测试计划或根据失败日志定位问题。使用时需要确认项目测试框架、运行命令和夹具数据,避免为了通过测试而改坏真实逻辑;涉及浏览器或外部服务时,应区分本地模拟、测试环境和生产环境。

总安装

618

周安装

26

GitHub Stars

2,078

下载量

216
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:running-chaos-tests(运行混沌测试)
来源仓库:https://github.com/jeremylongshore/claude-code-plugins-plus-skills
仓库路径:skills/running-chaos-tests
安装命令:
npx skills add https://github.com/jeremylongshore/claude-code-plugins-plus-skills --skill running-chaos-tests
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/jeremylongshore/claude-code-plugins-plus-skills --skill running-chaos-tests

简介

用于辅助测试设计、自动化测试、用例整理和回归验证。

  • 适合编写单元测试、端到端测试或根据失败日志定位问题。
  • 需确认项目测试框架、运行命令和夹具数据后使用。
  • 涉及浏览器或外部服务时,应区分本地模拟与生产环境。
  • running-chaos-tests 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

Chaos Engineering Toolkit

Overview

Execute controlled chaos engineering experiments to test system resilience, fault tolerance, and recovery capabilities. Injects failures including network latency, service crashes, resource exhaustion, and dependency outages to verify that systems degrade gracefully and recover automatically.

Prerequisites

  • Distributed system or microservice architecture deployed in a staging/test environment
  • Monitoring and alerting configured (Grafana, Datadog, CloudWatch, or Prometheus)
  • Rollback capability for the target environment (manual or automated)
  • Chaos engineering tool installed (toxiproxy, Pumba, Litmus, or Chaos Mesh)
  • Explicit approval from the team to run chaos experiments
  • Steady-state hypothesis defined (what "healthy" looks like in metrics)

Instructions

  1. Define the steady-state hypothesis:

- Identify measurable indicators of normal system behavior (e.g., p99 latency < 500ms, error rate < 0.1%, all health checks pass). - Record baseline metrics before injecting any failures. - Define the blast radius -- which services and users are affected by the experiment.

  1. Design chaos experiments by category:

- Network: Inject latency (200-2000ms), packet loss (5-50%), DNS failure, connection timeout. - Process: Kill a service instance, exhaust CPU or memory, fill disk. - Dependency: Block access to database, cache, or external API. - State: Corrupt data, introduce clock skew, simulate split-brain scenarios.

  1. Start with minimal impact and increase gradually:

- Begin with read-only experiments (network latency on non-critical path). - Progress to service-level failures (kill one instance of a multi-instance service). - Only move to data-level chaos after infrastructure chaos is validated.

  1. Execute each experiment with safeguards:

- Set a maximum experiment duration (5-15 minutes). - Configure automatic rollback triggers (error rate > 5% triggers abort). - Monitor system metrics in real-time during the experiment. - Have a manual kill switch ready (script to remove all injected failures immediately).

  1. Observe and record system behavior during the experiment:

- Did circuit breakers activate? How quickly? - Did auto-scaling trigger? How long until new instances were healthy? - Did retries succeed? Were they idempotent? - Did fallback mechanisms engage (cached responses, degraded mode)? - Were alerts triggered? Did on-call receive notification?

  1. After the experiment, verify full recovery:

- Remove all injected failures. - Verify steady-state hypothesis holds again within expected recovery time. - Check for data inconsistencies or orphaned state.

  1. Document findings and create action items for resilience improvements.

Output

  • Chaos experiment definition files (YAML or JSON) with hypothesis, method, and rollback
  • Experiment execution log with timeline of injected failures and observed effects
  • System behavior report covering circuit breakers, retries, fallbacks, and alerts
  • Recovery timeline showing time-to-detection and time-to-recovery
  • Action items for resilience improvements (retry policies, circuit breaker tuning, fallback additions)

Error Handling

ErrorCauseSolution
Experiment caused production outageBlast radius larger than expected or missing safeguardsAlways run in staging first; reduce scope; add automatic abort triggers; require approval
System did not recover after experimentAuto-healing mechanisms not configured or too slowAdd health-check-based restarts; configure auto-scaling; implement circuit breaker patterns
Monitoring missed the failureAlerting thresholds too lenient or wrong metrics monitoredTighten alert thresholds; add specific alerts for the failure mode tested; verify alert channels
Chaos tool cannot access targetNetwork segmentation or security policies blocking the toolDeploy chaos agent inside the target network; add security group rules for the chaos controller
Data corruption persists after rollbackStateful failure injection without transaction protectionUse read-only chaos first; snapshot databases before stateful experiments; implement compensating transactions

Examples

toxiproxy network latency injection:

set -euo pipefail
# Create a proxy for the database connection
toxiproxy-cli create postgres_proxy -l 0.0.0.0:15432 -u postgres-host:5432  # 15432: PostgreSQL port

# Inject 500ms latency
toxiproxy-cli toxic add postgres_proxy -t latency -a latency=500 -a jitter=100  # HTTP 500 Internal Server Error

# Run tests while latency is active
npm test -- --grep "handles slow database"

# Remove the toxic
toxiproxy-cli toxic remove postgres_proxy -n latency_downstream

Kubernetes pod kill experiment (Litmus Chaos):

apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: api-pod-kill
spec:
  appinfo:
    appns: default
    applabel: "app=api-server"
  chaosServiceAccount: litmus-admin
  experiments:
    - name: pod-delete
      spec:
        components:
          env:
            - name: TOTAL_CHAOS_DURATION
              value: "60"
            - name: CHAOS_INTERVAL
              value: "10"
            - name: FORCE
              value: "true"

Custom chaos script (process kill and verify recovery):

#!/bin/bash
set -euo pipefail
echo "=== Chaos Experiment: API server kill ==="
echo "Hypothesis: System recovers within 30 seconds"

# Record baseline
BASELINE=$(curl -s -o /dev/null -w '%{http_code}' http://app.test/health)
echo "Baseline health: $BASELINE"

# Kill one API instance
docker kill api-server-1

# Monitor recovery
for i in $(seq 1 30); do
  STATUS=$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 http://app.test/health)
  echo "T+${i}s: HTTP $STATUS"
  if [ "$STATUS" = "200" ]; then  # HTTP 200 OK
    echo "RECOVERED at T+${i}s"
    break
  fi
  sleep 1
done

Resources

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

33.58%
按下载量换算73

Claude

30.72%
按下载量换算66

Cursor

19.94%
按下载量换算43

Gemini CLI

10.04%
按下载量换算22

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

可疑

权限和风险

external-service

该 Skill 可能调用第三方服务、云服务或外部模型 API,使用前需要确认账号、额度、数据发送范围和服务条款。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills