Token导航 LogoToken导航TokenDH.com
研究检索权限需确认github未标认证来源可访问许可证需确认审计通过

test-flakiness测试片状

Agent Skill

用于辅助测试设计、自动化测试、用例整理和回归验证。它适合让 Agent 编写单元测试、端到端测试、测试计划或根据失败日志定位问题。使用时需要确认项目测试框架、运行命令和夹具数据,避免为了通过测试而改坏真实逻辑;涉及浏览器或外部服务时,应区分本地模拟、测试环境和生产环境。

总安装

442

周安装

19

GitHub Stars

16,636

下载量

155
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:test-flakiness(测试片状)
来源仓库:https://github.com/donchitos/claude-code-game-studios
仓库路径:skills/test-flakiness
安装命令:
npx skills add https://github.com/donchitos/claude-code-game-studios --skill test-flakiness
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/donchitos/claude-code-game-studios --skill test-flakiness

简介

用于辅助测试设计、自动化测试、用例整理和回归验证。

  • 适合让 Agent 编写单元测试、端到端测试、测试计划或根据失败日志定位问题。
  • 使用时需要确认项目测试框架、运行命令和夹具数据,避免为了通过测试而改坏真实逻辑。
  • 涉及浏览器或外部服务时,应区分本地模拟、测试环境和生产环境。
  • 安装方式:github,需通过 npx skills add 命令从指定仓库添加。

SKILL.md

Test Flakiness Detection

A flaky test is one that sometimes passes and sometimes fails without any code change. Flaky tests are worse than no tests in some ways — they train the team to ignore red CI runs, masking genuine failures. This skill identifies them, explains likely causes, and recommends whether to quarantine or fix each one.

Output: Updated tests/regression-suite.md quarantine section + optional production/qa/flakiness-report-[date].md

When to run:

  • Polish phase (tests have had many runs; statistical signal is reliable)
  • When developers start dismissing CI failures as "probably flaky"
  • After /regression-suite identifies quarantined tests that need diagnosis

1. Parse Arguments

Modes:

  • /test-flakiness [ci-log-path] — analyse a specific CI run log file
  • /test-flakiness scan — scan all available CI logs in .github/ or standard log output directories
  • /test-flakiness registry — read existing regression-suite.md quarantine section and provide remediation guidance for already-known flaky tests
  • No argument — auto-detect: run scan if CI logs are accessible, else registry

2. Locate CI Log Data

Option A — GitHub Actions (preferred)

Check for test result artifacts:

ls -t .github/ 2>/dev/null
ls -t test-results/ 2>/dev/null

For Godot projects: GdUnit4 outputs XML results compatible with JUnit format. Check test-results/ for .xml files.

For Unity projects: game-ci test runner outputs NUnit XML to test-results/ by default.

For Unreal projects: automation logs go to Saved/Logs/. Grep for Result: Success and Result: Fail patterns.

Option B — Local log files

If a path argument is provided, read that file directly.

Option C — No log data available

If no logs found:

"No CI log data found. To detect flaky tests, this skill needs test result history from multiple runs. Options: 1. Run the test suite at least 3 times and collect the output logs 2. Check CI pipeline output and save a log to test-results/ 3. Run /test-flakiness registry to review tests already flagged as flaky in tests/regression-suite.md"

Stop and ask the user which option to pursue.


3. Parse Test Results

For each CI log or result file found, parse:

JUnit XML format (GdUnit4 / Unity):

  • Grep for <testcase name= to get test names
  • Grep for <failure or <error to identify failures
  • Parse classname and name attributes for full test identifiers

Plain text logs:

  • Grep for pass/fail patterns:

- Godot: PASSED / FAILED adjacent to test names - Unreal: Result: Success / Result: Fail - Unity: Test passed / Test failed

Build a table: test_id → [run1_result, run2_result, run3_result,...]


4. Identify Flaky Tests

A test is flaky if it appears in the result history with both PASS and FAIL outcomes across runs with no code changes between them.

Flakiness thresholds:

  • High flakiness: Fails in >25% of runs — quarantine immediately
  • Moderate flakiness: Fails in 5–25% of runs — investigate and fix soon
  • Low/suspected flakiness: Fails in 1–5% of runs — monitor; may be genuinely rare failure

For each flaky test, classify the likely cause:

Cause classification

CauseSymptomsFix direction
Timing / asyncFails after awaiting signals or timers; pass rate correlates with system loadAdd explicit await/synchronisation; avoid time-based delays
Order dependencyFails when run after specific other tests; passes in isolationAdd proper setup/teardown; ensure test isolation
Random seedFails intermittently with no pattern; involves RNGPass explicit seed; don't use randf() in tests
Resource leakFails more often later in a test runFix cleanup in teardown; check orphan nodes (Godot) or object disposal (Unity)
External stateFails when a file, scene, or global exists from a prior testIsolate test from file system; use in-memory mocks
Floating pointFails on comparisons like == 0.5Use epsilon comparison (is_equal_approx, Assert.AreApproximately)
Scene/prefab load raceFails when scenes are not yet readyAwait one frame after instantiation; use await get_tree().process_frame

Use Grep to check the test file for timing calls, randf, global state access, or equality comparisons on floats to narrow down the cause.


5. Recommend Action

For each flaky test:

Quarantine (High flakiness):

"Quarantine this test immediately. Disable it in CI by adding @pytest.mark.skip / [Ignore] / GdUnitSkip annotation. Log it in tests/regression-suite.md quarantine section. The test is now opt-in only. Fix the root cause before removing quarantine."

Investigate and fix soon (Moderate):

"This test is intermittently unreliable. Root cause appears to be [cause]. Suggested fix: [specific fix based on cause classification]. Do not quarantine yet — fix the test directly."

Monitor (Low/suspected):

"This test shows suspected flakiness. Collect more run data before quarantining. Note it as 'suspected' in the regression suite."

6. Generate Reports

In-conversation summary

## Flakiness Detection Results

**Runs analysed**: [N]
**Tests tracked**: [N]

### Flaky Tests Found

| Test | System | Fail Rate | Likely Cause | Recommendation |
|------|--------|-----------|--------------|----------------|
| [test_name] | [system] | [N]% | Timing | Quarantine + fix async |
| [test_name] | [system] | [N]% | Float comparison | Fix: use epsilon compare |
| [test_name] | [system] | [N]% | Order dependency | Investigate teardown |

### Clean Tests (no flakiness detected)

[N] tests ran across [N] runs with consistent results — no flakiness detected.

### Data Limitations

[Note if fewer than 5 runs were available — fewer runs = less statistical confidence]

7. Update Regression Suite + Optional Report File

Ask: "May I update the quarantine section of tests/regression-suite.md with the flaky tests found?"

If yes: use Edit to append entries to the Quarantined Tests table. Never remove existing quarantine entries — only add new ones.

Ask (separately): "May I write a full flakiness report to production/qa/flakiness-report-[date].md?"

The full report includes per-test analysis with cause details and engine-specific fix snippets.

After writing:

  • For each quarantined test: "Add the engine-specific skip annotation to disable this test in CI. Re-enable after the root cause is fixed."
  • For fix-eligible tests: "The fix for [test] is straightforward — change the equality comparison on line [N] to use is_equal_approx."
  • Summary: "Once all quarantine annotations are applied, CI should run green. Schedule fix work for the [N] quarantined tests before the release gate."

Collaborative Protocol

  • Never delete test files — quarantine means annotate + list, not remove
  • Statistical confidence matters — with < 3 runs, flag findings as "suspected" not "confirmed"; ask if more run data is available
  • Fix is always the goal — quarantine is temporary; surface the fix direction even when recommending quarantine
  • Ask before writing — both the regression-suite update and the report file require explicit approval. On write: Verdict: COMPLETE — flakiness report written. On decline: Verdict: BLOCKED — user declined write.
  • Flakiness in CI is a team problem — surface the list and recommended actions clearly; do not just silently quarantine without the team knowing

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

32.89%
按下载量换算51

Claude

32.25%
按下载量换算50

Cursor

19.22%
按下载量换算30

Gemini CLI

9.59%
按下载量换算15

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

权限需确认

当前来源未能明确判断权限范围,默认进入异常复核队列。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills