Token导航 LogoToken导航TokenDH.com
研究检索只读github未标认证来源可访问许可证需确认审计通过

ab-test-analysisab 测试分析

Agent Skill

用于辅助测试设计、自动化测试、用例整理和回归验证。它适合让 Agent 编写单元测试、端到端测试、测试计划或根据失败日志定位问题。使用时需要确认项目测试框架、运行命令和夹具数据,避免为了通过测试而改坏真实逻辑;涉及浏览器或外部服务时,应区分本地模拟、测试环境和生产环境。

总安装

13,513

周安装

552

GitHub Stars

10,828

下载量

4,372
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:ab-test-analysis(ab 测试分析)
来源仓库:https://github.com/phuryn/pm-skills
仓库路径:skills/ab-test-analysis
安装命令:
npx skills add https://github.com/phuryn/pm-skills --skill ab-test-analysis
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/phuryn/pm-skills --skill ab-test-analysis

简介

ab-test-analysis 对 A/B 测试结果进行严谨统计分析,识别显著性并转化为产品决策建议。

  • 适用于电商、广告或用户行为实验的数据解读,支持 CSV/Excel 数据文件直接分析。
  • 自动验证样本量、运行时长与统计功效,生成置信区间与概率优于基线的结论。
  • 需确保原始数据完整准确,警惕辛普森悖论;建议结合业务背景判断结果合理性。
  • 可输出 Python 脚本复现分析过程,但不替代人工复核关键指标的实际意义。

SKILL.md

A/B Test Analysis

Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.

Context

You are analyzing A/B test results for $ARGUMENTS.

If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.

Instructions

  1. Understand the experiment:

- What was the hypothesis? - What was changed (the variant)? - What is the primary metric? Any guardrail metrics? - How long did the test run? - What is the traffic split?

  1. Validate the test setup:

- Sample size: Is the sample large enough for the expected effect size? - Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE² - Flag if the test is underpowered (<80% power) - Duration: Did the test run for at least 1-2 full business cycles? - Randomization: Any evidence of sample ratio mismatch (SRM)? - Novelty/primacy effects: Was there enough time to wash out initial behavior changes?

  1. Calculate statistical significance: If the user provides raw data, generate and run a Python script to calculate these.

- Conversion rate for control and variant - Relative lift: (variant - control) / control × 100 - p-value: Using a two-tailed z-test or chi-squared test - Confidence interval: 95% CI for the difference - Statistical significance: Is p < 0.05? - Practical significance: Is the lift meaningful for the business?

  1. Check guardrail metrics:

- Did any guardrail metrics (revenue, engagement, page load time) degrade? - A winning primary metric with degraded guardrails may not be a true win

  1. Interpret results: Outcome Recommendation Significant positive lift, no guardrail issues Ship it — roll out to 100% Significant positive lift, guardrail concerns Investigate — understand trade-offs before shipping Not significant, positive trend Extend the test — need more data or larger effect Not significant, flat Stop the test — no meaningful difference detected Significant negative lift Don't ship — revert to control, analyze why
  2. Provide the analysis summary: ## A/B Test Results: [Test Name] **Hypothesis**: [What we expected] **Duration**: [X days] | **Sample**: [N control / M variant] | Metric | Control | Variant | Lift | p-value | Significant? | |---|---|---|---|---|---| | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No | | [Guardrail] |... |... |... |... |... | **Recommendation**: [Ship / Extend / Stop / Investigate] **Reasoning**: [Why] **Next steps**: [What to do]

Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.


Further Reading

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

33.09%
按下载量换算1,447

Claude

29.26%
按下载量换算1,279

Cursor

17.33%
按下载量换算758

Gemini CLI

10.43%
按下载量换算456

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

只读

该 Skill 主要提供规则、说明或参考内容,本身偏只读;真正读写文件、联网或执行命令仍取决于宿主 Agent 的任务。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills