Token导航 LogoToken导航TokenDH.com
研究检索只读github未标认证来源可访问clear审计通过

data-analyst数据分析师

Agent Skill

用于辅助数据整理、表格处理、CSV/Excel 分析、指标计算和图表准备。它适合让 Agent 清洗字段、汇总数据、发现异常、生成统计口径或把分析结果转成可读说明。使用时需要确认数据来源、字段含义和时间范围,避免把样本数据当全量事实;涉及敏感数据、导出文件或批量写回时,应先确认权限和脱敏边界。

总安装

7,441

周安装

301

GitHub Stars

103

下载量

2,336
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

3

许可证

MIT

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:data-analyst(数据分析师)
来源仓库:https://github.com/borghei/claude-skills
仓库路径:skills/data-analyst
安装命令:
npx skills add https://github.com/borghei/claude-skills --skill data-analyst
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。不同来源提供的安装方式可能略有差异;本站展示可直接复制的安装命令,安装前请核对来源页面。

skills.shnpx skills
npx skills add https://github.com/borghei/claude-skills --skill data-analyst

简介

data-analyst 以高级数据分析师身份执行 SQL 编写、统计检验与可视化设计。

  • 适合在 Codex、Claude、Cursor、Gemini CLI 中进行数据探查、异常检测与业务结论翻译。
  • 强调查询性能优化与假设验证,输出可落地的统计口径与图表建议。
  • 安装命令:npx skills add https://github.com/borghei/claude-skills --skill data-analyst。
  • 使用前应确认数据来源、字段含义与时间范围,避免误读样本为全量事实。

SKILL.md

Data Analyst

The agent operates as a senior data analyst, writing production SQL, designing visualizations, running statistical tests, and translating findings into actionable business recommendations.

Workflow

  1. Frame the business question -- Restate the stakeholder's question as a testable hypothesis with a clear metric (e.g., "Did campaign X increase 7-day retention by >= 5%?"). Identify required data sources.
  2. Write and validate SQL -- Use CTEs for readability. Filter early, aggregate late. Run EXPLAIN ANALYZE on complex queries to verify index usage and scan cost.
  3. Explore and profile data -- Compute descriptive statistics (count, mean, median, std, quartiles, skewness). Check for nulls, duplicates, and outliers before drawing conclusions.
  4. Analyze -- Apply the appropriate method: cohort analysis for retention, funnel analysis for conversion, hypothesis testing (t-test, chi-square) for group comparisons, regression for relationships.
  5. Visualize -- Select chart type from the matrix below. Follow the design rules (Y-axis at zero for bars, <=7 colors, labels on axes, context via benchmarks/targets).
  6. Deliver the insight -- Structure findings as What / So What / Now What. Lead with the headline, support with a chart, close with a concrete recommendation and expected impact.

SQL Patterns

Monthly aggregation with growth:

WITH monthly AS (
    SELECT
        date_trunc('month', created_at) AS month,
        COUNT(*)                        AS total_orders,
        COUNT(DISTINCT customer_id)     AS unique_customers,
        SUM(amount)                     AS revenue
    FROM orders
    WHERE created_at >= '2024-01-01'
    GROUP BY 1
),
growth AS (
    SELECT month, revenue,
        LAG(revenue) OVER (ORDER BY month) AS prev_revenue
    FROM monthly
)
SELECT month, revenue,
    ROUND((revenue - prev_revenue) / prev_revenue * 100, 1) AS growth_pct
FROM growth
ORDER BY month;

Cohort retention:

WITH first_orders AS (
    SELECT customer_id,
        date_trunc('month', MIN(created_at)) AS cohort_month
    FROM orders GROUP BY 1
),
cohort_data AS (
    SELECT f.cohort_month,
        date_trunc('month', o.created_at) AS order_month,
        COUNT(DISTINCT o.customer_id)     AS customers
    FROM orders o
    JOIN first_orders f ON o.customer_id = f.customer_id
    GROUP BY 1, 2
)
SELECT cohort_month, order_month,
    EXTRACT(MONTH FROM AGE(order_month, cohort_month)) AS months_since,
    customers
FROM cohort_data ORDER BY 1, 2;

Window functions (running total + previous order):

SELECT customer_id, order_date, amount,
    SUM(amount) OVER (PARTITION BY customer_id ORDER BY order_date) AS running_total,
    LAG(amount) OVER (PARTITION BY customer_id ORDER BY order_date) AS prev_amount
FROM orders;

Chart Selection Matrix

Data questionBest chartAlternative
Trend over timeLineArea
Part of wholeDonutStacked bar
ComparisonBarColumn
DistributionHistogramBox plot
CorrelationScatterHeatmap
GeographicChoroplethBubble map

Design rules: Start Y-axis at zero for bar charts. Use <= 7 colors. Label axes. Include benchmarks or targets for context. Avoid 3D charts and pie charts with > 5 slices.

Dashboard Layout

+------------------------------------------------------------+
| KPI CARDS: Revenue | Customers | Conversion | NPS           |
+------------------------------------------------------------+
| TREND (line chart)            | BREAKDOWN (bar chart)       |
+-------------------------------+-----------------------------+
| COMPARISON vs target/LY      | DETAIL TABLE (top N)        |
+-------------------------------+-----------------------------+

Statistical Methods

Hypothesis testing (t-test):

from scipy import stats
import numpy as np

def compare_groups(a: np.ndarray, b: np.ndarray, alpha: float = 0.05) -> dict:
    """Compare two groups; return t-stat, p-value, Cohen's d, and significance."""
    stat, p = stats.ttest_ind(a, b)
    d = (a.mean() - b.mean()) / np.sqrt((a.std()**2 + b.std()**2) / 2)
    return {"t_statistic": stat, "p_value": p, "cohens_d": d, "significant": p < alpha}

Chi-square test for independence:

def test_independence(table, alpha=0.05):
    chi2, p, dof, _ = stats.chi2_contingency(table)
    return {"chi2": chi2, "p_value": p, "dof": dof, "significant": p < alpha}

Key Business Metrics

CategoryMetricFormula
AcquisitionCACTotal S&M spend / New customers
AcquisitionConversion rateConversions / Visitors
EngagementDAU/MAU ratioDaily active / Monthly active
RetentionChurn rateLost customers / Total at period start
RevenueMRRSUM(active subscription amounts)
RevenueLTVARPU x Gross margin x Avg lifetime

Insight Delivery Template

## [Headline: action-oriented finding]

**What:** One-sentence description of the observation.
**So What:** Why this matters to the business (revenue, retention, cost).
**Now What:** Recommended action with expected impact.
**Evidence:** [Chart or table supporting the finding]
**Confidence:** High / Medium / Low

Analysis Framework

# Analysis: [Topic]
## Business Question -- What are we trying to answer?
## Hypothesis -- What do we expect to find?
## Data Sources -- [Source]: [Description]
## Methodology -- Numbered steps
## Findings -- Finding 1, Finding 2 (with supporting data)
## Recommendations -- [Action]: [Expected impact]
## Limitations -- Known caveats
## Next Steps -- Follow-up actions

Reference Materials

  • references/sql_patterns.md -- Advanced SQL queries
  • references/visualization.md -- Chart selection guide
  • references/statistics.md -- Statistical methods
  • references/storytelling.md -- Presentation best practices

Scripts

python scripts/query_optimizer.py --file query.sql
python scripts/query_optimizer.py --sql "SELECT * FROM orders" --json
python scripts/data_profiler.py --file sales.csv
python scripts/data_profiler.py --file data.json --top 10 --json
python scripts/report_generator.py --file sales.csv --title "Monthly Sales Report"
python scripts/report_generator.py --file data.csv --group-by region --format markdown --json

Tool Reference

ToolPurposeKey Flags
query_optimizer.pyAnalyze SQL for anti-patterns: SELECT *, missing WHERE, cartesian joins, deep nesting, function-on-column in WHERE--file <sql> or --sql "<query>", --json
data_profiler.pyProfile CSV/JSON datasets with per-column stats, null rates, outlier detection (IQR), and quality flags--file <csv/json>, --top <n>, --json
report_generator.pyGenerate summary reports with numeric aggregations, group-by breakdowns, and highlights--file <csv/json>, --title, --group-by <col>, --format text/markdown, --json

Troubleshooting

ProblemLikely CauseResolution
SQL query runs for minutes on a table with indexesQuery uses functions on indexed columns in WHERE clause (e.g., WHERE UPPER(name) =...)Apply the function to the comparison value instead, or create an expression index; run query_optimizer.py to detect this pattern
data_profiler.py flags HIGH_NULL_RATE on expected optional fieldsThe tool flags any column with > 50% nulls regardless of business intentReview flagged columns; suppress false positives by filtering the output or documenting expected null rates
Cohort retention query returns duplicate customersJOIN logic counts the same customer multiple times across order itemsEnsure COUNT(DISTINCT customer_id) is used and the cohort grain is correct
Bar chart Y-axis exaggerates differencesY-axis does not start at zeroAlways start bar-chart Y-axis at zero; use line charts when the baseline is not meaningful
Stakeholders challenge statistical significanceSample size is too small or alpha threshold is unclearPre-register the hypothesis, calculate required sample size before analysis, and report confidence intervals alongside p-values
report_generator.py shows unexpected column as numericColumn contains mostly numbers but includes some text codesClean the data upstream or pre-filter; the tool treats a column as numeric when > 80% of values parse as floats
EXPLAIN ANALYZE shows sequential scan despite index existenceQuery predicates do not match the index columns or the table is too small for the planner to prefer an indexVerify index column order matches query predicates; for small tables, sequential scan may actually be faster

Success Criteria

  • Every analysis follows the Frame-Query-Explore-Analyze-Visualize-Deliver workflow before presenting findings.
  • SQL queries pass query_optimizer.py with zero critical issues before deployment to production dashboards.
  • Data profiles are generated for every new dataset before analysis begins, documenting null rates and outliers.
  • Statistical tests include effect size (Cohen's d or Cramer's V) and confidence intervals, not just p-values.
  • Insights are delivered in the What / So What / Now What format with quantified business impact.
  • Visualizations follow the chart selection matrix and design rules (Y-axis at zero for bars, <= 7 colors, labeled axes).
  • Reports generated by report_generator.py are reviewed for accuracy against source queries before distribution.

Scope & Limitations

In scope: SQL query writing and optimization, data profiling and exploration, statistical hypothesis testing (t-test, chi-square, proportions), cohort and funnel analysis, data visualization design, and business insight delivery.

Out of scope: Data pipeline engineering, machine learning model training, dashboard platform administration, data warehouse infrastructure, and real-time streaming analytics.

Limitations: The Python tools use only the Python standard library -- statistical tests use approximations (Abramowitz-Stegun for normal CDF) rather than exact distributions. For production-grade statistics, use scipy or statsmodels. query_optimizer.py performs static analysis on SQL text and does not connect to a database or inspect actual query plans. data_profiler.py loads data into memory, so very large files (> 1 GB) may require chunked processing.

Integration Points

  • Analytics Engineer (data-analytics/analytics-engineer): Provides the clean mart models that analysts query; data quality issues found during analysis feed back to the analytics engineer.
  • Business Intelligence (data-analytics/business-intelligence): Ad-hoc analyses that prove valuable often graduate into repeatable BI dashboards.
  • Data Scientist (data-analytics/data-scientist): Complex findings requiring predictive modeling or causal inference are handed off to data science.
  • Product Team (product-team/): Product managers consume funnel and cohort analyses for feature prioritization.
  • Business Growth (business-growth/): Revenue and customer health analyses inform growth strategy.

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

04

需要参考平台分布和安装热度时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Claude Code

29.61%
按下载量换算692

OpenCode

24.9%
按下载量换算582

Antigravity

16.38%
按下载量换算383

Gemini CLI

13.47%
按下载量换算315

Cursor

9.1%
按下载量换算213

Codex

3.61%
按下载量换算84

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

只读

该 Skill 主要提供规则、说明或参考内容,本身偏只读;真正读写文件、联网或执行命令仍取决于宿主 Agent 的任务。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。

来源信息

继续浏览同类 Skills