Token导航 LogoToken导航TokenDH.com
待分类敏感数据github未标认证来源可访问许可证需确认审计通过

tech-data-pipeline技术数据管道

Agent Skill

用于辅助数据整理、表格处理、CSV/Excel 分析、指标计算和图表准备。它适合让 Agent 清洗字段、汇总数据、发现异常、生成统计口径或把分析结果转成可读说明。使用时需要确认数据来源、字段含义和时间范围,避免把样本数据当全量事实;涉及敏感数据、导出文件或批量写回时,应先确认权限和脱敏边界。

总安装

353

周安装

15

GitHub Stars

125

下载量

124
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:tech-data-pipeline(技术数据管道)
来源仓库:https://github.com/asgard-ai-platform/skills
仓库路径:skills/tech-data-pipeline
安装命令:
npx skills add https://github.com/asgard-ai-platform/skills --skill tech-data-pipeline
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/asgard-ai-platform/skills --skill tech-data-pipeline

简介

用于辅助数据整理和表格分析处理。tech-data-pipeline 属于待分类类 Skill,可作为该场景下的辅助能力补充。

  • 适合清洗字段、汇总指标或生成统计口径。
  • 使用时需确认数据来源和时间范围,避免误判全量事实。
  • 涉及敏感数据时应先确认脱敏边界。适用宿主包括 Codex、Claude、Cursor、Gemini CLI,接入前应确认版本、权限和运行环境要求。
  • 导出文件前建议核对权限和操作影响。

SKILL.md

Data Pipeline Design

Framework

IRON LAW: Data Quality Checks at Every Stage

A pipeline that moves bad data fast is worse than no pipeline — it
corrupts downstream analytics and decisions. Every pipeline stage
(extract, transform, load) must have data quality checks:
row counts, null checks, schema validation, freshness checks.

"Garbage in, garbage out" is not a warning — it's a guarantee.

ETL vs ELT

AspectETL (Extract, Transform, Load)ELT (Extract, Load, Transform)
Transform where?Before loading (in pipeline)After loading (in warehouse)
Best forStructured data, compliance-heavyCloud warehouses (BigQuery, Snowflake)
FlexibilityLess (transform logic is fixed)More (transform in SQL after loading)
CostCompute in pipelineCompute in warehouse
TrendLegacy/on-premModern/cloud-native

Pipeline Architecture

[Sources] → [Extract] → [Stage] → [Transform] → [Load] → [Serve]
   ↑                                                          ↓
   |              [Quality Checks at every stage]        [Dashboard]
   |              [Monitoring & Alerting]                [API]
   └──────────────── [Orchestrator (Airflow/Prefect)] ──────┘

Data Source Types

SourceExtraction MethodChallenges
DatabaseCDC (Change Data Capture), bulk query, replicationSchema changes, performance impact on source
APIREST/GraphQL polling, webhooksRate limits, pagination, auth token refresh
FilesS3/GCS pickup, SFTP, email attachmentFormat inconsistency, encoding issues
StreamingKafka, Kinesis, Pub/SubOrdering, exactly-once processing
SaaS toolsPre-built connectors (Fivetran, Airbyte)API changes, data model complexity

Orchestration Tools

ToolTypeBest ForComplexity
AirflowPython DAGsComplex pipelines, team of engineersHigh
PrefectPython, modern APISimpler than Airflow, good DXMedium
dbtSQL transforms onlyTransform layer in ELTLow-Medium
CronSimple schedulingSingle script, low complexityLow
Fivetran/AirbyteManaged connectorsExtract + Load (no transform)Low

Data Quality Framework

CheckWhat It ValidatesWhen
Row countExpected number of rows (within ±10% of prior run)After extract, after load
Null checkCritical columns have no unexpected nullsAfter extract
Schema validationColumn names, types match expectedAfter extract
FreshnessData is recent (not stale)After load
UniquenessNo duplicate primary keysAfter load
Range checkValues within expected boundsAfter transform
Referential integrityForeign keys match parent tablesAfter load

Pipeline Design Steps

  1. Map sources and destinations: What data, from where, to where?
  2. Define freshness requirements: Real-time? Hourly? Daily?
  3. Choose architecture: ETL or ELT based on tools and team
  4. Build incrementally: Start with one source, one destination, one schedule
  5. Add quality checks: At minimum: row count + null check + freshness
  6. Set up monitoring: Alert on failure, quality check violations, latency
  7. Document: Data dictionary, pipeline diagram, SLAs

Output Format

# Data Pipeline Design: {Project}

## Sources & Destinations
| Source | Type | Destination | Freshness | Volume |
|--------|------|-----------|-----------|--------|
| {source} | DB/API/File | {dest} | {daily/hourly} | {rows/day} |

## Architecture
- Pattern: ETL / ELT
- Orchestrator: {tool}
- Transform: {tool/SQL}
- Quality: {tool/custom checks}

## Pipeline Diagram
{Source} → {Extract} → {Stage} → {Transform} → {Load} → {Serve}

## Quality Checks
| Stage | Check | Threshold | Alert |
|-------|-------|-----------|-------|
| Extract | Row count | ±10% of prior | Slack alert |
| Load | Freshness | < 6 hours old | PagerDuty |

## Schedule
| Pipeline | Frequency | Start Time | SLA |
|----------|-----------|-----------|-----|
| {name} | {daily/hourly} | {time} | Data ready by {time} |

Gotchas

  • Idempotency is essential: A pipeline that runs twice should produce the same result as running once. Use upsert (not insert) and date-partitioned loads.
  • Schema drift: Source systems change schemas without warning. Build schema detection and alerting.
  • Backfill capability: When a pipeline fails for 3 days, can you rerun for those days without duplicating data? Design for this from day 1.
  • Don't build what you can buy: Fivetran/Airbyte handle 200+ source connectors. Writing a custom Salesforce extractor is rarely worth the engineering time.
  • Data warehouse vs data lake: Warehouse (BigQuery, Snowflake) = structured, SQL-queryable. Lake (S3, GCS) = raw, any format. Most modern stacks use both (lakehouse pattern).

References

  • For dbt project structure, see references/dbt-guide.md
  • For data warehouse modeling (star schema), see references/dimensional-modeling.md

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

32.18%
按下载量换算40

Claude

29.49%
按下载量换算37

Cursor

19.35%
按下载量换算24

Gemini CLI

9.6%
按下载量换算12

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

敏感数据

该 Skill 可能接触密钥、Token、环境变量或敏感配置,应进入高风险复核队列,默认不自动发布。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills