Token导航 LogoToken导航TokenDH.com
前端设计只读github未标认证来源可访问许可证需确认审计通过

data-python数据 Python

Agent Skill

用于辅助 Python 项目开发、测试、依赖管理和常见框架工作流。它适合让 Agent 阅读 Python 代码、定位测试问题、整理运行命令、生成脚本或分析数据处理逻辑。使用时需要确认项目虚拟环境、依赖版本和测试入口;涉及执行脚本、读写文件、访问数据库或调用外部 API 时,应先明确运行目录和输入输出范围,避免误改生产数据。

总安装

519

周安装

21

GitHub Stars

1

下载量

163
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:data-python(数据 Python)
来源仓库:https://github.com/alexanderstephenthompson/claude-hub
仓库路径:skills/data-python
安装命令:
npx skills add https://github.com/alexanderstephenthompson/claude-hub --skill data-python
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/alexanderstephenthompson/claude-hub --skill data-python

简介

针对 Python 数据处理项目的最佳实践指导,强调向量化操作。

  • 避免 iterrows() 和链式索引等低效写法,关注数据类型优化。
  • 提供 pandas、polars 和 pyspark 等主流库的性能优化建议。
  • 处理真实数据前务必在样本集上验证代码效率与正确性。
  • data-python 属于前端设计类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

Data Python Skill

Version: 1.0 Stack: Python (pandas, polars, pyspark)

Python makes it easy to write data processing code that works on sample data and fails on real data. iterrows() takes 30 seconds on 10K rows and 30 minutes on 10M. A DataFrame without explicit dtypes uses 8x the memory it needs. Chained indexing creates silent copies that lose your changes. These aren't edge cases — they're the default behavior of pandas when you write it like regular Python.

Vectorized operations, explicit schemas, and proper dtypes mean your code scales from prototype to production without rewriting.


Scope and Boundaries

This skill covers:

  • DataFrame patterns (pandas, polars, pyspark)
  • Data type handling and validation
  • Memory-efficient processing
  • Vectorized operations over loops
  • Method chaining patterns
  • Error handling for data pipelines

Defers to other skills:

  • code-quality: General code structure, testing, naming
  • data-sql: Query patterns when using SQL interfaces
  • data-pipelines: Orchestration and ETL architecture

Use this skill when: Writing Python code that processes data.


Core Principles

  1. Vectorize, Don't Loop — Use DataFrame operations, not row iteration.
  2. Fail Fast on Bad Data — Validate early, reject invalid data at boundaries.
  3. Memory Awareness — Know your data size, use appropriate dtypes.
  4. Immutable Transforms — Chain operations, don't mutate in place.
  5. Explicit Schemas — Define expected columns and types upfront.

Patterns

Method Chaining

# Good - readable pipeline
result = (
    df
    .query("status == 'active'")
    .assign(total=lambda x: x["quantity"] * x["price"])
    .groupby("category")
    .agg({"total": "sum"})
    .sort_values("total", ascending=False)
)

# Bad - intermediate variables obscure flow
filtered = df[df["status"] == "active"]
filtered["total"] = filtered["quantity"] * filtered["price"]
grouped = filtered.groupby("category")
result = grouped.agg({"total": "sum"})
result = result.sort_values("total", ascending=False)

Schema Validation

EXPECTED_COLUMNS = {"id", "name", "value", "timestamp"}
REQUIRED_COLUMNS = {"id", "value"}

def validate_schema(df: pd.DataFrame) -> pd.DataFrame:
    missing = REQUIRED_COLUMNS - set(df.columns)
    if missing:
        raise ValueError(f"Missing required columns: {missing}")
    return df[list(EXPECTED_COLUMNS & set(df.columns))]

Type Optimization

def optimize_dtypes(df: pd.DataFrame) -> pd.DataFrame:
    """Downcast numeric types to reduce memory."""
    for col in df.select_dtypes(include=["int"]).columns:
        df[col] = pd.to_numeric(df[col], downcast="integer")
    for col in df.select_dtypes(include=["float"]).columns:
        df[col] = pd.to_numeric(df[col], downcast="float")
    return df

Anti-Patterns

Anti-PatternProblemFix
for row in df.iterrows()Slow, defeats vectorizationUse vectorized operations
df["col"] = df.apply(...)Usually slower than vectorizedUse np.where or df.assign
Chained indexing df["a"]["b"]SettingWithCopyWarningUse df.loc[:, "a"]
Loading entire file to check schemaWastes memoryUse nrows=100 or chunking
Ignoring dtypes on readMemory bloatSpecify dtype= parameter

Checklist

  • No iterrows() or itertuples() for computation
  • Explicit dtypes on file reads
  • Schema validation at boundaries
  • Method chaining for transforms
  • Memory profiled for large datasets

References

  • references/vectorization.md — Vectorized operations and performance
  • references/memory-optimization.md — Memory optimization techniques

Assets

  • assets/pandas-cheatsheet.md — Quick reference for pandas operations

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

34.74%
按下载量换算57

Claude

29.53%
按下载量换算48

Cursor

19.9%
按下载量换算32

Gemini CLI

10.07%
按下载量换算16

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

只读

该 Skill 主要提供规则、说明或参考内容,本身偏只读;真正读写文件、联网或执行命令仍取决于宿主 Agent 的任务。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills