Token导航 LogoToken导航TokenDH.com
待分类需要联网github未标认证来源可访问许可证需确认审计通过

parquet-analysis镶木地板分析

Agent Skill

parquet-analysis 用于处理 GitHub 仓库、Issue、Pull Request 和代码协作信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要围绕仓库状态、代码变更或协作事项进行整理时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

285

周安装

12

GitHub Stars

公开资料未说明

下载量

1
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:parquet-analysis(镶木地板分析)
来源仓库:https://github.com/brojonat/llmsrules
仓库路径:skills/parquet-analysis
安装命令:
npx skills add https://github.com/brojonat/llmsrules --skill parquet-analysis
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/brojonat/llmsrules --skill parquet-analysis

简介

parquet-analysis 用于处理 GitHub 仓库、Issue、Pull Request 和代码协作信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要围绕仓库状态、代码变更或协作事项进行整理时使用。

  • 适用于待分类任务,支持 Parquet 格式相关的协作流程管理。
  • 通过 npx skills add 命令从指定 GitHub 仓库安装,需确认权限范围和联网能力。
  • 建议结合原始 README 核验具体用法,注意维护状态及是否触发文件读写或命令执行。
  • 适用宿主包括 Codex、Claude、Cursor、Gemini CLI,接入前应确认版本、权限和运行环境要求。

SKILL.md

Parquet Analysis with Python and Ibis

Analyze parquet files using Ibis, a database-agnostic Python DataFrame API. Ibis translates Python operations into optimized queries for the underlying backend (DuckDB by default for parquet files).

Quick Start

Basic workflow

import ibis

# Connect to DuckDB (optimized for parquet)
con = ibis.duckdb.connect()

# Read parquet file
table = con.read_parquet("data.parquet")

# Explore
print(table.schema())      # View schema
print(table.head(10))      # Preview data
print(table.describe())    # Summary statistics

# Filter and aggregate
summary = (
    table
    .filter(table.amount > 100)
    .group_by("category")
    .aggregate(
        total=table.amount.sum(),
        avg=table.amount.mean(),
        count=table.count()
    )
)
print(summary.execute())

# Export
con.to_parquet(summary, "output.parquet")

Core Operations

Select and filter

# Select columns
selected = table.select("id", "amount", "date")

# Filter rows
filtered = table.filter(
    (table.amount > 100) &
    (table.date >= "2024-01-01")
)

# Sort
sorted_data = table.order_by(table.amount.desc())

Transform

# Add computed columns
enriched = table.mutate(
    revenue=table.quantity * table.unit_price,
    year=table.date.year(),
    size=ibis.case()
        .when(table.amount < 100, "small")
        .when(table.amount < 1000, "medium")
        .else_("large")
        .end()
)

Aggregate

# Group by and summarize
by_category = (
    table.group_by("category")
    .aggregate(
        total=table.amount.sum(),
        avg=table.amount.mean(),
        count=table.count()
    )
)

Join

# Read and join multiple files
customers = con.read_parquet("customers.parquet")
orders = con.read_parquet("orders.parquet")

joined = (
    orders
    .join(customers, orders.customer_id == customers.id, how="left")
    .select(
        orders.order_id,
        orders.amount,
        customers.name
    )
)

Export

# To parquet
con.to_parquet(result, "output.parquet")

# To CSV (via pandas)
df = result.execute()
df.to_csv("output.csv", index=False)

Common Patterns

Data quality checks

# Row count
row_count = table.count().execute()

# Check for nulls
null_counts = table.select([
    col.isnull().sum().name(f"{col}_nulls")
    for col in table.columns
]).execute()

# Value distribution
table.group_by("category").aggregate(count=table.count()).execute()

Time-based analysis

# Monthly aggregation
monthly = (
    table.mutate(month=table.date.truncate("M"))
    .group_by("month")
    .aggregate(total=table.amount.sum())
    .order_by("month")
)

# Extract date components
dated = table.mutate(
    year=table.date.year(),
    month=table.date.month(),
    quarter=table.date.quarter()
)

Window functions

# Rank within groups
ranked = table.mutate(
    rank=table.amount.rank().over(
        ibis.window(group_by="category", order_by=table.amount.desc())
    )
)

# Running total
with_cumsum = table.mutate(
    cumulative=table.amount.sum().over(
        ibis.window(order_by="date", rows=(None, 0))
    )
)

Best Practices

  1. Filter early: Apply filters before aggregations to reduce data volume
  2. Use lazy evaluation: Ibis operations don't execute until .execute() is called - chain operations before executing
  3. Handle nulls: Check for and handle null values explicitly
  4. Leverage selectors for column operations (see REFERENCE.md)

Detailed Resources

Scripts

Complete analysis workflow

Run the full example script:

python scripts/analyze.py

This demonstrates:

  • Data exploration and quality checks
  • Transformations and enrichment
  • Aggregations and filtering
  • Joins (if multiple files available)
  • Multiple export formats
  • Optional visualization

Quick-start template

Copy and customize the template:

cp scripts/template.py my_analysis.py
# Edit with your file paths and logic
python my_analysis.py

Installation

Install required packages:

uv add "ibis-framework[duckdb]"

# Optional for visualization
uv add matplotlib

Troubleshooting

Import error: Ensure ibis-framework[duckdb] is installed

Large files: DuckDB handles large parquet files efficiently. For very large datasets:

  • Filter early and aggressively
  • Use selective column reading: con.read_parquet("file.parquet", columns=["id", "amount"])
  • Process in chunks or use more selective queries

Schema mismatches in joins: Ensure column types match using .cast():

table = table.mutate(id=table.id.cast("int64"))

Performance: For complex queries, check the generated SQL with ibis.to_sql(table) to understand what's being executed

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

34.15%
按下载量换算0

Claude

32.12%
按下载量换算0

Cursor

17.94%
按下载量换算0

Gemini CLI

9.97%
按下载量换算0

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills