Token导航 LogoToken导航TokenDH.com
研究检索需要联网github未标认证来源可访问许可证需确认审计提醒

dataset-transformation数据集转换

Agent Skill

用于辅助数据整理、表格处理、CSV/Excel 分析、指标计算和图表准备。它适合让 Agent 清洗字段、汇总数据、发现异常、生成统计口径或把分析结果转成可读说明。使用时需要确认数据来源、字段含义和时间范围,避免把样本数据当全量事实;涉及敏感数据、导出文件或批量写回时,应先确认权限和脱敏边界。

总安装

955

周安装

39

GitHub Stars

634

下载量

306
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:dataset-transformation(数据集转换)
来源仓库:https://github.com/awslabs/agent-plugins
仓库路径:skills/dataset-transformation
安装命令:
npx skills add https://github.com/awslabs/agent-plugins --skill dataset-transformation
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/awslabs/agent-plugins --skill dataset-transformation

简介

用于将用户提供数据集转换为指定格式,输出 Jupyter notebook 形式的转换代码。

  • 适合 SageMaker 模型训练或评估前的数据处理、清洗和格式化任务。
  • 每次只推进一个决策,生成生产就绪代码,支持正式审核流程。
  • 需确认数据集来源、字段含义和时间范围,避免误用样本代表整体。
  • 涉及敏感数据或批量操作时,应先验证权限与脱敏规则。

SKILL.md

Dataset Transformation Agent

Transforms a data set provided by the user into their desired format. All transformation code is delivered as a Jupyter notebook.

When to Use

  • User needs to generate code for transforming datasets for SageMaker model training or model evaluation.
  • A dataset requires processing, cleaning, or formatting before training or evaluation.
  • Workflow requires a formal review and approval cycle before execution.

Principles

  1. One thing at a time. Each response advances exactly one decision. Never combine multiple questions or recommendations in a single turn.
  2. Confirm before proceeding. Wait for the user to agree before moving to the next step. You are a guide, not a runaway train.
  3. Don't read files until you need them. Only read reference files when you've reached the workflow step that requires them and the user has confirmed the direction. Never read ahead.
  4. No narration. Don't explain what you're about to do or what you just did. Share outcomes and ask questions. Keep responses short and focused.
  5. No repetition. If you said something before a tool call, don't repeat it after. Only share new information.
  6. Do not deviate from the Workflow. The steps listed in the workflow should be followed exactly as described. Progress from Step 1 to Step 11 to complete the task. Do not deviate from the workflow!
  7. Always end with a question. Whenever you pause for user input, acknowledgment, or feedback, your response must end with a question. Never leave the user with a statement and expect them to know they need to respond.
  8. Default output format is JSONL. Unless the user explicitly requests a different file format, the transformed dataset should be written as .jsonl (JSON Lines — one JSON object per line).

Known Dataset Formats Reference

This skill supports two transformation purposes — training data and evaluation data — each with its own format resolution path. The purpose is determined in Step 1 of the workflow.

Training Data Formats

When the transformation is for model training, resolve the target format using the reference file ../dataset-evaluation/references/strategy_data_requirements.md. The required format depends on both the model type (Open Weights like Llama/Qwen vs Nova) and the finetuning technique (SFT, DPO, RLVR) — make sure to match on both dimensions. If either the model type or technique is not yet known, ask the user before resolving the format.

Evaluation Data Formats

When the transformation is for model evaluation, resolve the target format using this order:

  1. Try fetching the live documentation at https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-evaluation-dataset-formats.html to get the latest evaluation dataset schema definitions.
  2. If the fetch fails (e.g., no internet access, VPC environment), fall back to the offline copy at references/sagemaker_dataset_formats.md. Inform the user that the format schemas are from an offline copy and may be outdated.

Use whichever source you successfully access as the source of truth for the target format. Do not rely on memorized schemas.

Workflow

Step 1: Determine transformation purpose

Your first response should determine whether this transformation is for model training or model evaluation. If the context already makes this clear (e.g., the user said "I need to prep my training data" or "I need to format my eval dataset"), confirm your understanding and move on. Otherwise, ask:

"Is this dataset transformation for model training or model evaluation? This helps me look up the right target format for you."
  • Training → format resolution will use the local training data requirements reference (model type + finetuning technique dependent).
  • Evaluation → format resolution will use the live AWS documentation (with offline fallback).

Remember this choice — it determines how the target format is resolved in Step 3.

⏸ Wait for user.

Step 2: Set expectations

Acknowledge the user's request and state what this skill can do:

"I can help you transform your dataset's format! Here's my plan: I will first need to understand the format of your dataset and the transformation requirements. Once I have that, I will generate a dataset transformation function that we can refine together. After the dataset transformation function is refined to your liking, I will perform the transformation task and upload it to your desired location! Does this sound good?"

⏸ Wait for user.

Step 3: Understand the dataset transformation task

For this step, you need to know: what dataset format the user would like to transform their dataset from and what dataset format they would like to transform it in to. If you know this already, skip this step. If not, ask the user:

"What's the dataset format you would like to transform it into?"

Resolve the target format based on the purpose determined in Step 1:

  • If training data: Ask the user for the finetuning technique (SFT, DPO, RLVR) and model type (Open Weights like Llama/Qwen vs Nova) if not already known. Then look up the required format from the "Training Data Formats" section in the Known Dataset Formats Reference above.
  • If evaluation data: If the user mentions a well-known format name (e.g., "OpenAI format", "SageMaker format"), fetch the schema from the live documentation as described in the "Evaluation Data Formats" section above. If a well-known format is fetched, confirm with the user:
"I've found a SageMaker dataset format: {sagemaker-dataset-format-name} with schema: {sagemaker-dataset-format-schema}. Is this what you were referring to?"

If the user describes a custom format not listed in the reference doc, ask them to provide a sample record of the desired output format.

⏸ Wait for user.

Step 4: Get the dataset from the user

For this step, you need: the location of the user's dataset. If you know this already, skip this step. If not, ask the user:

"Where can I find your dataset? Either a local directory or S3 location works!"

⏸ Wait for user.

Step 5: Examine sample data

Read 1–2 sample records from the user's dataset and show them so the user can confirm the source schema. Do not run format detection — that is handled by the planning skill before this skill is invoked.

Do not show a side-by-side mapping to the target format here — the detailed mapping will be handled in Step 7 when generating the transformation function.

⏸ Wait for user.

Step 6: Get the dataset output location

For this step, you need: to understand where to output the transformed dataset to. It could be an S3 URI or local directory If you already know where the dataset is supposed to be output to, skip this step. If not, ask the user:

"Where should I output your transformed dataset to? Either a local directory or S3 location works!"

If the user provides a directory (not a full file path), construct the output filename using the pattern {original_name}_{target_format}.jsonl (e.g., gen_qa_100k_openai.jsonl).

⏸ Wait for user.

Step 7: Generate and validate the transformation function

For this step, you need: to generate a python function that transforms the dataset from the format in Step 5 to the format in Step 3

Read the reference guide at references/dataset_transformation_code.md and follow its skeleton exactly when generating the transformation function.

The python function should be in the form of:

def transform_dataset(df: pd.DataFrame) -> pd.DataFrame:

Add a %%writefile <project-dir>/scripts/transform_fn.py code cell to the notebook AND write the file to disk for testing. The <project-dir> is the project directory established by the directory-management skill (e.g., dpo-to-rlvr-conversion). All notebooks go in <project-dir>/notebooks/ and all scripts go in <project-dir>/scripts/.

Continue iterating with the user's feedback — update the notebook cell in place on each revision rather than showing code inline.

If sample data was collected in Step 5, test the function against the sample records:

  1. Generate the transformation function.
  2. Write the sample data to a temporary JSONL file (e.g., /tmp/test_input.jsonl), then run: python3 -c "import sys; sys.path.insert(0, '<project-dir>/scripts'); from transform_fn import transform_dataset; import pandas as pd; df = pd.read_json('/tmp/test_input.jsonl', lines=True); result = transform_dataset(df); print(result.to_json(orient='records', lines=True))"
  3. If the test fails, fix and re-test until it passes.
  4. Show the user the function and transformed sample output for review.

If no sample data, present the function for review and refinement.

⏸ Wait for user.

Step 8: Determine notebook target

Check if the project notebook already exists at <project-dir>/notebooks/<project-name>.ipynb.

  • If it exists → ask: *"Would you like me to append the transformation cells to the existing notebook, or create a new one?"*
  • If it doesn't exist → create it

When appending, add a markdown header cell ## Dataset Transformation as a section divider before the new cells.

⏸ Wait for user.

Step 9: Generate the execution cells in the notebook

Before writing the notebook, read:

  • references/notebook_structure.md (cell order, placeholders, and content)
  • references/notebook_writing_guide.md (Jupyter notebook JSON formatting)

Generate the execution logic as code cells in the notebook.

  • Add a %%writefile <project-dir>/scripts/<script_name>.py code cell to the notebook AND write the file to disk for testing.
  • The script must import transform_dataset from transform_fn.
  • Replace placeholders with the actual input/output paths.

Read the reference guide at references/dataset_transformation_code.md and follow its execution script skeleton exactly.

If sample data was collected in Step 5, test the full pipeline:

  1. Write the sample records to a temporary JSONL file (e.g., /tmp/test_input.jsonl).
  2. Run: python3 <project-dir>/scripts/<script_name> --input /tmp/test_input.jsonl --output /tmp/test_output.jsonl
  3. If it fails, debug and fix, then re-run until successful.
  4. Show the user the output for review.

If no sample data, present the notebook for review and refinement.

⏸ Wait for user.

Step 10: Determine and confirm execution mode

Check the size of the input dataset:

  • If the dataset is in S3, use the AWS MCP tool head-object (S3 service) with the bucket and key to get ContentLength.
  • If the dataset is local, check the file size.

Decision criteria:

  • Dataset < 50 MB → recommend local execution
  • Dataset ≥ 50 MB → recommend SageMaker Processing Job

Inform the user of the recommendation and get their approval:

If local:

"Your dataset is {size} MB — since it's under 50 MB, I'd recommend running the transformation locally. Would you like to proceed with local execution, or would you prefer a SageMaker Processing Job instead?"

If SageMaker Processing Job:

"Your dataset is {size} MB — since it's over 50 MB, I'd recommend running this as a SageMaker Processing Job for better performance. Would you like to proceed with a SageMaker Processing Job, or would you prefer to run it locally instead?"

Do not execute until the user approves. If the user rejects the recommendation, switch to the alternative and get their explicit approval before proceeding.

⏸ Wait for user.

After user confirms, add an execution cell to the notebook. Do NOT run the full transformation — only generate the cell for the user to execute themselves:

If local execution:

  • Add a cell that runs the transformation by importing from the .py files already on disk (written by the agent during Steps 7 and 9): import transform_dataset from transform_fn, load the dataset, transform, and save output. Scripts are located in <project-dir>/scripts/.

If SageMaker Processing Job:

  • Add a cell that submits and monitors the Processing Job inline using the V3 SageMaker SDK directly (FrameworkProcessor, ProcessingInput, ProcessingOutput, etc.). Create a FrameworkProcessor with the SKLearn 1.2-1 image, configure inputs/outputs, and call processor.run(wait=True, logs=True) to block the cell and stream logs until the job completes. See scripts/transformation_tools.py for reference implementation details.
  • Inform the user they can run this cell to kick off and monitor the job.

Important: The agent must NOT execute the full dataset transformation itself. The notebook cells are generated for the user to review and run. Only sample data (from Steps 7 and 9) should be transformed by the agent for validation purposes.

"I've added the execution cell to the notebook. You can run it to transform the full dataset. Would you like to review the notebook before running it?"

⏸ Wait for user.

Step 11: Verify and confirm with the user

For this step, you need: to verify the output looks correct and confirm with the user.

  • Read 1–2 sample records from the output to show the user.
  • Report the total number of records transformed.
  • Ask the user if the output looks good.

⏸ Wait for user to confirm.

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

35.07%
按下载量换算107

Claude

32.6%
按下载量换算100

Cursor

19.77%
按下载量换算60

Gemini CLI

9.25%
按下载量换算28

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

可疑

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills