Token导航 LogoToken导航TokenDH.com
前端设计敏感数据github未标认证来源可访问clear审计提醒

ai-multimodalAI 多式联运

Agent Skill

ai-multimodal 用于处理 GitHub 仓库、Issue、Pull Request 和代码协作信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要围绕仓库状态、代码变更或协作事项进行整理时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

713

周安装

30

GitHub Stars

6

下载量

250
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

3

许可证

MIT

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:ai-multimodal(AI 多式联运)
来源仓库:https://github.com/duc01226/easyplatform
仓库路径:skills/ai-multimodal
安装命令:
npx skills add https://github.com/duc01226/easyplatform --skill ai-multimodal
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。不同来源提供的安装方式可能略有差异;本站展示可直接复制的安装命令,安装前请核对来源页面。

skills.shnpx skills
npx skills add https://github.com/duc01226/easyplatform --skill ai-multimodal

简介

ai-multimodal 融合文本、图像与语音等多种输入输出形式,实现跨模态智能处理。

  • 适用于 Codex、Claude、Cursor、Gemini CLI 中复杂感知任务的统一解决方案。
  • 采用任务拆解机制确保每个环节都有明确责任人(AI 或人类)的参与确认。
  • 严禁将推测当作事实呈现,所有结论必须具备可验证的来源支撑链条。
  • 适用宿主包括 Codex、Claude、Cursor、Gemini CLI,接入前应确认版本、权限和运行环境要求。

SKILL.md

[IMPORTANT] Use TaskCreate to break ALL work into small tasks BEFORE starting — including tasks for each file read. This prevents context loss from long files. For simple tasks, AI MUST ATTENTION ask user whether to skip.
Critical Thinking Mindset — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act. Anti-hallucination: Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.
AI Mistake Prevention — Failure modes to avoid on every task: - Check downstream references before deleting. Deleting components causes documentation and code staleness cascades. Map all referencing files before removal. - Verify AI-generated content against actual code. AI hallucinates APIs, class names, and method signatures. Always grep to confirm existence before documenting or referencing. - Trace full dependency chain after edits. Changing a definition misses downstream variables and consumers derived from it. Always trace the full chain. - Trace ALL code paths when verifying correctness. Confirming code exists is not confirming it executes. Always trace early exits, error branches, and conditional skips — not just happy path. - When debugging, ask "whose responsibility?" before fixing. Trace whether bug is in caller (wrong data) or callee (wrong handling). Fix at responsible layer — never patch symptom site. - Assume existing values are intentional — ask WHY before changing. Before changing any constant, limit, flag, or pattern: read comments, check git blame, examine surrounding code. - Verify ALL affected outputs, not just the first. Changes touching multiple stacks require verifying EVERY output. One green check is not all green checks. - Holistic-first debugging — resist nearest-attention trap. When investigating any failure, list EVERY precondition first (config, env vars, DB names, endpoints, DI registrations, data preconditions), then verify each against evidence before forming any code-layer hypothesis. - Surgical changes — apply the diff test. Bug fix: every changed line must trace directly to the bug. Don't restyle or improve adjacent code. Enhancement task: implement improvements AND announce them explicitly. - Surface ambiguity before coding — don't pick silently. If request has multiple interpretations, present each with effort estimate and ask. Never assume all-records, file-based, or more complex path.

Quick Summary

Goal: Process and generate multimedia content (images, audio, video, documents) using Google Gemini API via Python scripts.

Workflow:

  1. Identify Modality — Match input type to task (analyze, transcribe, extract, generate)
  2. Check Limits — Inline max 20MB, File API max 2GB; split large audio at 15min chunks
  3. Execute — Run gemini_batch_process.py with appropriate task and files
  4. Post-Process — Format output as markdown with timestamps, save generated content

Key Rules:

  • Requires GEMINI_API_KEY environment variable
  • Always request specific nodes/files, avoid full-file downloads
  • Use media_optimizer.py to compress/split files exceeding limits

Be skeptical. Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence percentages (Idea should be more than 80%).

AI Multimodal

Purpose

Process audio, images, videos, and documents or generate images/videos using Google Gemini's multimodal API via bundled Python scripts.

When to Use

  • Analyzing images or screenshots (Gemini vision is preferred over Claude's built-in vision for complex tasks)
  • Transcribing audio files (meetings, podcasts, interviews)
  • Extracting data from PDFs, scanned documents, or charts
  • Processing video content (scene detection, temporal Q&A)
  • Generating images with Imagen 4 or videos with Veo 3
  • Converting documents to markdown with visual understanding

When NOT to Use

  • Simple text-only LLM calls -- use Claude directly
  • Reading a file Claude can already read (code, markdown, JSON) -- use Read tool
  • Building AI-powered application features -- use api-design or frontend-design
  • Music composition workflows -- load references/music-generation.md only when specifically requested
  • General prompt engineering -- use ai-artist skill

Prerequisites

export GEMINI_API_KEY="your-key"  # From https://aistudio.google.com/apikey
pip install google-genai python-dotenv pillow
python scripts/check_setup.py  # Verify setup

Optional: API key rotation for rate limits (set GEMINI_API_KEY_2, GEMINI_API_KEY_3).

Workflow

Step 1: Identify Modality

Input TypeTaskCommand
Image (PNG/JPG/WEBP)Analyze, caption, OCR--task analyze
Audio (WAV/MP3/AAC)Transcribe, summarize--task transcribe
Video (MP4/MOV)Scene detection, Q&A--task analyze
PDF/DocumentExtract tables, forms--task extract
Text promptGenerate image--task generate
Text promptGenerate video--task generate-video

Step 2: Check Limits

  • Inline upload: max 20MB
  • File API: max 2GB (auto-used for large files)
  • Audio transcription: split at 15-minute chunks for full transcript
  • Video transcription: extract audio first, then split and transcribe
  • Formats: Audio (WAV/MP3/AAC, up to 9.5h), Images (PNG/JPEG/WEBP, up to 3.6k), Video (MP4/MOV, up to 6h), PDF (up to 1k pages)

IF file exceeds limits, use scripts/media_optimizer.py to compress/split first.

Step 3: Execute

Quick check: If gemini CLI is available, use: "<prompt>" | gemini -y -m gemini-2.5-flash

Standard: Use the batch processing script:

# Analyze media
python scripts/gemini_batch_process.py --files <file> --task <analyze|transcribe|extract>

# Generate content
python scripts/gemini_batch_process.py --task generate --prompt "description"
python scripts/gemini_batch_process.py --task generate-video --prompt "description"

Stdin support: cat image.png | python scripts/gemini_batch_process.py --task analyze --prompt "Describe this"

Step 4: Post-Processing

  • For transcripts: output in markdown with [HH:MM:SS -> HH:MM:SS] timestamps
  • For document extraction: save as structured markdown under docs/assets/
  • For generated images/videos: save to working directory with descriptive filename

Step 5: Verification

  • Confirm output matches expected format and completeness
  • For long transcripts: verify no truncation occurred (check chunk boundaries)
  • For generated content: verify quality meets prompt requirements

Models

PurposeModelNotes
Analysis (fast)gemini-2.5-flashRecommended default
Analysis (advanced)gemini-2.5-proComplex reasoning tasks
Image generationimagen-4.0-generate-001Standard quality
Image generation (quality)imagen-4.0-ultra-generate-001Best quality
Image generation (speed)imagen-4.0-fast-generate-001Fastest
Video generationveo-3.1-generate-preview8s clips with audio

Scripts Reference

  • gemini_batch_process.py -- CLI orchestrator for all tasks, auto-resolves API keys and models
  • media_optimizer.py -- Compress/resize/split media to fit Gemini limits
  • document_converter.py -- Convert PDFs/images/Office docs to markdown
  • check_setup.py -- Verify environment, dependencies, and API key

Use --help on any script for full options.

Examples

Example 1: Transcribe a Meeting Recording

Input: 45-minute meeting audio file meeting-2025-01-15.mp3

Steps:

  1. File is >15min, so split first: python scripts/media_optimizer.py --input meeting-2025-01-15.mp3 --split-duration 900
  2. Transcribe each chunk: python scripts/gemini_batch_process.py --files meeting-part-*.mp3 --task transcribe
  3. Output: Markdown file with timestamps, speaker detection, and metadata (duration, topics covered)

Example 2: Extract Data from a PDF Report

Input: Quarterly HR report PDF with tables, charts, and forms

Steps:

  1. Convert and extract: python scripts/document_converter.py --input quarterly-report.pdf --output docs/assets/
  2. Output: Structured markdown with tables preserved, chart descriptions, and form field values extracted

Detailed References

Load for in-depth guidance:

TopicFile
Audio processingreferences/audio-processing.md
Vision/image analysisreferences/vision-understanding.md
Image generationreferences/image-generation.md
Video analysisreferences/video-analysis.md
Video generationreferences/video-generation.md
Music generationreferences/music-generation.md

Related Skills

  • ai-artist -- for prompt engineering and optimization (not media processing)
  • media-processing -- for FFmpeg-based audio/video encoding without AI
  • pdf-to-markdown -- for simple PDF text extraction without vision AI

Closing Reminders

  • MANDATORY IMPORTANT MUST ATTENTION break work into small todo tasks using TaskCreate BEFORE starting
  • MANDATORY IMPORTANT MUST ATTENTION search codebase for 3+ similar patterns before creating new code
  • MANDATORY IMPORTANT MUST ATTENTION cite file:line evidence for every claim (confidence >80% to act)
  • MANDATORY IMPORTANT MUST ATTENTION add a final review todo task to verify work quality
  • MUST ATTENTION apply critical thinking — every claim needs traced proof, confidence >80% to act. Anti-hallucination: never present guess as fact.
  • MUST ATTENTION apply AI mistake prevention — holistic-first debugging, fix at responsible layer, surface ambiguity before coding, re-read files after compaction.

[TASK-PLANNING] Before acting, analyze task scope and systematically break it into small todo tasks and sub-tasks using TaskCreate.

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

04

需要参考平台分布和安装热度时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Claude Code

28.17%
按下载量换算70

windsurf

25.05%
按下载量换算63

OpenCode

17.49%
按下载量换算44

Codex

12.37%
按下载量换算31

Antigravity

8.21%
按下载量换算21

Gemini CLI

3.4%
按下载量换算9

安全审计

Gen Agent Trust Hub

可疑

Socket

通过

Snyk

可疑

权限和风险

敏感数据

该 Skill 可能接触密钥、Token、环境变量或敏感配置,应进入高风险复核队列,默认不自动发布。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。

来源信息

继续浏览同类 Skills