Token导航 LogoToken导航TokenDH.com
效率需要联网clawhub未标认证来源可访问clear审计通过

midasheng-audio-generate米达声音频生成

Agent Skill

用于辅助音频、音乐、语音转写、语音合成或声音素材处理。它适合让 Agent 生成配乐说明、整理音频流程、调用语音工具或处理播客和视频配音素材。使用时需要确认输入音频来源、输出格式、时长和模型限制;涉及人声克隆、版权音乐或公开发布时,应先核对授权和合规边界。

总安装

7,128

周安装

297

GitHub Stars

3

下载量

2,376
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:midasheng-audio-generate(米达声音频生成)
来源仓库:https://github.com/jimbozhang/midasheng-audio-generate
安装命令:
openclaw skills install midasheng-audio-generate
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install midasheng-audio-generate

简介

通过文本描述生成身临其境的音频场景,包括语音、音效、音乐和环境声音。

  • 适用于音频、音乐、语音转写、语音合成或声音素材处理等场景。
  • 使用 openclaw skills install midasheng-audio-generate 命令安装。
  • 使用时需确认输入音频来源、输出格式、时长和模型限制;涉及人声克隆、版权音乐或公开发布时,应先核对授权和合规边界。
  • 适用宿主包括 OpenClaw,接入前应确认版本、权限和运行环境要求。

SKILL.md

name
midasheng-audio-generate
description
Developed by Xiaomi and Shanghai Jiao Tong University. Transform text into high‑quality audio scenes with speech, SFX, music, and ambiance. Demo: https://nieeim.github.io/Dasheng-AudioGen-Web/
endpoint
purpose
Primary audio generation interface
purpose
Query service queue status
authentication
none
privacy
data_sent
User-provided text prompts and any agent-generated enhancements
data_retained
Unknown (dependent on third-party provider's policy)
recommendation
>
links
url
https://nieeim.github.io/Dasheng-AudioGen-Web/
url
https://huggingface.co/spaces/mispeech/Dasheng-AudioGen
url
https://github.com/NieeiM/Dasheng-Audiogen
requirements

midasheng-audio-generate

Audio scene generation from text descriptions. Generates WAV audio with speech, sound effects, music, and environmental sounds.

1. Trigger

Use this skill when the user requests audio, sound effects, or music generation based on a text description.

2. Execution Steps

Step 1: Design the Audio Scene (Prompt Refinement)

Before calling the API, you must act as an expert Audio Scene Architect and Foley Designer. Deeply understand the user's natural language input (which may be in any language) and translate it into a highly structured tagged string based on real-world acoustic logic and scene realism.

Prompt Tag Definition:

  • <|caption|>: The overall, comprehensive description of the audio scene.
  • <|speech|>: Speaker identity (e.g., middle-aged man, energetic girl) and speaking style.
  • <|asr|>: The actual transcript / spoken dialogue.
  • <|sfx|>: Specific sound effects present in the audio (e.g., footsteps, doorbell, dog barking).
  • <|music|>: Description of background music (e.g., soft jazz, tense orchestral).
  • <|env|>: Environmental or ambient background noise (e.g., city bustle, forest wind and crickets).

Crucial Generation Rules:

  1. Scene Enrichment: Do not merely copy the user's input! Act as a sound designer and logically enrich the scene.
  2. Speech & Dialogue Generation: If the user explicitly mentions speech or implies a speaking scenario, creatively generate a reasonable and vivid transcript for the <|speech|> and <|asr|> fields.
  3. Strict ASR Formatting: For the <|asr|> tag, output only the raw spoken text. Do not include any speaker labels or narration, such as “man:”, “speaker1:”, or “a man says”.
  4. Omit Missing Elements: If any element is not relevant, directly omit its corresponding tag.
  5. Language & Case Constraint: The entire generated prompt string MUST be in lowercase English, including <|asr|> content.
  6. Strict Output: Output ONLY the formatted tagged string internally for the next step.

Step 2: Execute Command

curl -X POST "https://llmplus.ai.xiaomi.com/dasheng/audio/gen" \
  -H "Content-Type: application/json" \
  -d "{\"text\": \"<FORMATTED_PROMPT_STRING>\"}" \
  -o <FILENAME.wav>

3. Queue Status

Query Command

curl -X POST "https://llmplus.ai.xiaomi.com/metrics?path=/dasheng/audio/gen"

Returned Fields

  • active: Number of currently active requests
  • avg_latency_ms: Average processing latency (milliseconds)
  • Estimated wait time = active × avg_latency_ms

When to Call

  1. When the IM is about to timeout but the audiogen service has not returned a result: Check the queue status and inform the user, asking them to inquire again later.
  2. When the user asks about task progress later but the service still hasn't returned: Check the latest queue status and report it back to the user.

Status Levels

  • 🟢 active=0 or estimated wait <5s → Service idle
  • 🟡 Estimated wait 5-30s → Slight queue
  • 🔴 Estimated wait >30s → Queue is long, recommend trying again later

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

89.15%
按下载量换算2,118

安全审计

VirusTotal

通过

ClawScan

通过

Static analysis

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills