Token导航 LogoToken导航TokenDH.com
研究检索敏感数据clawhub未标认证来源可访问clear审计提醒

audioclaw-skills-voice-reply音频爪技能语音回复

Agent Skill

用于辅助音频、音乐、语音转写、语音合成或声音素材处理。它适合让 Agent 生成配乐说明、整理音频流程、调用语音工具或处理播客和视频配音素材。使用时需要确认输入音频来源、输出格式、时长和模型限制;涉及人声克隆、版权音乐或公开发布时,应先核对授权和合规边界。

总安装

6,937

周安装

298

GitHub Stars

公开资料未说明

下载量

2,432
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:audioclaw-skills-voice-reply(音频爪技能语音回复)
来源仓库:https://github.com/kikidouloveme79/audioclaw-skills-voice-reply
安装命令:
openclaw skills install audioclaw-skills-voice-reply
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install audioclaw-skills-voice-reply

简介

发送带有可切换 voice_id、情绪预设或说话风格的语音回复。

  • 适用于飞书、Lark 等平台集成语音交互场景。
  • 支持运行时动态调整音色与表达风格以满足不同需求。
  • 需配置正确的 voice_id 参数与平台 API 对接。
  • 建议在使用前验证目标平台对语音格式的兼容性。audioclaw-skills-voice-reply 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

name
audioclaw-skills-voice-reply
description
Use when AudioClaw Skills, Feishu, or Lark needs to send AudioClaw voice replies with runtime-switchable voice_id, emotion preset, or speaking style, including per-message speaker overrides, voice-family emotion routing, cache reuse, and safe fallback when a requested voice is unavailable.

AudioClaw Skills Voice Reply

When to use

Use this skill when AudioClaw already has the final reply text and now needs a voice version that can change on demand.

Common triggers:

  • A chat bot should answer in different tones such as calm, warm, cheerful, serious, or promo.
  • The caller wants to switch voice_id or voice_family in a single request without editing code.
  • The same AudioClaw workflow should support both free voices and paid or custom voices.
  • The runtime should keep working even when a requested voice is not available to the current key.
  • A user has already said "以后一直给我发语音", and AudioClaw should remember that preference for future turns.
  • AudioClaw needs a workspace-local file plus a stable way to deliver it as a Feishu audio message.
  • The caller already has a cloned AudioClaw voice_id such as vc-... and wants the runtime to use it directly without falling back first.

Do not use this skill for ASR intake or long-form digest generation.

Workflow

  1. Start from the final user-facing text. Do not pass hidden reasoning or raw markdown tables.
  2. Build an AudioClaw voice request with:

- text - optional scene - optional voice_id - optional voice_family - optional emotion - optional speed, pitch, volume

  1. Run scripts/openclaw_voice_switchboard.py.

- For AudioClaw on Feishu or Lark, prefer scripts/picoclaw_voice_reply.py.

  1. Let the script resolve the request in this order:

- exact voice_id - voice_family plus matching emotion variant - scene plus emotion preset - validated fallback voice - Important: if the exact voice_id is a clone-style id such as vc-..., this skill now tries that id directly first, even if it is not part of the built-in official voice catalog.

  1. If preference_key is present, the script can remember:

- reply_mode - default voice_id - default emotion - default scene

  1. If the result should be sent through AudioClaw media upload, pass:

- --out inside the AudioClaw workspace - --openclaw-workspace-root pointing at the workspace root - --delivery-profile feishu_voice when the downstream channel prefers .ogg/.opus - optional --chmod 644 if you want to be explicit, though this skill now defaults to 0644 - if --openclaw-workspace-root is set and --out is omitted, this skill now writes to workspace/state/audio/ automatically

  1. Use the returned JSON manifest in AudioClaw to:

- prefer scripts/picoclaw_voice_reply.py for AudioClaw on Feishu - let the wrapper send the Feishu audio message directly - do not send the local path or the MEDIA:... line through the message tool - only use send_file when you intentionally opt out of direct Feishu sending - log trace_id - persist the resolved voice choice for the next turn if desired

  1. If the requested voice is not available to the current key, let the fallback stand unless the user explicitly requires strict failure.
  2. If you want AudioClaw to remember a clone voice, either:

- set it as the user's default voice with --set-default-voice-id vc-... - or register it explicitly with --register-clone-voice-id vc-...

Voice discovery

When AudioClaw needs to find a usable voice or confirm a voice_id, use this order:

  1. Check the local catalog first:

- run python3 scripts/openclaw_voice_switchboard.py --list-voices - use this for fast lookup of built-in voices, known emotion variants, and clone voices already registered locally

  1. If the user asks for the official public voice list, package availability, or a voice not found locally, check the official voice page:

- https://senseaudio.cn/docs/voice_api - page title: API 音色服务说明

  1. Prefer the official page when AudioClaw needs to confirm:

- whether a voice_id is a free, VIP, or SVIP voice - whether a named speaker has multiple emotion variants - whether the request likely needs selected-voice purchase or custom voice authorization

  1. After finding a likely voice_id, still let the runtime validate access at synthesis time, because account permissions can differ by key.

Practical rule:

  • Local --list-voices is the first-stop runtime catalog.
  • https://senseaudio.cn/docs/voice_api is the canonical fallback reference for official voice names, voice_id, and package tier notes.

AudioClaw rule

When this skill is used inside AudioClaw for Feishu or Lark voice replies:

  1. Run scripts/picoclaw_voice_reply.py.
  2. Let the wrapper upload the generated .ogg/.opus file to Feishu and send it as msg_type=audio.
  3. Do not call the send_file tool for that audio unless you explicitly passed --skip-direct-send.
  4. Do not call the message tool with the local path or the MEDIA:... reference. AudioClaw will send them as plain text.
  5. After the audio is sent, prefer no extra text confirmation.
  6. If the host runtime still requires one final assistant message to finish the turn, send one short natural Chinese line such as 我已经用语音回复你了。 instead of leaving the turn empty.
  7. Use media_reference only as debug metadata or future AudioClaw compatibility data.

This rule matters because this AudioClaw environment does not render MEDIA:... as media, and the generic send_file tool sends Feishu voice notes as plain files instead of audio messages. The reliable path here is direct Feishu upload plus msg_type=audio.

Runtime model

The official public TTS API exposes:

  • voice_setting.voice_id
  • voice_setting.speed
  • voice_setting.vol
  • voice_setting.pitch
  • audio_setting.format
  • audio_setting.sample_rate
  • one HTTP endpoint with two modes:

- non-stream with stream=false - SSE with stream=true

Important constraint:

  • The public TTS API docs do not expose a standalone emotion request field.
  • Emotion switching is therefore handled by choosing a matching voice_id when one exists, or by keeping the voice and shaping speed / pitch / vol.
  • This skill requests final-file TTS in non-stream mode by default, because AudioClaw only needs the completed file and this avoids stream assembly edge cases.
  • For this server-side HTTP TTS path, the official docs still use Authorization: Bearer API_KEY. The generated Public Key is not required by this skill.
  • If the requested voice_id looks like a clone id such as vc-..., this skill now auto-routes TTS to SenseAudio-TTS-1.5 and records audio.model_used in the manifest.

API key lookup

This skill now treats SENSEAUDIO_API_KEY as the default API key source again.

Runtime rules:

  • If the host app injects SENSEAUDIO_API_KEY as an AudioClaw login token such as v2.public..., the shared bootstrap will replace it with the real sk-... value from ~/.audioclaw/workspace/state/senseaudio_credentials.json before TTS starts.
  • --api-key-env still works, but the default runtime path is SENSEAUDIO_API_KEY.

If you need the exact same speaker timbre across many emotions, use a purchased multi-variant voice family or an authorized custom voice. Otherwise this skill will approximate the requested emotion with the best available voice and tuning.

Request contract

Minimal JSON request:

{
  "text": "我们已经收到你的需求,今天下午会把结果发给你。",
  "scene": "assistant",
  "emotion": "calm"
}

Full request:

{
  "text": "新品今晚八点开售,现在下单还有首发赠品。",
  "scene": "sales",
  "voice_id": "male_0027_b",
  "voice_family": "male_0027",
  "emotion": "promo",
  "speed": 1.08,
  "pitch": 1,
  "volume": 1.05,
  "audio_format": "mp3",
  "sample_rate": 32000,
  "preference_key": "feishu:ou_xxx",
  "reply_mode": "voice",
  "allow_fallback": true,
  "strict_voice": false,
  "cache_dir": "/tmp/openclaw-voice-cache"
}

Clone voice example:

{
  "text": "这是你的克隆音色回复测试。",
  "voice_id": "vc-yxdCFUKyNLPexxJ66jaXWk",
  "emotion": "calm",
  "allow_fallback": false,
  "strict_voice": true
}

Supported emotion presets:

  • neutral
  • calm
  • warm
  • cheerful
  • serious
  • promo
  • sad
  • angry
  • analytical

Supported scene hints:

  • assistant
  • customer_support
  • briefing
  • sales
  • marketing
  • narration
  • education
  • gaming
  • warning

AudioClaw integration pattern

Recommended handoff:

  1. AudioClaw generates final reply text.
  2. AudioClaw decides whether this turn should speak and what mood it wants.
  3. AudioClaw calls scripts/openclaw_voice_switchboard.py with a request JSON.
  4. The script returns a manifest with:

- requested voice - resolved voice - emotion strategy - effective speed / pitch / volume - delivery profile and file mode - local audio path - AudioClaw-friendly media reference when the output is under a workspace root - trace_id

  1. AudioClaw uploads the resulting file to Feishu or any downstream channel, or lets the AudioClaw wrapper do that directly for Feishu.

Operational rules:

  • Cache by resolved voice plus text plus audio settings.
  • If a paid voice is unavailable, allow fallback unless the request is marked strict.
  • Always log the resolved voice and trace_id.
  • Force generated audio files to mode 0644 so the AudioClaw sender can read them reliably.
  • When --openclaw-workspace-root is set and --out stays inside that root, expose delivery.openclaw_media_reference.
  • When --delivery-profile feishu_voice is enabled, synthesize with AudioClaw first and then transcode to ogg/opus with system ffmpeg or imageio-ffmpeg.
  • This publishable skill intentionally does not bundle ffmpeg. Install ffmpeg or run python3 -m pip install imageio-ffmpeg on the target machine.
  • Avoid ad-hoc temp filenames under the workspace root. Prefer workspace/state/audio/, which this skill will now use automatically when --openclaw-workspace-root is given without --out.
  • For AudioClaw on Feishu, scripts/picoclaw_voice_reply.py now uses the local Feishu app credentials from ~/.audioclaw/config.json, uploads the audio through the official /open-apis/im/v1/files endpoint, and sends it as msg_type=audio.
  • The wrapper infers the active Feishu chat_id from the latest agent_main_feishu_direct_*.jsonl session log unless you pass --chat-id explicitly.

Commands

List voices:

python3 scripts/openclaw_voice_switchboard.py --list-voices

Check the official voice catalog page:

https://senseaudio.cn/docs/voice_api

List emotion presets:

python3 scripts/openclaw_voice_switchboard.py --list-emotions

Enable permanent voice reply for one user:

python3 scripts/openclaw_voice_switchboard.py \
  --preference-key "feishu:ou_xxx" \
  --set-reply-mode voice \
  --set-default-voice-id male_0004_a \
  --set-default-emotion calm \
  --set-default-scene assistant

Enable permanent cloned-voice reply for one user:

python3 scripts/openclaw_voice_switchboard.py \
  --preference-key "feishu:ou_xxx" \
  --set-reply-mode voice \
  --set-default-voice-id vc-yxdCFUKyNLPexxJ66jaXWk \
  --set-default-emotion calm \
  --set-default-scene assistant

Register a prepared clone voice so the runtime can list and reuse it:

python3 scripts/openclaw_voice_switchboard.py \
  --register-clone-voice-id vc-yxdCFUKyNLPexxJ66jaXWk \
  --register-clone-name "我的克隆音色"

Show registered clone voices:

python3 scripts/openclaw_voice_switchboard.py --show-clone-voices

Show current voice preference:

python3 scripts/openclaw_voice_switchboard.py \
  --preference-key "feishu:ou_xxx" \
  --show-preference

Generate one AudioClaw turn from a JSON request:

python3 scripts/openclaw_voice_switchboard.py \
  --request-file /path/to/request.json \
  --out /tmp/openclaw_reply.mp3

Direct CLI example:

python3 scripts/openclaw_voice_switchboard.py \
  --text "我们已经处理好了,稍后把结果发给你。" \
  --scene assistant \
  --emotion warm \
  --preference-key "feishu:ou_xxx" \
  --delivery-profile feishu_voice \
  --openclaw-workspace-root ~/.audioclaw/workspace \
  --out ~/.audioclaw/workspace/audioclaw_warm.ogg

AudioClaw Feishu one-step example:

python3 scripts/picoclaw_voice_reply.py \
  --text "这是一次可以直接发给飞书的语音回复。" \
  --scene assistant \
  --emotion calm \
  --workspace-root ~/.audioclaw/workspace

Only generate, do not send:

python3 scripts/picoclaw_voice_reply.py \
  --text "这是一次只生成不直发的测试语音。" \
  --scene assistant \
  --emotion calm \
  --workspace-root ~/.audioclaw/workspace \
  --skip-direct-send

Resources

  • scripts/senseaudio_tts_client.py

- Small importable client for https://api.senseaudio.cn/v1/t2a_v2 - Handles SSE chunks and writes audio bytes

  • references/openclaw_voice_switchboard.md

- TTS capability summary plus the official voice catalog reference at https://senseaudio.cn/docs/voice_api

  • scripts/openclaw_voice_switchboard.py

- Main runtime for AudioClaw - Resolves voice, emotion, fallback, caching, and output manifest

  • scripts/picoclaw_voice_reply.py

- AudioClaw-first wrapper - Generates audio and, by default, sends it to Feishu as a real audio message

  • scripts/feishu_audio_sender.py

- Direct Feishu sender for .ogg/.opus - Uses ~/.audioclaw/config.json app credentials by default, infers the active chat, uploads the file, and sends msg_type=audio

  • references/openclaw_voice_switchboard.md

- Official docs summary, request examples, and AudioClaw capability boundaries

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

93.01%
按下载量换算2,262

安全审计

VirusTotal

可疑

ClawScan

可疑

Static analysis

通过

权限和风险

敏感数据

该 Skill 可能接触密钥、Token、环境变量或敏感配置,应进入高风险复核队列,默认不自动发布。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills