Token导航 LogoToken导航TokenDH.com
开发执行命令github未标认证来源可访问许可证需确认审计异常

voice-mode语音模式

Agent Skill

用于辅助音频、音乐、语音转写、语音合成或声音素材处理。它适合让 Agent 生成配乐说明、整理音频流程、调用语音工具或处理播客和视频配音素材。使用时需要确认输入音频来源、输出格式、时长和模型限制;涉及人声克隆、版权音乐或公开发布时,应先核对授权和合规边界。

总安装

318

周安装

13

GitHub Stars

1

下载量

102
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:voice-mode(语音模式)
来源仓库:https://github.com/llblab/skills
仓库路径:skills/voice-mode
安装命令:
npx skills add https://github.com/llblab/skills --skill voice-mode
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/llblab/skills --skill voice-mode

简介

提供灵活的语音处理模式切换,适应不同任务类型的参数配置。

  • 可用于实时语音分析、批量转写或交互式语音调试场景。
  • 由 LBLab 开源项目维护,支持主流 AI 代码编辑器集成。
  • 建议根据输入音频长度与质量动态选择最优处理模式。
  • voice-mode 属于开发类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

Voice Mode (Super-Skill)

Purpose

This skill unifies voice output and voice input in one place:

  • say — text-to-speech (TTS)
  • listen — speech-to-text (STT)
  • duplex mode — agent orchestration (saylisten) built on the atomic scripts

Use say and listen independently, or let the agent combine them into continuous duplex dialogue. This is an offline-first skill: STT runs locally via faster-whisper, and TTS uses local piper models after the initial voice download.

Atomic Commands

1) Speak

say "text to announce"
say --lang ru "<text in Russian>"
# short alias is also supported:
say -l ru "<text in Russian>"

2) Listen

listen

3) Duplex mode (agent orchestration)

say --lang ru "<spoken reply in the conversation language>"
listen -l ru -d 0 -s 1

Duplex mode is not a standalone shell script in this skill. Core protocol remains atomic: say then listen. In duplex sessions, prefer listen -d 0 -s 1: no hard timeout, stop by user pause.

Operating Modes

Mode A: Selective Voice (default)

  • Use say only for short, high-value moments (greeting, warning, key conclusion).
  • Keep code, tables, and long technical details in text.

Mode B: Full Voice Output (screenless)

When explicitly requested by the user:

  1. Use say for every response.
  2. Speak the entire assistant reply through say, not just a short follow-up question.
  3. Do not duplicate full spoken content in chat.
  4. For code/tables: describe briefly by voice (language, purpose, size), avoid reading raw code line by line.

Mode C: Voice Input On-Demand

  • Call listen when the user wants to dictate the next prompt.
  • listen prints recognized text to stdout.

Mode D: Duplex Continuous Dialogue (saylisten)

When the user enables duplex mode (e.g. "turn on duplex", "full voice mode"):

  1. Generate the full assistant response first.
  2. Speak the full response via say.
  3. Immediately call listen -d 0 -s 1 in the same conversation language.
  4. Treat recognized text as the next user prompt.
  5. Normalize the recognized text and stop when a stop phrase intent is heard: стоп, выключи прослушивание, выключи дуплекс, stop listening.

Canonical agent loop:

answer = full assistant reply
say --lang <lang> "<answer>"
heard = listen -l <lang> -d 0 -s 1
if heard matches a stop phrase intent:
  exit duplex mode

This is a hands-free conversational flow owned by the agent, not by a dedicated shell helper. Never keep the substantive reply only in chat while sending a shorter handoff question to speech.

Mode E: Autonomous Voice Alerts (optional)

Short proactive announcements are allowed for:

  • long-running operations,
  • critical blockers/security issues,
  • required confirmation to proceed safely.

Keep alerts brief and informative.

Voice Guard + Listen Guard

Before say: ask if silence would hide important information. If not, do not speak.

Before listen: ask if voice input is actually needed right now. Do not invoke speculatively.

Language Memory

  • Preferred language is stored in ~/.pi_voice_lang.
  • Use short language codes: ru, en, de,... (not ru_RU, en_US).
  • In duplex mode, keep say and listen -l <lang> aligned.
  • say auto-downloads missing Piper model on first use.

Initialization (Linux & macOS)

Run bootstrap once:

"${SKILL_DIR}/scripts/_bootstrap"

Bootstrap installs to ~/.local/bin:

  • say
  • listen
  • listen-server

Platform Support

  • Linux: piper + aplay, faster-whisper, arecord/pyaudio
  • macOS: piper + afplay, faster-whisper, sox/pyaudio

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

34.89%
按下载量换算36

Claude

32.18%
按下载量换算33

Cursor

17.68%
按下载量换算18

Gemini CLI

9.54%
按下载量换算10

安全审计

Gen Agent Trust Hub

未通过

Socket

通过

Snyk

通过

权限和风险

执行命令

安装流程涉及命令执行,可能通过 npx skills add https://github.com/llblab/skills --skill voice-mode 联网下载 Skill 或依赖。用户安装前应确认命令来源、仓库内容和执行环境。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills