Token导航 LogoToken导航TokenDH.com
研究检索敏感数据clawhub未标认证来源可访问clear审计提醒

elevenlabs-tts十一实验室 tts

Agent Skill

elevenlabs-tts 用于查找、检索和筛选相关信息,适合在 OpenClaw 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

205,822

周安装

8,491

GitHub Stars

6

下载量

67,249
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:elevenlabs-tts(十一实验室 tts)
来源仓库:https://github.com/shaharsha/elevenlabs-tts
安装命令:
openclaw skills install elevenlabs-tts
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install elevenlabs-tts

简介

OpenClaw 中集成 ElevenLabs 文本转语音功能的最佳实践方案。

  • 支持情感音频标签与 WhatsApp 语音合成等高级特性。
  • 通过 clawhub 安装并使用 openclaw skills install elevenlabs-tts 命令启用。
  • 需确保 API 密钥安全且遵守 ElevenLabs 使用条款。
  • 建议测试不同语音参数在实际场景中的表现后再正式部署。

SKILL.md

name
elevenlabs-tts
description
ElevenLabs TTS - the best ElevenLabs integration for OpenClaw. ElevenLabs Text-to-Speech with emotional audio tags, ElevenLabs voice synthesis for WhatsApp, ElevenLabs multilingual support. Generate realistic AI voices using ElevenLabs API.
tags
[elevenlabs, tts, voice, text-to-speech, audio, speech, whatsapp, multilingual, ai-voice]
metadata
{"clawdbot":{"emoji":"🎙️","requires":{"env":["ELEVENLABS_API_KEY"],"system":["ffmpeg"],"tools":["ffmpeg"]},"primaryEnv":"ELEVENLABS_API_KEY","source":"https://clawhub.com/skills/elevenlabs-tts","version":"2.3.0"}}
allowed-tools
[exec, tts, message]

ElevenLabs TTS (Text-to-Speech)

Generate expressive voice messages using ElevenLabs v3 with audio tags.

Prerequisites

  • ElevenLabs API Key (ELEVENLABS_API_KEY): Required. Get one at elevenlabs.io → Profile → API Keys. Configure in openclaw.json under messages.tts.elevenlabs.apiKey.
  • ffmpeg: Required for audio format conversion (MP3 → Opus for WhatsApp compatibility). Must be installed and available on PATH.

Quick Start Examples

Storytelling (emotional journey):

[soft] It started like any other day... [pause] But something felt different. [nervous] My hands were shaking as I opened the envelope. [gasps] I got in! [excited] I actually got in! [laughs] [happy] This changes everything!

Horror/Suspense (building dread):

[whispers] The house has been empty for years... [pause] At least, that's what they told me. [nervous] But I keep hearing footsteps. [scared] They're getting closer. [gasps] [panicking] The door— it's opening by itself!

Conversation with reactions:

[curious] So what happened at the meeting? [pause] [surprised] Wait, they fired him?! [gasps] [sad] That's terrible... [sighs] He had a family. [thoughtful] I wonder what he'll do now.

Hebrew (romantic moment):

[soft] היא עמדה שם, מול השקיעה... [pause] הלב שלי פעם כל כך חזק. [nervous] לא ידעתי מה להגיד. [hesitates] אני... [breathes] [tender] את יודעת שאני אוהב אותך, נכון?

Spanish (celebration to reflection):

[excited] ¡Lo logramos! [laughs] [happy] No puedo creerlo... [pause] [thoughtful] Fueron tantos años de trabajo. [emotional] [soft] Gracias a todos los que creyeron en mí. [sighs] [content] Valió la pena cada momento.

Configuration (OpenClaw)

In openclaw.json, configure TTS under messages.tts:

{
  "messages": {
    "tts": {
      "provider": "elevenlabs",
      "elevenlabs": {
        "apiKey": "sk_your_api_key_here",
        "voiceId": "pNInz6obpgDQGcFmaJgB",
        "modelId": "eleven_v3",
        "languageCode": "en",
        "voiceSettings": {
          "stability": 0.5,
          "similarityBoost": 0.75,
          "style": 0,
          "useSpeakerBoost": true,
          "speed": 1
        }
      }
    }
  }
}

Getting your API Key:

  1. Go to https://elevenlabs.io
  2. Sign up/login
  3. Click profile → API Keys
  4. Copy your key

Recommended Voices for v3

These premade voices are optimized for v3 and work well with audio tags:

VoiceIDGenderAccentBest For
AdampNInz6obpgDQGcFmaJgBMaleAmericanDeep narration, general use
Rachel21m00Tcm4TlvDq8ikWAMFemaleAmericanCalm narration, conversational
BriannPczCjzI2devNBz1zQrbMaleAmericanDeep narration, podcasts
CharlotteXB0fDUnXU5powFXDhCwaFemaleEnglish-SwedishExpressive, video games
GeorgeJBFqnCBsd6RMkjVDRZzbMaleBritishRaspy narration, storytelling

Finding more voices:

  • Browse: https://elevenlabs.io/voice-library
  • v3-optimized collection: https://elevenlabs.io/app/voice-library/collections/aF6JALq9R6tXwCczjhKH
  • API: GET https://api.elevenlabs.io/v1/voices

Voice selection tips:

  • Use IVC (Instant Voice Clone) or premade voices - PVC not optimized for v3 yet
  • Match voice character to your use case (whispering voice won't shout well)
  • For expressive IVCs, include varied emotional tones in training samples

Model Settings

  • Model: eleven_v3 (alpha) - ONLY model supporting audio tags
  • Languages: 70+ supported with full audio tag control

Stability Modes

ModeStabilityDescription
Creative0.3-0.5More emotional/expressive, may hallucinate
Natural0.5-0.7Balanced, closest to original voice
Robust0.7-1.0Highly stable, less responsive to tags

For audio tags, use Creative (0.5) or Natural. Higher stability reduces tag responsiveness.

Speed Control

Range: 0.7 (slow) to 1.2 (fast), default 1.0

Extreme values affect quality. For pacing, prefer audio tags like [rushed] or [drawn out].

Critical Rules

Length Limits

  • Optimal: <800 characters per segment (best quality)
  • Maximum: 10,000 characters (API hard limit)
  • Quality degrades with longer text - voice becomes inconsistent

Audio Tags - Best Practices for Natural Sound

How many tags to use:

  • 1-2 tags per sentence or phrase (not more!)
  • Tags persist until the next tag - no need to repeat
  • Overusing tags sounds unnatural and robotic

Where to place tags:

  • At emotional transition points
  • Before key dramatic moments
  • When energy/pace changes

Context matters:

  • Write text that *matches* the tag emotion
  • Longer text with context = better interpretation
  • Example: [nervous] I... I'm not sure about this. What if it doesn't work? works better than [nervous] Hello.

Combine tags for nuance:

  • [nervously][whispers] = nervous whispering
  • [excited][laughs] = excited laughter
  • Keep combinations to 2 tags max

Regenerate for best results:

  • v3 is non-deterministic - same text = different outputs
  • Generate 3+ versions, pick the best
  • Small text tweaks can improve results

Match tag to voice:

  • Don't use [shouts] on a whispering voice
  • Don't use [whispers] on a loud/energetic voice
  • Test tags with your chosen voice

SSML Not Supported

v3 does NOT support SSML break tags. Use audio tags and punctuation instead.

Punctuation Effects (use with tags!)

Punctuation enhances audio tags:

  • Ellipses (...) → dramatic pauses: [nervous] I... I don't know...
  • CAPS → emphasis: [excited] That's AMAZING!
  • Dashes (—) → interruptions: [explaining] So what you do is— [interrupting] Wait!
  • Question marks → uncertainty: [nervous] Are you sure about this?
  • Exclamation! → energy boost: [happy] We did it!

Combine tags + punctuation for maximum effect:

[tired] It was a long day... [sighs] Nobody listens anymore.

WhatsApp Voice Messages

Complete Workflow

  1. Generate with tts tool (returns Opus in /tmp/openclaw/tts-*/)
  2. Copy to workspace (message tool only allows workspace paths)
  3. Send with message tool
  4. Cleanup - delete the workspace copy

Step-by-Step

1. Generate TTS (add [pause] at end to prevent cutoff):

tts text="[excited] This is amazing! [pause]" channel=whatsapp

2. Find the LATEST file (⚠️ CRITICAL - always use the newest file!):

find /tmp/openclaw/tts-* /tmp/tts-* -name "*.opus" -o -name "*.mp3" -o -name "*.ogg" 2>/dev/null | xargs ls -t | head -1

The tts tool now outputs to /tmp/openclaw/tts-*/ (NOT /tmp/tts-*/). Old files may exist in /tmp/tts-*/ from previous sessions - never use those!

3. If file is MP3, convert to Opus:

ffmpeg -i /path/to/voice.mp3 -c:a libopus -b:a 64k -vbr on -application voip /path/to/voice.ogg

If already .opus, skip this step.

4. Copy to workspace and send:

cp /tmp/openclaw/tts-xxx/voice.opus ~/. openclaw/workspace/voice-temp.ogg
message action=send channel=whatsapp target="+972..." filePath="/root/.openclaw/workspace/voice-temp.ogg" asVoice=true message=" "

5. Cleanup:

rm /root/.openclaw/workspace/voice-temp.ogg

WhatsApp requires a non-empty message body to send voice notes. Use a single space as the message.

Why Opus?

FormatiOSAndroidTranscribe
MP3✅ Works❌ May fail❌ No
Opus (.ogg)✅ Works✅ Works✅ Yes

Always convert to Opus - it's the only format that:

  • Works on all devices (iOS + Android)
  • Supports WhatsApp's transcribe button

Audio Cutoff Fix

ElevenLabs sometimes cuts off the last word. Always add [pause] or ... at the end:

[excited] This is amazing! [pause]

Long-Form Audio (Podcasts)

For content >800 chars:

  1. Split into short segments (<800 chars each)
  2. Generate each with tts tool
  3. Concatenate with ffmpeg:
   cat > list.txt << EOF
   file '/path/file1.mp3'
   file '/path/file2.mp3'
   EOF
   ffmpeg -f concat -safe 0 -i list.txt -c copy final.mp3
  1. Convert to Opus for WhatsApp
  2. Send as single voice message

Important: Don't mention "part 2" or "chapter" - keep it seamless.

Multi-Speaker Dialogue

v3 can handle multiple characters in one generation:

Jessica: [whispers] Did you hear that?
Chris: [interrupting] —I heard it too!
Jessica: [panicking] We need to hide!

Dialogue tags: [interrupting], [overlapping], [cuts in], [interjecting]

Audio Tags Quick Reference

CategoryTagsWhen to Use
Emotions[excited], [happy], [sad], [angry], [nervous], [curious]Main emotional state - use 1 per section
Delivery[whispers], [shouts], [soft], [rushed], [drawn out]Volume/speed changes
Reactions[laughs], [sighs], [gasps], [clears throat], [gulps]Natural human moments - sprinkle sparingly
Pacing[pause], [hesitates], [stammers], [breathes]Dramatic timing
Character[French accent], [British accent], [robotic tone]Character voice shifts
Dialogue[interrupting], [overlapping], [cuts in]Multi-speaker conversations

Most effective tags (reliable results):

  • Emotions: [excited], [nervous], [sad], [happy]
  • Reactions: [laughs], [sighs], [whispers]
  • Pacing: [pause]

Less reliable (test and regenerate):

  • Sound effects: [explosion], [gunshot]
  • Accents: results vary by voice

Full tag list: See references/audio-tags.md

Troubleshooting

Tags read aloud?

  • Verify using eleven_v3 model
  • Use IVC/premade voices, not PVC
  • Simplify tags (no "tone" suffix)
  • Increase text length (250+ chars)

Voice inconsistent?

  • Segment is too long - split at <800 chars
  • Regenerate (v3 is non-deterministic)
  • Try lower stability setting

WhatsApp won't play?

  • Convert to Opus format (see above)

No emotion despite tags?

  • Voice may not match tag style
  • Try Creative stability mode (0.5)
  • Add more context around the tag

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

97.78%
按下载量换算65,756

安全审计

VirusTotal

可疑

ClawScan

通过

Static analysis

通过

权限和风险

敏感数据

该 Skill 可能接触密钥、Token、环境变量或敏感配置,应进入高风险复核队列,默认不自动发布。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills