SenseAudio Language Tutor
Create interactive language-learning audio with official SenseAudio TTS endpoints and parameters.
What This Skill Does
- Generate pronunciation examples in supported voices
- Create bilingual vocabulary and sentence practice audio
- Produce slowed-speed listening drills for learners
- Build short dialogue exercises with repetition pauses
- Export lesson audio files and companion study notes
Credential and Dependency Rules
- Read the API key from
SENSEAUDIO_API_KEY.
- Send auth only as
Authorization: Bearer <API_KEY>.
- Do not place API keys in query parameters, logs, or saved examples.
- If Python helpers are used, this skill expects
python3, requests, and pydub.
pydub may also require a local audio backend such as ffmpeg; if unavailable, prefer writing individual audio files instead of merging them.
Official TTS Constraints
Use the official SenseAudio TTS rules summarized below:
- HTTP endpoint:
POST https://api.senseaudio.cn/v1/t2a_v2
- Model:
SenseAudio-TTS-1.0
- Max text length:
10000 characters
voice_setting.voice_id is required
voice_setting.speed range: 0.5-2.0
- Optional audio format values:
mp3, wav, pcm, flac
- Optional sample rates:
8000, 16000, 22050, 24000, 32000, 44100
- Optional MP3 bitrates:
32000, 64000, 128000, 256000
- Optional channels:
1 or 2
Recommended Workflow
- Prepare lesson content:
- Split vocabulary, example sentences, and dialogues into short chunks.
- Keep each API call comfortably below the
10000 character limit.
- Build minimal TTS requests:
- Send
model, text, stream, and voice_setting.voice_id.
- Add
speed, pitch, vol, and audio_setting only when needed.
- Decode and save audio safely:
- HTTP responses return hex-encoded audio in
data.audio; decode before saving.
- Keep filenames deterministic and avoid exposing secrets in paths or logs.
- Compose lessons carefully:
- If
pydub and an audio backend are available, merge clips and insert silence.
- Otherwise, emit per-word or per-sentence clips and a manifest/Markdown study guide.
- Handle failures and traceability:
- Check HTTP status and provider error payloads before decoding audio.
- Record
trace_id only for troubleshooting and avoid showing it unless needed.
Minimal Helper
import binascii
import os
import requests
API_KEY = os.environ["SENSEAUDIO_API_KEY"]
API_URL = "https://api.senseaudio.cn/v1/t2a_v2"
def generate_tts(text, voice_id="male_0004_a", speed=1.0, stream=False):
payload = {
"model": "SenseAudio-TTS-1.0",
"text": text,
"stream": stream,
"voice_setting": {
"voice_id": voice_id,
"speed": speed,
},
"audio_setting": {
"format": "mp3",
"sample_rate": 32000,
"bitrate": 128000,
"channel": 2,
},
}
response = requests.post(
API_URL,
headers={
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
},
json=payload,
timeout=60,
)
response.raise_for_status()
data = response.json()
audio_hex = data["data"]["audio"]
return binascii.unhexlify(audio_hex), data.get("trace_id")
Patterns
Vocabulary Drill
- Generate one clip for the target word
- Generate one clip for an example sentence
- Optionally generate a slower clip at
speed=0.8
- Save clips separately or merge with pauses
Bilingual Lesson
- Alternate source phrase and translated phrase
- Use short pauses (
1000-2000ms) between clips
- Consider different
voice_id values for source and translation when helpful
Dialogue Practice
- Create one clip per line of dialogue
- Insert repetition pauses after each line
- Prefer shorter turns for easier debugging and regeneration
Output Options
- Individual MP3 clips for words, sentences, or dialogue turns
- Merged lesson audio if local audio tooling is available
- Markdown study guide with transcript, translation, and file manifest
Safety Notes
- Do not hardcode credentials.
- Do not claim unsupported language-selection parameters for TTS unless the official docs add them.
- Avoid assuming raw bytes can be passed directly to
pydub.AudioSegment; decode and load through a supported container format.
1---2name: senseaudio-language-tutor3description: Create language learning audio with SenseAudio TTS, including pronunciation drills, bilingual lessons, slowed speech practice, and dialogue exercises. Use when users need language learning materials, pronunciation practice, or vocabulary audio.4---56# SenseAudio Language Tutor78Create interactive language-learning audio with official SenseAudio TTS endpoints and parameters.910## What This Skill Does1112- Generate pronunciation examples in supported voices13- Create bilingual vocabulary and sentence practice audio14- Produce slowed-speed listening drills for learners15- Build short dialogue exercises with repetition pauses16- Export lesson audio files and companion study notes1718## Credential and Dependency Rules1920- Read the API key from `SENSEAUDIO_API_KEY`.21- Send auth only as `Authorization: Bearer <API_KEY>`.22- Do not place API keys in query parameters, logs, or saved examples.23- If Python helpers are used, this skill expects `python3`, `requests`, and `pydub`.24- `pydub` may also require a local audio backend such as `ffmpeg`; if unavailable, prefer writing individual audio files instead of merging them.2526## Official TTS Constraints2728Use the official SenseAudio TTS rules summarized below:2930- HTTP endpoint: `POST https://api.senseaudio.cn/v1/t2a_v2`31- Model: `SenseAudio-TTS-1.0`32- Max text length: `10000` characters33- `voice_setting.voice_id` is required34- `voice_setting.speed` range: `0.5-2.0`35- Optional audio format values: `mp3`, `wav`, `pcm`, `flac`36- Optional sample rates: `8000`, `16000`, `22050`, `24000`, `32000`, `44100`37- Optional MP3 bitrates: `32000`, `64000`, `128000`, `256000`38- Optional channels: `1` or `2`3940## Recommended Workflow41421. Prepare lesson content:43- Split vocabulary, example sentences, and dialogues into short chunks.44- Keep each API call comfortably below the `10000` character limit.45462. Build minimal TTS requests:47- Send `model`, `text`, `stream`, and `voice_setting.voice_id`.48- Add `speed`, `pitch`, `vol`, and `audio_setting` only when needed.49503. Decode and save audio safely:51- HTTP responses return hex-encoded audio in `data.audio`; decode before saving.52- Keep filenames deterministic and avoid exposing secrets in paths or logs.53544. Compose lessons carefully:55- If `pydub` and an audio backend are available, merge clips and insert silence.56- Otherwise, emit per-word or per-sentence clips and a manifest/Markdown study guide.57585. Handle failures and traceability:59- Check HTTP status and provider error payloads before decoding audio.60- Record `trace_id` only for troubleshooting and avoid showing it unless needed.6162## Minimal Helper6364```python65import binascii66import os6768import requests6970API_KEY = os.environ["SENSEAUDIO_API_KEY"]71API_URL = "https://api.senseaudio.cn/v1/t2a_v2"727374def generate_tts(text, voice_id="male_0004_a", speed=1.0, stream=False):75 payload = {76 "model": "SenseAudio-TTS-1.0",77 "text": text,78 "stream": stream,79 "voice_setting": {80 "voice_id": voice_id,81 "speed": speed,82 },83 "audio_setting": {84 "format": "mp3",85 "sample_rate": 32000,86 "bitrate": 128000,87 "channel": 2,88 },89 }90 response = requests.post(91 API_URL,92 headers={93 "Authorization": f"Bearer {API_KEY}",94 "Content-Type": "application/json",95 },96 json=payload,97 timeout=60,98 )99 response.raise_for_status()100 data = response.json()101 audio_hex = data["data"]["audio"]102 return binascii.unhexlify(audio_hex), data.get("trace_id")103```104105## Patterns106107### Vocabulary Drill108109- Generate one clip for the target word110- Generate one clip for an example sentence111- Optionally generate a slower clip at `speed=0.8`112- Save clips separately or merge with pauses113114### Bilingual Lesson115116- Alternate source phrase and translated phrase117- Use short pauses (`1000-2000ms`) between clips118- Consider different `voice_id` values for source and translation when helpful119120### Dialogue Practice121122- Create one clip per line of dialogue123- Insert repetition pauses after each line124- Prefer shorter turns for easier debugging and regeneration125126## Output Options127128- Individual MP3 clips for words, sentences, or dialogue turns129- Merged lesson audio if local audio tooling is available130- Markdown study guide with transcript, translation, and file manifest131132## Safety Notes133134- Do not hardcode credentials.135- Do not claim unsupported language-selection parameters for TTS unless the official docs add them.136- Avoid assuming raw bytes can be passed directly to `pydub.AudioSegment`; decode and load through a supported container format.