Operational Steps
- 确认输入参数完整
- 执行核心操作(参考本目录下的 scripts/ 或 references/)
- 验证输出符合契约
- 保存结果并报告
IO_CONTRACT
- input:
request: str, context: dict— 用户请求描述、上下文信息 - output:
result: dict — 技能执行结果(结构因技能而异)
对应原则:P2(机械原子暴露输入输出规范)
HeartMuLa - Open-Source Music Generation
Overview
HeartMuLa is a family of open-source music foundation models (Apache-2.0) that generates music conditioned on lyrics and tags, with multilingual support. Generates full songs from lyrics + tags. Comparable to Suno for open-source. Includes:
- HeartMuLa - Music language model (3B/7B) for generation from lyrics + tags
- HeartCodec - 12.5Hz music codec for high-fidelity audio reconstruction
- HeartTranscriptor - Whisper-based lyrics transcription
- HeartCLAP - Audio-text alignment model
When to Use
- User wants to generate music/songs from text descriptions
- User wants an open-source Suno alternative
- User wants local/offline music generation
- User asks about HeartMuLa, heartlib, or AI music generation
Hardware Requirements
- Minimum: 8GB VRAM with
--lazy_load true(loads/unloads models sequentially) - Recommended: 16GB+ VRAM for comfortable single-GPU usage
- Multi-GPU: Use
--mula_device cuda:0 --codec_device cuda:1to split across GPUs - 3B model with lazy_load peaks at ~6.2GB VRAM
Installation Steps
1. Clone Repository
cd ~/ # or desired directory
git clone https://github.com/HeartMuLa/heartlib.git
cd heartlib
2. Create Virtual Environment (Python 3.10 required)
uv venv --python 3.10 .venv
. .venv/bin/activate
uv pip install -e .
3. Fix Dependency Compatibility Issues
IMPORTANT: As of Feb 2026, the pinned dependencies have conflicts with newer packages. Apply these fixes:
# Upgrade datasets (old version incompatible with current pyarrow)
uv pip install --upgrade datasets
# Upgrade transformers (needed for huggingface-hub 1.x compatibility)
uv pip install --upgrade transformers
4. Patch Source Code (Required for transformers 5.x)
Patch 1 - RoPE cache fix in src/heartlib/heartmula/modeling_heartmula.py:
In the setup_caches method of the HeartMuLa class, add RoPE reinitialization after the reset_caches try/except block and before the with device: block:
# Re-initialize RoPE caches that were skipped during meta-device loading
from torchtune.models.llama3_1._position_embeddings import Llama3ScaledRoPE
for module in self.modules():
if isinstance(module, Llama3ScaledRoPE) and not module.is_cache_built:
module.rope_init()
module.to(device)
Why: from_pretrained creates model on meta device first; Llama3ScaledRoPE.rope_init() skips cache building on meta tensors, then never rebuilds after weights are loaded to real device.
Patch 2 - HeartCodec loading fix in src/heartlib/pipelines/music_generation.py:
Add ignore_mismatched_sizes=True to ALL HeartCodec.from_pretrained() calls (there are 2: the eager load in __init__ and the lazy load in the codec property).
Why: VQ codebook initted buffers have shape [1] in checkpoint vs [] in model. Same data, just scalar vs 0-d tensor. Safe to ignore.
5. Download Model Checkpoints
cd heartlib # project root
hf download --local-dir './ckpt' 'HeartMuLa/HeartMuLaGen'
hf download --local-dir './ckpt/HeartMuLa-oss-3B' 'HeartMuLa/HeartMuLa-oss-3B-happy-new-year'
hf download --local-dir './ckpt/HeartCodec-oss' 'HeartMuLa/HeartCodec-oss-20260123'
All 3 can be downloaded in parallel. Total size is several GB.
GPU / CUDA
HeartMuLa uses CUDA by default (--mula_device cuda --codec_device cuda). No extra setup needed if the user has an NVIDIA GPU with PyTorch CUDA support installed.
- The installed
torch==2.4.1includes CUDA 12.1 support out of the box torchtunemay report version0.4.0+cpu— this is just package metadata, it still uses CUDA via PyTorch- To verify GPU is being used, look for "CUDA memory" lines in the output (e.g. "CUDA memory before unloading: 6.20 GB")
- No GPU? You can run on CPU with
--mula_device cpu --codec_device cpu, but expect generation to be extremely slow (potentially 30-60+ minutes for a single song vs4 minutes on GPU). CPU mode also requires significant RAM (12GB+ free). If the user has no NVIDIA GPU, recommend using a cloud GPU service (Google Colab free tier with T4, Lambda Labs, etc.) or the online demo at https://heartmula.github.io/ instead.
Usage
Basic Generation
cd heartlib
. .venv/bin/activate
python ./examples/run_music_generation.py \
--model_path=./ckpt \
--version="3B" \
--lyrics="./assets/lyrics.txt" \
--tags="./assets/tags.txt" \
--save_path="./assets/output.mp3" \
--lazy_load true
Input Formatting
Tags (comma-separated, no spaces):
piano,happy,wedding,synthesizer,romantic
or
rock,energetic,guitar,drums,male-vocal
Lyrics (use bracketed structural tags):
[Intro]
[Verse]
Your lyrics here...
[Chorus]
Chorus lyrics...
[Bridge]
Bridge lyrics...
[Outro]
Key Parameters
| Parameter | Default | Description |
|---|---|---|
--max_audio_length_ms |
240000 | Max length in ms (240s = 4 min) |
--topk |
50 | Top-k sampling |
--temperature |
1.0 | Sampling temperature |
--cfg_scale |
1.5 | Classifier-free guidance scale |
--lazy_load |
false | Load/unload models on demand (saves VRAM) |
--mula_dtype |
bfloat16 | Dtype for HeartMuLa (bf16 recommended) |
--codec_dtype |
float32 | Dtype for HeartCodec (fp32 recommended for quality) |
Performance
- RTF (Real-Time Factor) ≈ 1.0 — a 4-minute song takes ~4 minutes to generate
- Output: MP3, 48kHz stereo, 128kbps
Pitfalls
-
-
Verification
-
-
- Do NOT use bf16 for HeartCodec — degrades audio quality. Use fp32 (default).
- Tags may be ignored — known issue (#90). Lyrics tend to dominate; experiment with tag ordering.
- Triton not available on macOS — Linux/CUDA only for GPU acceleration.
- RTX 5080 incompatibility reported in upstream issues.
- The dependency pin conflicts require the manual upgrades and patches described above.
Links
- Repo: https://github.com/HeartMuLa/heartlib
- Models: https://huggingface.co/HeartMuLa
- Paper: https://arxiv.org/abs/2601.10547
- License: Apache-2.0
验证清单 · VERIFICATION
- venv 使用 Python 3.10,依赖升级(datasets/transformers)及 RoPE、HeartCodec 两处补丁均已应用
- 三个模型 checkpoint(HeartMuLaGen、HeartMuLa-oss-3B、HeartCodec-oss)已下载至
./ckpt - VRAM 满足要求(≥8GB 且
--lazy_load true,或 ≥16GB),输出日志出现 "CUDA memory" 行 - 歌词文件使用括号结构标签([Verse]/[Chorus] 等),tags 为无空格逗号分隔格式
- 生成的 MP3 输出存在且非零字节,时长符合
--max_audio_length_ms上限 - HeartCodec 使用
--codec_dtype float32(未用 bf16),音质未退化
约束规则 · RULES
- 输入约束: 参数类型、范围、格式必须校验
- 输出约束: 返回值结构、编码、命名必须一致
- 异常约束: 错误信息必须包含上下文和恢复建议
- 安全约束: 不执行未验证的任意代码,不暴露内部状态
Golden 集合 · GOLDEN SET
- Golden Input: 标准输入样本(覆盖正常路径)
- Golden Output: 预期输出(精确匹配或格式校验)
- Golden Error: 预期错误信息(覆盖失败路径)
Golden 集合是测试的单一真理来源。所有改进必须通过 golden 测试。
违反规则的操作视为不安全,必须拒绝或隔离。
每项验证必须可执行、可记录、可复现。验证失败时记录原因和修复。
Heartmula
Genes (策略基因)
紧凑策略表示。条件→策略。需要深度时参考完整文档。
- [HEAR-001] 当 VRAM 低于 16GB 或需单卡运行大模型时 → 启用
--lazy_load true以按需加载/卸载模型,将峰值显存控制在 ~6.2GB - [HEAR-002] 当使用 transformers 5.x 且模型在 meta-device 初始化时 → 手动修补 RoPE 缓存,在权重加载后重新初始化
Llama3ScaledRoPE以修复位置编码 - [HEAR-003] 当加载 HeartCodec 遇到 VQ codebook 形状不匹配(scalar vs 0-d tensor)时 → 在
from_pretrained中设置ignore_mismatched_sizes=True以安全忽略缓冲区差异 - [HEAR-004] 当需要保证高保真音频重建质量时 → 强制 HeartCodec 使用
float32精度,严禁使用bfloat16以避免音质退化 - [HEAR-005] 当用户无 NVIDIA GPU 或处于 macOS 环境时 → 推荐云端 GPU 服务或在线 Demo,避免本地 CPU 运行导致的 30-60 分钟极慢生成速度
- [HEAR-006] 当输入包含歌词和风格标签时 → 采用括号结构标签(如 [Verse], [Chorus])格式化歌词,并使用无空格逗号分隔标签,同时注意标签可能被歌词主导的问题