Tuna Unified Visual Multimodal

Cascaded VAE+SigLIP encoders creating single continuous representation space supporting both vision understanding and generation, trained jointly on both tasks without format mismatches. Deploy for unified multimodal models where understanding and generation enhance each other.

adu2021 Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/tuna-unified-visual-multimodal commit 6c750b3d3b

Frequently asked questions

npx skillmds@latest add adu2021/tuna-unified-visual-multimodal