Vision Sft

Fine-tune vision-language models (VLMs) with supervised learning on image+text data. Use when adapting a VLM to a visual domain or task, configuring frozen-vision-tower LoRA, or debugging a VLM fine-tune that trains without learning.

wshobson d47610e 2 files · 13.9 KB Updated

File contents

wshobson/agents commit d47610e602

Frequently asked questions

npx skillmds@latest add wshobson/vision-sft