Stepfun Vision Skill

Read and understand images for the user, but ONLY when the main model is DeepSeek (model is deepseek-v4-flash or deepseek-v4-pro in config.toml). Use this skill whenever the user sends or pastes an image, attaches a screenshot, references an image file ("看下这张图", "read this image", "screenshot shows..."), or when a message contains the placeholder "image content omitted because you do not support image input". Also use it when you need to inspect image content (OCR, screenshots, diagrams, photos) but your current model cannot process image input. Do not use this skill when the main model is any other provider (e.g. gpt-5.6* on the relay, Gemini, etc.) — those models can see images directly.

jwangkun b40865f 2 files · 5.0 KB Updated

File contents

jwangkun/stepfun-vision-skill/tree/main/skills commit b40865faad

Frequently asked questions

npx skillmds@latest add jwangkun/stepfun-vision-skill