Image Comprehension Ollama

If you are a vision enabled model then you do not need this skill. Use this skill to analyze image files on disk via a local vision model (default: moondream:1.8b via Ollama). Invoke it whenever you encounter an image provided as a file path — such as a screenshot, photo, diagram, chart, or scan — and need a text description of its contents. The local model processes the image and returns a description to stdout, supplementing your workflow when images are not directly viewable in chat or when a local analysis is preferred. Supports PNG, JPEG, GIF, WebP, BMP.

aosama Updated

File contents

aosama/image-comprehension-ollama/tree/main/skills/image-comprehension-ollama commit e2f374531f

Frequently asked questions

npx skillmds@latest add aosama/image-comprehension-ollama