Vision

Query images with a local Ollama vision model without loading the image into the main agent context. Use when you need to describe a screenshot, check whether rendered content is present, detect overlapping elements, or ask any visual question about a PNG/JPEG/WebP file. Requires Ollama running locally with the Gemma 4 multimodal model (`gemma4` on Ollama). Script: .agents/skills/vision/scripts/ask.py. Trigger phrases: "describe image", "what does this screenshot show", "does the canvas contain content", "check screenshot visually", "look at this image", "any overlapping elements", "vision query".

gridaco 84463bc 2 files · 15.6 KB Updated

File contents

gridaco/nothing/tree/main/.agents/skills/vision commit 84463bc887

Frequently asked questions

npx skillmds@latest add gridaco/vision