Multimodal Eye

SKILL: Multimodal Eye

Moeabdelaziz007 Updated

File contents

SKILL: Multimodal Eye

Description

Provides the agent with "Vision". Can analyze images or video frames provided via URL or local path.

Capabilities

  1. analyze_image(image_path):
    • Describes the content of an image.
    • Extracts text (OCR).
    • Identifies UI elements (for frontend coding).

Usage Prompt

"Agent, look at screenshot.png and tell me what the error message says."

Moeabdelaziz007/Gemini-3.0-Superpowers-and-Pesona/tree/main/skills/multimodal_eye commit 2b391c5950

Frequently asked questions

npx skillmds@latest add moeabdelaziz007/multimodal-eye