# Image

> Extract text from images using a vision LLM

- Skill: `gabrielmoreira/image-3` (Agent Skill)
- Install (CLI): `npx skillmds@latest add gabrielmoreira/image-3`
- Raw SKILL.md: https://api.skillmd.com/api/skills/gabrielmoreira/image-3/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- License: AGPL-3.0-or-later
- Author: gabrielmoreira (https://skillmd.com/u/gabrielmoreira)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/gabrielmoreira/image-3

---


# Image Skill

Base64-encodes the image and passes it to a vision-capable LLM that extracts
all text and key information. Returns the LLM's response as `result.text`.

## Setup

No pip dependency — the skill uses only the Python standard library plus a
LLM provider you supply at construction time. The provider can be any object
that implements the `complete()` interface (see below).

## Standalone usage

```python
import asyncio
from synthadoc.skills.image.scripts.main import ImageSkill

# ImageSkill REQUIRES a vision-capable provider — calling extract() without
# one raises ValueError immediately.
skill = ImageSkill(provider=my_provider)

async def main():
    result = await skill.extract("/path/to/screenshot.png")
    print(result.text)          # extracted text from the image
    print(result.metadata)      # {"tokens_input": N, "tokens_output": N}

asyncio.run(main())
```

**Provider interface** — any object with this async method:

```python
async def complete(
    messages: list,             # list of Message objects from synthadoc.skills.base
    system: str | None = None,
    temperature: float = 0.0,
    max_tokens: int = 4096,
) -> object                     # must have .text (str), .input_tokens (int), .output_tokens (int)
```

Build the provider with any vision-capable model. `Message` is importable
from `synthadoc.skills.base` — no dependency on `synthadoc.providers`:

```python
from synthadoc.skills.base import Message
```

**Supported image formats:** `.png`, `.jpg`/`.jpeg`, `.webp`, `.gif`, `.tiff`

## When this skill is used

- Source path ends with `.png`, `.jpg`, `.jpeg`, `.webp`, `.gif`, or `.tiff`
- User intent contains: `image`, `screenshot`, `diagram`, `photo`

