# Oatda Vision Analysis

> Analyze images, photos, screenshots, or diagrams using vision-capable AI models (GPT-4o, GPT-5, Claude Sonnet 4.5, Claude Opus 4.5, Gemini 3, GLM-4.6V) through OATDA's unified AI gateway. Triggers when the user wants image analysis, OCR, photo understanding, chart reading, screenshot description, or computer-vision tasks via a single API key.

- Skill: `devcsde/oatda-vision-analysis-2` (Agent Skill)
- Install (CLI): `npx skillmds@latest add devcsde/oatda-vision-analysis-2`
- Raw SKILL.md: https://api.skillmd.com/api/skills/devcsde/oatda-vision-analysis-2/raw
- Safety review: pending (external: skill-scanner PASS, skillspector CAUTION)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Integrations & APIs
- Author: devcsde (https://skillmd.com/u/devcsde)
- Updated: 2026-09-22
- Page: https://skillmd.com/skills/devcsde/oatda-vision-analysis-2

---


# OATDA Vision Analysis

Analyze images using vision-capable AI models through OATDA's unified API.

## API Key Resolution

All commands need the OATDA API key. Resolve it inline for each `exec` call:

```bash
export OATDA_API_KEY="${OATDA_API_KEY:-$(cat ~/.oatda/credentials.json 2>/dev/null | jq -r '.profiles[.defaultProfile].apiKey' 2>/dev/null)}"
```

If the key is empty or `null`, tell the user to get one at https://oatda.com and configure it.

**Security**: Never print the full API key. Only verify existence or show first 8 chars.

## Model Mapping

| User says | Provider | Model |
|-----------|----------|-------|
| gpt-4o (default) | openai | gpt-4o |
| gpt-4o-mini | openai | gpt-4o-mini |
| gpt-5 | openai | gpt-5 |
| claude, sonnet | anthropic | claude-sonnet-4-5-20250929 |
| opus | anthropic | claude-opus-4-5-20251101 |
| gemini | google | gemini-3-pro-preview |
| gemini-2.5 | google | gemini-2.5-pro |
| glm-v | zai | glm-4.6v |

**Default**: `openai` / `gpt-4o` if no model specified.

> ⚠️ Models update frequently. Run `oatda-list-models` with `type=vision` for the latest vision models.

> ⚠️ Models update frequently. If a model ID fails, query `oatda-list-models` with `?type=chat` for the latest vision-capable models.

## Image URL Validation

- **Accept**: `https://` URLs or `data:image/` base64 data URIs
- **Reject**: `http://` URLs, local file paths, internal IPs (localhost, 127.0.0.1, 169.254.x.x)
- If user provides a local file, suggest converting to base64 first

## API Call

**CRITICAL**: The canonical endpoint is `/api/v1/llm` (NOT `/api/v1/llm/generate-image` — that's for image generation). The body uses a `contents` array, NOT a simple `prompt` string. The legacy `/api/v1/llm/image` alias is deprecated but accepts the same body for backward compatibility.

```bash
export OATDA_API_KEY="${OATDA_API_KEY:-$(cat ~/.oatda/credentials.json 2>/dev/null | jq -r '.profiles[.defaultProfile].apiKey' 2>/dev/null)}" && \
curl -s -X POST "https://oatda.com/api/v1/llm" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OATDA_API_KEY" \
  -d '{
    "provider": "<PROVIDER>",
    "model": "<MODEL>",
    "contents": [
      {"type": "text", "text": "<ANALYSIS_PROMPT>"},
      {"type": "image", "image": {"url": "<IMAGE_URL>", "detail": "auto"}}
    ]
  }'
```

### Optional Parameters (add to body)

- `temperature`: 0-2, default 0.7
- `maxTokens`: Max response tokens

### Image Detail Levels

- `"auto"` — Let the model decide (default)
- `"low"` — Faster, cheaper, less detail
- `"high"` — More detail, higher cost (recommended for OCR)

## Response Format

```json
{
  "success": true,
  "provider": "openai",
  "model": "gpt-4o",
  "response": "The image shows a sunset over...",
  "usage": {
    "promptTokens": 800,
    "completionTokens": 200,
    "totalTokens": 1000
  },
  "costs": {
    "inputCost": 0.004,
    "outputCost": 0.006,
    "totalCost": 0.01,
    "currency": "USD"
  }
}
```

Present the `response` field to the user. Optionally mention token usage and cost.

## Error Handling

| HTTP Status | Meaning | Action |
|-------------|---------|--------|
| 401 | Invalid API key | Tell user to check their key |
| 400 | Bad request | Check image URL is valid HTTPS, model supports vision |
| 429 | Rate limited | Wait 5 seconds and retry once |

## Example

User: "Describe this image: https://example.com/photo.jpg"

```bash
export OATDA_API_KEY="${OATDA_API_KEY:-$(cat ~/.oatda/credentials.json 2>/dev/null | jq -r '.profiles[.defaultProfile].apiKey' 2>/dev/null)}" && \
curl -s -X POST "https://oatda.com/api/v1/llm" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OATDA_API_KEY" \
  -d '{
    "provider": "openai",
    "model": "gpt-4o",
    "contents": [
      {"type": "text", "text": "Describe this image in detail"},
      {"type": "image", "image": {"url": "https://example.com/photo.jpg", "detail": "auto"}}
    ]
  }'
```

## Notes

- Canonical endpoint is `/api/v1/llm` — NOT `/api/v1/llm/generate-image` (that's for generation). Legacy `/api/v1/llm/image` alias is deprecated but accepts the same body.
- Body uses `contents` array format, NOT a simple prompt string
- Only HTTPS image URLs accepted — no HTTP, no local paths
- Image tokens are included in prompt token count and affect cost
- For OCR tasks, use `"detail": "high"`
- Use `oatda-generate-image` for creating images
- Use `oatda-list-models` for available vision models

