Vision Bot
Analyze images for detailed descriptions, object detection, and OCR text extraction. Accepts images via URL or base64. Auto-detects the right mode from your task — OCR for text extraction, counting for quantity questions, or full description by default.
When to Use
- Describing image contents for accessibility
- Extracting text from screenshots, signs, or photos (OCR)
- Counting objects in images
- Identifying objects in images
- Analyzing charts, diagrams, or visual data
Usage Flow
- Provide an
image_url (JPEG, PNG, GIF, WebP) or image_base64 encoded image
- Optionally specify a
task — mention "read", "OCR", or "license plate" for text extraction; "count" or "how many" for counting mode
- AIProx routes to the vision-bot agent
- Returns description, objects array, extracted text, and detected mode
Security Manifest
| Permission |
Scope |
Reason |
| Network |
aiprox.dev |
API calls to orchestration endpoint |
| Env Read |
AIPROX_SPEND_TOKEN |
Authentication for paid API |
Make Request
curl -X POST https://aiprox.dev/api/orchestrate \
-H "Content-Type: application/json" \
-H "X-Spend-Token: $AIPROX_SPEND_TOKEN" \
-d '{
"task": "extract all text from this image",
"image_url": "https://example.com/photo.jpg"
}'
Response
{
"description": "A modern office workspace with a standing desk and dual monitors.",
"objects": ["desk", "monitors", "keyboard", "mouse", "plant", "window", "headphones"],
"text_found": "Visual Studio Code - main.js",
"mode": "ocr"
}
Trust Statement
Vision Bot fetches and analyzes images via URL or base64 input. Images are processed transiently using Claude's vision capabilities via LightningProx. No images are stored. Your spend token is used for payment only.
1---2name: vision-bot3description: Analyze images via URL or base64. Auto-detects mode: OCR, object counting, or full description.4---56# Vision Bot78Analyze images for detailed descriptions, object detection, and OCR text extraction. Accepts images via URL or base64. Auto-detects the right mode from your task — OCR for text extraction, counting for quantity questions, or full description by default.910## When to Use1112- Describing image contents for accessibility13- Extracting text from screenshots, signs, or photos (OCR)14- Counting objects in images15- Identifying objects in images16- Analyzing charts, diagrams, or visual data1718## Usage Flow19201. Provide an `image_url` (JPEG, PNG, GIF, WebP) **or** `image_base64` encoded image212. Optionally specify a `task` — mention "read", "OCR", or "license plate" for text extraction; "count" or "how many" for counting mode223. AIProx routes to the vision-bot agent234. Returns description, objects array, extracted text, and detected mode2425## Security Manifest2627| Permission | Scope | Reason |28|------------|-------|--------|29| Network | aiprox.dev | API calls to orchestration endpoint |30| Env Read | AIPROX_SPEND_TOKEN | Authentication for paid API |3132## Make Request3334```bash35curl -X POST https://aiprox.dev/api/orchestrate \36 -H "Content-Type: application/json" \37 -H "X-Spend-Token: $AIPROX_SPEND_TOKEN" \38 -d '{39 "task": "extract all text from this image",40 "image_url": "https://example.com/photo.jpg"41 }'42```4344### Response4546```json47{48 "description": "A modern office workspace with a standing desk and dual monitors.",49 "objects": ["desk", "monitors", "keyboard", "mouse", "plant", "window", "headphones"],50 "text_found": "Visual Studio Code - main.js",51 "mode": "ocr"52}53```5455## Trust Statement5657Vision Bot fetches and analyzes images via URL or base64 input. Images are processed transiently using Claude's vision capabilities via LightningProx. No images are stored. Your spend token is used for payment only.