AI Vision

Multimodal UI understanding and single-step planning via OpenAI-compatible Responses APIs. Use when you need AIQuery/AIAssert and plan-next to extract UI element coordinates, validate UI assertions, summarize screenshots, or decide the next UI action from an image. External agents handle execution via adb/hdc and multi-step loops. Defaults to Doubao models but can be pointed at other multimodal providers via base URL, API key, and model name.

httprunner Updated

File contents

httprunner/skills/tree/main/ai-vision commit 24b7e8eebd

Frequently asked questions

npx skillmds@latest add httprunner/ai-vision