# Use OmniParser for vision-based GUI parsing

> Parse screenshots into structured UI elements so computer-use agents can reason about controls before acting.

- Skill: `agentskillexchange/use-omniparser-for-vision-based-gui-parsing` (Agent Skill)
- Install (CLI): `npx skillmds@latest add agentskillexchange/use-omniparser-for-vision-based-gui-parsing`
- Raw SKILL.md: https://api.skillmd.com/api/skills/agentskillexchange/use-omniparser-for-vision-based-gui-parsing/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: agentskillexchange (https://skillmd.com/u/agentskillexchange)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/agentskillexchange/use-omniparser-for-vision-based-gui-parsing

---


# Use OmniParser for vision-based GUI parsing

Parse screenshots into structured UI elements so computer-use agents can reason about controls before acting.

## Prerequisites

Python 3.12; conda; Hugging Face model weights; optional Gradio demo

## Installation

Use the upstream install or setup path that matches your environment:
- conda create -n "omni" python==3.12
- conda activate omni
- pip install -r requirements.txt

Requirements and caveats from upstream:
- python
- python weights/convert_safetensor_to_pt.py
- python gradio_demo.py

Basic usage or getting-started notes:
- First clone the repo, and then install environment:
- cd OmniParser
- Ensure you have the V2 weights downloaded in weights folder (ensure caption weights folder is called icon_caption_florence). If not download them with:

- Source: https://github.com/microsoft/OmniParser
- Extracted from upstream docs: https://raw.githubusercontent.com/microsoft/OmniParser/HEAD/README.md

## Documentation

- https://microsoft.github.io/OmniParser/

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/use-omniparser-for-vision-based-gui-parsing/)

