Use OmniParser for vision-based GUI parsing

Parse screenshots into structured UI elements so computer-use agents can reason about controls before acting.

agentskillexchange Updated 28 repo stars

File contents

Use OmniParser for vision-based GUI parsing

Parse screenshots into structured UI elements so computer-use agents can reason about controls before acting.

Prerequisites

Python 3.12; conda; Hugging Face model weights; optional Gradio demo

Installation

Use the upstream install or setup path that matches your environment:

  • conda create -n "omni" python==3.12
  • conda activate omni
  • pip install -r requirements.txt

Requirements and caveats from upstream:

  • python
  • python weights/convert_safetensor_to_pt.py
  • python gradio_demo.py

Basic usage or getting-started notes:

Documentation

Source

agentskillexchange/skills/tree/main/skills/use-omniparser-for-vision-based-gui-parsing commit 70c0e37f9c

Frequently asked questions

npx skillmds@latest add agentskillexchange/use-omniparser-for-vision-based-gui-parsing