multimodal page grounding

Use this skill when the user wants tasks where the page layout, screenshot, button position, visual state, or on-screen cues actually matter. Trigger it for requests like “make it rely on what’s visible on the page,” “the agent should need the screenshot,” “buttons and layout should matter,” or “don’t let text alone be enough.” It is the right skill for browser tasks where perception must combine HTML-like structure with visual grounding.

dingxingdi Updated

File contents

dingxingdi/paper_fast_search_backup/tree/main/examples/evol_ability/20260325_170549/profiles/web/skills/multimodal-page-grounding commit 3f85aefaaf

Frequently asked questions

npx skillmds@latest add dingxingdi/multimodal-page-grounding-2