Minimax M3 Multimodal Input

How to use MiniMax M3's native multimodal input (image, video) for grounded decisions in coding work. Covers reading attached images/frames, treating them as ground truth for visual claims, screenshot diffing, design parity from mockups, and routing visual evidence through reports and PRs. Load when the user attaches an image, screenshot, mockup, frame, or short clip — or when the task involves "make it look like this", "match this design", "why does this UI look wrong", or "read this error screenshot". Use when this capability is needed.

tomevault-io Updated

File contents

tomevault-io/skills-registry/tree/main/madebyaris--advance-minimax-m2-cursor-rules--advance-minimax-m2-cursor-rules commit 192c8ec2f0

Frequently asked questions

npx skillmds@latest add tomevault-io/minimax-m3-multimodal-input